Skip to content

Add support for storage element depopulation - #1394

Open
blktests-ci-kpd[bot] wants to merge 9 commits into
for-next_basefrom
series/1179504=>for-next
Open

blktests-ci-kpd[bot] wants to merge 9 commits into
for-next_basefrom
series/1179504=>for-next

Conversation

@blktests-ci-kpd

Copy link
Copy Markdown

Pull request for series with
subject: Add support for storage element depopulation
version: 1
url: https://patchwork.kernel.org/series/1179504/

@blktests-ci-kpd

Copy link
Copy Markdown
Author

Upstream branch: 67cd826
series: https://patchwork.kernel.org/series/1179504/
version: 1

@blktests-ci-kpd

Copy link
Copy Markdown
Author

Upstream branch: 67cd826
series: https://patchwork.kernel.org/series/1179504/
version: 1

@blktests-ci-kpd
blktests-ci-kpd Bot force-pushed the series/1179504=>for-next branch from a035e95 to 007956d Compare October 5, 2026 10:52
@blktests-ci-kpd

Copy link
Copy Markdown
Author

Upstream branch: 67cd826
series: https://patchwork.kernel.org/series/1179504/
version: 1

@blktests-ci-kpd

Copy link
Copy Markdown
Author

Upstream branch: 0f7ca97
series: https://patchwork.kernel.org/series/1179504/
version: 1

@blktests-ci-kpd

Copy link
Copy Markdown
Author

Upstream branch: c71f470
series: https://patchwork.kernel.org/series/1179504/
version: 1

@blktests-ci-kpd
blktests-ci-kpd Bot force-pushed the series/1179504=>for-next branch from 4bb3509 to d1814e3 Compare October 6, 2026 12:47
@blktests-ci-kpd

Copy link
Copy Markdown
Author

Upstream branch: c71f470
series: https://patchwork.kernel.org/series/1180291/
version: 2

@blktests-ci-kpd
blktests-ci-kpd Bot force-pushed the series/1179504=>for-next branch from d1814e3 to 5c01234 Compare October 6, 2026 12:50
@blktests-ci-kpd

Copy link
Copy Markdown
Author

Upstream branch: c71f470
series: https://patchwork.kernel.org/series/1180719/
version: 3

@blktests-ci-kpd blktests-ci-kpd Bot added V3 and removed V2 labels Oct 7, 2026
@blktests-ci-kpd
blktests-ci-kpd Bot force-pushed the series/1179504=>for-next branch from 5c01234 to ee8e746 Compare October 7, 2026 08:30
@blktests-ci-kpd

Copy link
Copy Markdown
Author

Upstream branch: 1f50498
series: https://patchwork.kernel.org/series/1180719/
version: 3

@blktests-ci-kpd
blktests-ci-kpd Bot force-pushed the series/1179504=>for-next branch from ee8e746 to 9a73c18 Compare October 7, 2026 23:26
@blktests-ci-kpd

Copy link
Copy Markdown
Author

Upstream branch: 1f50498
series: https://patchwork.kernel.org/series/1182229/
version: 4

@blktests-ci-kpd blktests-ci-kpd Bot removed the V3 label Oct 9, 2026
Introduce the helper function disk_zone_is_offline() for testing if a zone
of a gendisk has the offline condition. The zone to test is identified
using a sector number that must belong to the target zone to check.

Signed-off-by: Damien Le Moal <dlemoal@kernel.org>
Split out most of the code of blk_revalidate_disk_zones() into the
internal function disk_revalidate_zones(). This will allow calling this
new helper together with other code under the zone revalidation mutex.
No functional change.

Signed-off-by: Damien Le Moal <dlemoal@kernel.org>
Recent SCSI (SBC) and ATA (ACS) standards define the storage element
depopulation feature. This feature is intended for managing hard-disks
heads, either combined read-write heads or pairs of read and write heads,
allowing to keep disks with defective heads longer in production by
allowing "depopulating" (removing) defective heads.

This feature comes in two different flavors:
 - A destructive version which removes a head and reformats the disk at a
   lower capacity point, restoaring a fully functional contiguous LBA
   address space.
 - A data preserving version restricted to host-managed zoned disks, which
   marks the zones served by a removed head as offline or read-only.

In preparation for using the data-preserving flavor of the depopulation
feature in file systems natively supporting zoned block devices, introduce
a set of storage element management operations and functions to define
generic calls into block device drivers for managing storage elements.

The set of operations is defined with struct blk_se_ops and includes three
operations:
 - report_elements: get information on a device storage elements state
   reported with the new struct blk_se.
 - remove_element: depopulate a defective storage element
 - restore_elements: restore depopulated storage elements

Each operation is called from the functions
bdev_report_storage_elements(), bdev_remove_storage_element() and
bdev_restore_storage_elements().

Removing a healthy storage element from a device is possible and useful
for testing. The storage element restoration operation restore_elements
allows repopulating such healthy element. Repopulating defective storage
elements is generally not allowed by devices. The remove_element and
restore_elements operations are executed under the zone revalidation lock
so that they are serialized iand also serialized against user initiated
device revalidation, thus allowing to always have a consistent view of the
device zone conditions.

The device drivers of zoned block devices can indicate support for the
storage element depopulation feature by specifying the storage element
management operations with the se_ops field of struct
block_device_operations.

Signed-off-by: Damien Le Moal <dlemoal@kernel.org>
Define the new ioctl commands BLKGETNRSTORELEMS, BLKREPORTSTORELEMS,
BLKREMOVESTORELEM, and BLKRESTORESTORELEMS to provide users accessing
zoned block devices directly with an interface to the storage elements
management operations of the block layer.

The ioctls BLKGETNRSTORELEMS and BLKREPORTSTORELEMS are defined to
respectively get the number of storage elements of a zoned device and to
get an array of struct blk_storage_element describing the current state of
the device storage elements.

The ioctl BLKREMOVESTORELEM can be used to remove (depopulate) a storage
element that has a degraded status (i.e. BLK_SE_STS_DEGRADED).

Finally, the BLKRESTORESTORELEMS ioctl interfaces with
bdev_restore_storage_elements() to restore the depopulated storage
elements of a zoned device.

Signed-off-by: Damien Le Moal <dlemoal@kernel.org>
Emulate the storage element depopulation feature of SMR disks in the
zloop driver. This emulation is controlled using the new stor_elements
option.

This option can take several values:
 - ZLOOP_STOR_ELEMENTS_NONE (0): no emulation (default)
 - ZLOOP_STOR_ELEMENTS_RDWR (1): emulate all access storage elements
   (e.g.  read+write heads)
 - ZLOOP_STOR_ELEMENTS_PAIRS (2): emulate fractional access storage
   elements (e.g. pairs of read and write heads)

If enabled with the value 1 or 2, the number of storage elements, or of
pairs of fractional access storage elements, is automatically calculated
based on the number of zones of the device so that we have at least 2 and
at most 32 storage elements (or 2 pairs of fractional access storage
elements).

The mapping of zones to storage elements is defined simply as the modulo
of a zone number and a storage element ID. E.g, removing a storage
element from a set of 4 storage elements will offline 1 zone every 4
zones.

The storage element management operations are specified with the
zloop_se_ops (struct blk_storage_elements_ops).

Signed-off-by: Damien Le Moal <dlemoal@kernel.org>
Allow users to mark storage elements of a zloop device as degraded using
the new "degrade_element" control command. The element to degrade is
indicated using the element_id option. Example:

echo "degrade_element id=0,element_id=2" > /dev/zloop-control

If the element ID identifies an all access storage element, read and write
operations targeting a zone served by the degraded lement are failed.
For a partial access storage element, read or write operations are failed
depending on the storage element type. This check for an element status
is done without taking the storage elements mutex lock, so in order to
avoid memory access ordering issues, storage element status access is
changed to use READ_ONCE() and WRITE_ONCE()

Signed-off-by: Damien Le Moal <dlemoal@kernel.org>
Issuing read commands to offline zones of a ZBC disk will result in
failures. So do not issue such command and fail then early in
sd_setup_read_write_cmnd(). This change also checks write commands, which
is fine but should actually never happen as the block layer zone write
plugging code already fails write BIOs to offline zones early in the
submission path.

Signed-off-by: Damien Le Moal <dlemoal@kernel.org>
@blktests-ci-kpd

Copy link
Copy Markdown
Author

Upstream branch: 4da93ac
series: https://patchwork.kernel.org/series/1182229/
version: 4

…ulation

sd_zbc_revalidate_zones() skips revalidating the zones of a ZBC device if
the zone size and total number of zones of the disk has not changed. This
is to avoid a call to the rather slow blk_revalidate_disk_zones().

However, for ZBC devices that support data preserving head depopulation
(REMOVE ELEMENT AND MODIFY ZONES command), a disk capacity and number of
zones does not change after a head is depopulated but the condition of
zones changes as the zones served by the head that was depopulated become
either read-only or offline. In this case, not calling
blk_revalidate_disk_zones() prevents the block layer from taking
appropriate actions on the zone write plugs of the disk for the zones that
became read-only or offline.

Avoid any issue with the block layer view of the zone conditions by not
skipping the call to blk_revalidate_disk_zones() for disks that support
the REMOVE ELEMENT AND MODIFY ZONES command. This check is done from
sd_zbc_read_zones() using the helper function sd_zbc_check_modify_zones().
The new scsi disk flag modify_zones_supported is defined to remember the
result of this check.

Signed-off-by: Damien Le Moal <dlemoal@kernel.org>
Define the storage element management operations using struct
blk_storage_elements_ops. These operations are valid only on SMR disks
supporting the storage element depopulation feature.

Signed-off-by: Damien Le Moal <dlemoal@kernel.org>
@blktests-ci-kpd
blktests-ci-kpd Bot force-pushed the series/1179504=>for-next branch from 52eeaf6 to a5f66d4 Compare October 10, 2026 02:05
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant