[PATCH 0/7] fs: preserve superblock inode walk positions across lock drops

Jan Kara jack at suse.cz
Wed Sep 9 12:49:33 UTC 2026


On Wed 09-09-26 17:01:05, Julian Sun wrote:
> [Motivation]
> 
> We observed hung tasks in production during disk hotplug operations. A
> kernel thread spends a long time in evict_inodes() while holding s_umount,
> blocking other users of that lock and causing further stalls.
> 
> The problem is that evict_inodes() restarts its walk from the head of
> sb->s_inodes every time it reschedules. When a large number of referenced
> inodes remain near the head, each restart scans those inodes again without
> making progress through that part of the list. The repeated scans can
> delay eviction long enough to trigger hung-task reports. This is the same
> problem that [1] attempted to address.

Thanks fro the patch! What's really unexpected is that there are many
inodes with incremented refcount at the time when evict_inodes() is called.
Did you have a look who's holding those references? Or is evict_inodes()
called through the fs_bdev_mark_dead() call (as you mention "disk hotplug
operations")? There the filesystem is in fact in use so what you describe
makes some sense.

								Honza

> [Approach]
> 
> This series introduces sb_for_each_inodes() for two purposes:
> 
>   1. Consolidate open-coded s_inodes walks behind a common entry point.
>   2. Retain each walk's position across drops of s_inode_list_lock.
> 
> The iterator mechanism follows the approach used by cgroup task iteration,
> such as css_task_iter_next(). Active iterators are registered on a separate
> list, sb->s_inodes_iters. Before removing an inode from s_inodes, the
> removal path advances any iterator whose next position points to that
> inode. These updates are protected by s_inode_list_lock, so a walker can
> drop the lock and later resume from its saved position.
> 
> Existing walkers, such as drop_pagecache_sb() and add_dquot_ref(), already
> contain their own position-preserving logic: they carry an inode reference
> across iterations so that they can resume after dropping the list lock.
> Moving that responsibility into sb_for_each_inodes() simplifies these
> callers and lets their callbacks focus on the per-inode work.
> 
> Patch 1 removes trailing whitespace from include/linux/fs.h.
> Patch 2 introduces sb_for_each_inodes().
> The remaining patches convert existing walks to the new interface.
> remove_dquot_ref() and nr_blockdev_pages() are left unchanged: their
> walks are simple and do not require the inode->i_lock locking imposed
> by the callback interface. Converting them would add an unnecessary
> lock/unlock overhead for every inode.
> 
> [Testing]
> 
> I tested this series with approximately 20 hours of xfstests case
> execution, repeatedly running the auto group on ext4 and XFS, no new
> issues were observed. And with this patch applied, the hung task that
> previously occurred on every run no longer occurs.
> 
> [1] https://lore.kernel.org/all/20241118114508.1405494-1-yebin@huaweicloud.com/
> 
> Julian Sun (7):
>   fs: remove trailing whitespace from include/linux/fs.h
>   fs: introduce sb_for_each_inodes().
>   block: use sb_for_each_inodes() in sync_bdevs()
>   fs: use sb_for_each_inodes() API.
>   gfs2: use sb_for_each_inodes() for cooperative eviction
>   quota: use sb_for_each_inodes() in add_dquot_ref()
>   landlock: use sb_for_each_inodes() when detaching a superblock
> 
>  block/bdev.c                   |  85 +++++++++----------
>  fs/drop_caches.c               |  44 +++++-----
>  fs/gfs2/ops_fstype.c           |  38 ++++-----
>  fs/inode.c                     | 142 +++++++++++++++++++++++--------
>  fs/quota/dquot.c               |  72 ++++++----------
>  fs/super.c                     |   1 +
>  include/linux/fs.h             |  29 +++++--
>  include/linux/fs/super_types.h |   3 +-
>  security/landlock/fs.c         | 150 +++++++++++++--------------------
>  9 files changed, 295 insertions(+), 269 deletions(-)
> 
> -- 
> 2.39.5
> 
-- 
Jan Kara <jack at suse.com>
SUSE Labs, CR



More information about the Linux-security-module-archive mailing list