[RFC PATCH bpf-next 09/12] namei: add an rcu-walk mode to the ancestor walk

sashiko-bot at kernel.org sashiko-bot at kernel.org
Tue Oct 6 00:33:04 UTC 2026


Thank you for your contribution! Sashiko AI review found 2 potential issue(s) to consider:
- [High] vfs: TOCTOU race in vfs_walk_handover() via inverted seqcount check
- [High] vfs: vfs_walk_ancestors() skips seqlock validation on early exit
--

commit aae071e59d5189d756d48145018ec259d19ab6d4
Author: Justin Suess <utilityemal77 at gmail.com>
Subject: namei: add an rcu-walk mode to the ancestor walk

This commit introduces an RCU-walk mode for vfs_walk_ancestors(), allowing
lockless path traversal. It adds vfs_walk_handover() to transition a lockless
walk into a referenced walk. The changes provide a lockless stepwise path
iteration mechanism intended to be driven by consumers like BPF.

> diff --git a/fs/namei.c b/fs/namei.c
> index 73f25152d9173..31f96602d4e94 100644
> --- a/fs/namei.c
> +++ b/fs/namei.c
[ ... ]
> @@ -2317,7 +2396,59 @@ int vfs_walk_next(struct vfs_ancestor_walk *aw)
[ ... ]
> +bool vfs_walk_handover(struct vfs_ancestor_walk *to,
> +		       struct vfs_ancestor_walk *from)
> +{
> +	struct path pos = from->pos;
> +	int err;
> +
> +	err = __legitimize_mnt(pos.mnt, from->m_seq);
> +	if (unlikely(err)) {
> +		if (err < 0)
> +			mnt_undo_legitimize(real_mount(pos.mnt));
> +		goto dead;
> +	}
> +	if (unlikely(read_seqcount_retry(&pos.dentry->d_seq, from->seq) ||
> +		     !lockref_get_not_dead(&pos.dentry->d_lockref))) {

[Severity: High]
Is there a race condition here due to the order of operations?

Because of the short-circuiting logical OR, read_seqcount_retry() executes
before lockref_get_not_dead(). This creates a window where a concurrent
rename or unlink could modify the dentry and increment its sequence counter
immediately after the sequence check passes but before the reference is
acquired.

Could this allow the handover to return a stale or concurrently modified
dentry to the caller? Acquiring the reference first and then validating the
sequence counter would prevent this TOCTOU race.

> +		mnt_undo_legitimize(real_mount(pos.mnt));
> +		goto dead;
> +	}
> +	to->pos = pos;
> +	to->flags = 0;
> +	to->pos_flags = 0;
> +	return true;
> +
> +dead:
> +	memset(to, 0, sizeof(*to));
> +	return false;
>  }
[ ... ]
> @@ -2347,6 +2485,8 @@ int vfs_walk_ancestors(const struct path *path,
>  	struct vfs_ancestor_walk aw;
>  	int ret;
>  
> +	if (flags & VFS_WALK_RCU)
> +		rcu_read_lock();
>  	vfs_walk_start(&aw, path, flags);
>  	for (;;) {
>  		ret = vfs_walk_next(&aw);
[ ... ]
> @@ -2365,6 +2505,8 @@ int vfs_walk_ancestors(const struct path *path,
>  		}
>  	}
>  	vfs_walk_end(&aw);
> +	if (flags & VFS_WALK_RCU)
> +		rcu_read_unlock();

[Severity: High]
Does this early exit path need sequence validation when running in RCU mode?

When the callback cb() inside this loop returns VFS_WALK_STOP or a negative
error, the loop breaks early and returns the result without validating the
seqlock.

Since VFS_WALK_RCU makes the walk lockless, the callback might operate on
torn or corrupted data during a concurrent modification. If the callback
incorrectly aborts the walk based on that corrupted state, the function
will return success or an error rather than returning -ECHILD to trigger
a retry.

Should there be a final sequence check on the current dentry before
returning the callback's result?

>  	return ret;
>  }

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20261006002020.2890858-1-utilityemal77@gmail.com?part=9



More information about the Linux-security-module-archive mailing list