[RFC PATCH bpf-next 09/12] namei: add an rcu-walk mode to the ancestor walk

bot+bpf-ci at kernel.org bot+bpf-ci at kernel.org
Tue Oct 6 01:10:57 UTC 2026


> namei: add an rcu-walk mode to the ancestor walk

TL;DR: vfs_walk_handover() may drop the MOUNTPOINT flag on a disconnected
root, so an escalated walk could report different position flags and give
Landlock a different verdict than the lockless walk.

> diff --git a/fs/namei.c b/fs/namei.c
> index 73f25152d917..31f96602d4e9 100644
> --- a/fs/namei.c
> +++ b/fs/namei.c

[ ... ]

> @@ -2286,6 +2330,35 @@ static int vfs_walk_step_ref(struct vfs_ancestor_walk *aw)

[ ... ]

> +	/* A crossing onto a disconnected root, as in vfs_walk_step_ref(). */
> +	aw->pos_flags = parent == aw->pos.dentry ?
> +		VFS_WALK_POS_DISCONNECTED | VFS_WALK_POS_MOUNTPOINT :
> +		vfs_walk_pos_flags(aw->pos.mnt, parent);
> +	aw->pos.dentry = parent;
> +	aw->seq = next_seq;
> +	return 0;
> +}
> +
>  /**

[ ... ]

> @@ -2308,6 +2383,10 @@ int vfs_walk_next(struct vfs_ancestor_walk *aw)
>  		aw->flags |= VFS_WALK_STARTED;
>  		aw->pos_flags = vfs_walk_pos_flags(aw->pos.mnt, aw->pos.dentry);
>  	}

[ ... ]

> @@ -2317,7 +2396,59 @@ int vfs_walk_next(struct vfs_ancestor_walk *aw)

[ ... ]

> +bool vfs_walk_handover(struct vfs_ancestor_walk *to,
> +		       struct vfs_ancestor_walk *from)
> +{
> +	struct path pos = from->pos;
> +	int err;
> +
> +	err = __legitimize_mnt(pos.mnt, from->m_seq);
> +	if (unlikely(err)) {
> +		if (err < 0)
> +			mnt_undo_legitimize(real_mount(pos.mnt));
> +		goto dead;
> +	}
> +	if (unlikely(read_seqcount_retry(&pos.dentry->d_seq, from->seq) ||
> +		     !lockref_get_not_dead(&pos.dentry->d_lockref))) {
> +		mnt_undo_legitimize(real_mount(pos.mnt));
> +		goto dead;
> +	}
> +	to->pos = pos;
> +	to->flags = 0;
> +	to->pos_flags = 0;
> +	return true;

Can the first vfs_walk_next() on @to yield the same VFS_WALK_POS_* flags
that @from reported for this position?

vfs_walk_step_rcu() can leave @from on a disconnected root with
two flags set:

    aw->pos_flags = parent == aw->pos.dentry ?
            VFS_WALK_POS_DISCONNECTED | VFS_WALK_POS_MOUNTPOINT :
            vfs_walk_pos_flags(aw->pos.mnt, parent);

vfs_walk_handover() then sets @to up as a walk that has not started
(to->flags = 0, to->pos_flags = 0).  The first vfs_walk_next() on @to takes
the not-started branch and recomputes the flags from the position alone:

    aw->flags |= VFS_WALK_STARTED;
    aw->pos_flags = vfs_walk_pos_flags(aw->pos.mnt, aw->pos.dentry);

vfs_walk_pos_flags() can only return VFS_WALK_POS_DISCONNECTED or 0.  So if
the handover happens while @from sits on such a position, @from reports
DISCONNECTED | MOUNTPOINT, but @to yields the same dentry as DISCONNECTED
only.  The kernel-doc of vfs_walk_handover() says this position is what "the
first vfs_walk_next() on it yields".

Whether a disconnected root is a mountpoint a crossing landed on is walk
state, not something that can be recomputed from the position.  The
bpf_path_ancestors_pos_flags() changelog (8dd2a68fe0e7, "bpf: add a path
ancestor iterator") says a program needs it "to reproduce Landlock's
evaluation of disconnected positions", and check_access_path_walk() in
security/landlock/fs.c skips rule matching on MOUNTPOINT positions but
matches rules on positions flagged only DISCONNECTED.

A program that escalates with bpf_path_ancestors_legitimize() (47d25368b3cc,
later in this series) at such a position would see a plain disconnected
root through bpf_path_ancestors_pos_flags().  It could then match rules
against the mountpoint dentry, so the same path can get a different verdict
depending on whether the walk escalated.

The position is reachable, since d_set_mounted() accepts an IS_ROOT
disconnected dentry as a mountpoint, for example one obtained through
open_by_handle_at().

Should the flags be carried over (to->pos_flags = from->pos_flags), with the
not-started branch of vfs_walk_next() ORing in vfs_walk_pos_flags() instead
of overwriting the flags?  vfs_walk_start() already zeroes pos_flags, so
ordinary walks would be unaffected.

No later commit in 3cc40137a219..70352ce0817c touches fs/namei.c.


---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md

CI run summary: https://github.com/kernel-patches/bpf/actions/runs/37395354107


More information about the Linux-security-module-archive mailing list