[RFC PATCH bpf-next 09/12] namei: add an rcu-walk mode to the ancestor walk
bot+bpf-ci at kernel.org
bot+bpf-ci at kernel.org
Tue Oct 6 01:10:57 UTC 2026
> namei: add an rcu-walk mode to the ancestor walk
TL;DR: vfs_walk_handover() may drop the MOUNTPOINT flag on a disconnected
root, so an escalated walk could report different position flags and give
Landlock a different verdict than the lockless walk.
> diff --git a/fs/namei.c b/fs/namei.c
> index 73f25152d917..31f96602d4e9 100644
> --- a/fs/namei.c
> +++ b/fs/namei.c
[ ... ]
> @@ -2286,6 +2330,35 @@ static int vfs_walk_step_ref(struct vfs_ancestor_walk *aw)
[ ... ]
> + /* A crossing onto a disconnected root, as in vfs_walk_step_ref(). */
> + aw->pos_flags = parent == aw->pos.dentry ?
> + VFS_WALK_POS_DISCONNECTED | VFS_WALK_POS_MOUNTPOINT :
> + vfs_walk_pos_flags(aw->pos.mnt, parent);
> + aw->pos.dentry = parent;
> + aw->seq = next_seq;
> + return 0;
> +}
> +
> /**
[ ... ]
> @@ -2308,6 +2383,10 @@ int vfs_walk_next(struct vfs_ancestor_walk *aw)
> aw->flags |= VFS_WALK_STARTED;
> aw->pos_flags = vfs_walk_pos_flags(aw->pos.mnt, aw->pos.dentry);
> }
[ ... ]
> @@ -2317,7 +2396,59 @@ int vfs_walk_next(struct vfs_ancestor_walk *aw)
[ ... ]
> +bool vfs_walk_handover(struct vfs_ancestor_walk *to,
> + struct vfs_ancestor_walk *from)
> +{
> + struct path pos = from->pos;
> + int err;
> +
> + err = __legitimize_mnt(pos.mnt, from->m_seq);
> + if (unlikely(err)) {
> + if (err < 0)
> + mnt_undo_legitimize(real_mount(pos.mnt));
> + goto dead;
> + }
> + if (unlikely(read_seqcount_retry(&pos.dentry->d_seq, from->seq) ||
> + !lockref_get_not_dead(&pos.dentry->d_lockref))) {
> + mnt_undo_legitimize(real_mount(pos.mnt));
> + goto dead;
> + }
> + to->pos = pos;
> + to->flags = 0;
> + to->pos_flags = 0;
> + return true;
Can the first vfs_walk_next() on @to yield the same VFS_WALK_POS_* flags
that @from reported for this position?
vfs_walk_step_rcu() can leave @from on a disconnected root with
two flags set:
aw->pos_flags = parent == aw->pos.dentry ?
VFS_WALK_POS_DISCONNECTED | VFS_WALK_POS_MOUNTPOINT :
vfs_walk_pos_flags(aw->pos.mnt, parent);
vfs_walk_handover() then sets @to up as a walk that has not started
(to->flags = 0, to->pos_flags = 0). The first vfs_walk_next() on @to takes
the not-started branch and recomputes the flags from the position alone:
aw->flags |= VFS_WALK_STARTED;
aw->pos_flags = vfs_walk_pos_flags(aw->pos.mnt, aw->pos.dentry);
vfs_walk_pos_flags() can only return VFS_WALK_POS_DISCONNECTED or 0. So if
the handover happens while @from sits on such a position, @from reports
DISCONNECTED | MOUNTPOINT, but @to yields the same dentry as DISCONNECTED
only. The kernel-doc of vfs_walk_handover() says this position is what "the
first vfs_walk_next() on it yields".
Whether a disconnected root is a mountpoint a crossing landed on is walk
state, not something that can be recomputed from the position. The
bpf_path_ancestors_pos_flags() changelog (8dd2a68fe0e7, "bpf: add a path
ancestor iterator") says a program needs it "to reproduce Landlock's
evaluation of disconnected positions", and check_access_path_walk() in
security/landlock/fs.c skips rule matching on MOUNTPOINT positions but
matches rules on positions flagged only DISCONNECTED.
A program that escalates with bpf_path_ancestors_legitimize() (47d25368b3cc,
later in this series) at such a position would see a plain disconnected
root through bpf_path_ancestors_pos_flags(). It could then match rules
against the mountpoint dentry, so the same path can get a different verdict
depending on whether the walk escalated.
The position is reachable, since d_set_mounted() accepts an IS_ROOT
disconnected dentry as a mountpoint, for example one obtained through
open_by_handle_at().
Should the flags be carried over (to->pos_flags = from->pos_flags), with the
not-started branch of vfs_walk_next() ORing in vfs_walk_pos_flags() instead
of overwriting the flags? vfs_walk_start() already zeroes pos_flags, so
ordinary walks would be unaffected.
No later commit in 3cc40137a219..70352ce0817c touches fs/namei.c.
---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
CI run summary: https://github.com/kernel-patches/bpf/actions/runs/37395354107
More information about the Linux-security-module-archive
mailing list