[RFC PATCH bpf-next 11/12] bpf: add a lockless path ancestor iterator

bot+bpf-ci at kernel.org bot+bpf-ci at kernel.org
Tue Oct 6 01:10:47 UTC 2026


> bpf: add a lockless path ancestor iterator

TL;DR:
- Verifier may give a plain trusted dentry from the lockless path iterator,
  risking use-after-free in sleepable kfuncs after RCU exit
- A failed legitimize may look like a finished walk, hiding a lost race

> diff --git a/fs/bpf_fs_kfuncs.c b/fs/bpf_fs_kfuncs.c
> index 265cb414a08a..9dad65df4267 100644
> --- a/fs/bpf_fs_kfuncs.c
> +++ b/fs/bpf_fs_kfuncs.c

[ ... ]

> @@ -633,6 +656,79 @@ bpf_path_ancestors_pos_flags(struct bpf_iter_path_ancestors *it__iter)

[ ... ]

> +/*
> + * Unlike the referenced variant, this hands out the walk's own position:
> + * a lockless iteration holds no references to pass on, and the verifier
> + * keeps the whole of it inside one RCU read-side critical section.
> + */
> +__bpf_kfunc struct path *
> +bpf_iter_path_ancestors_rcu_next(struct bpf_iter_path_ancestors_rcu *it)
> +{
> +	return bpf_path_ancestors_step((void *)it);
> +}

The commit message says that "nothing read out of the position may be passed
to a kfunc demanding a trusted argument", and the comment above struct
bpf_iter_path_ancestors says the same.  Does the verifier actually enforce
this?

The position returned by bpf_iter_path_ancestors_rcu_next() is typed
PTR_TO_BTF_ID | MEM_RCU, because check_kfunc_call() adds MEM_RCU to the
iter_next result of a KF_RCU_PROTECTED iterator.  Commit 7ed5aa71ad77 ("bpf:
mark struct path trusted") added:

    BTF_TYPE_SAFE_TRUSTED(struct path) {
            struct dentry *dentry;
    };

In check_ptr_to_btf_access(), the branch taken for
is_trusted_reg(env, reg) || is_rcu_reg(reg) applies type_is_trusted() and
sets PTR_TRUSTED on a safe-trusted field even when the base pointer is only
MEM_RCU.

So loading the dentry field of a lockless position yields a plain
PTR_TRUSTED dentry.  It does not keep MEM_RCU, so
invalidate_rcu_protected_refs() leaves it alone at bpf_rcu_read_unlock().

The safe-trusted rule for struct path assumes the path pins its dentry.  That
holds for the referenced variant, but the lockless walk takes no reference on
the dentry, and an ancestor's lifetime is not tied to the original path once
a concurrent rename moves the descendant.

After the program leaves the RCU read-side section, could a sleepable kfunc
that takes a trusted dentry (the xattr kfuncs in this set, for example) be
handed a dentry that has neither a reference nor RCU protection?

Would it work to keep MEM_RCU on fields loaded from a MEM_RCU struct path
(for example by applying the struct path safe-trusted rule only to trusted
bases), or to avoid exposing struct path as the return type of the lockless
iterator?  A negative selftest showing that such a load is rejected would
also help.

[ ... ]

> + * Return: 0, -%ENOENT if the lockless iteration was not on a position, or
> + * -%ECHILD if it lost the race to acquire one; %BPF_PATH_ANCESTORS_RETRY is
> + * then also flagged, and the program has reached no conclusion about the
> + * ancestry.  @it__uninit is initialized whatever this returns, so a program
> + * need not branch on the result: a walk that could not be escalated simply
> + * yields no position.
> + */
> +__bpf_kfunc int
> +bpf_path_ancestors_legitimize(struct bpf_iter_path_ancestors *it__uninit,
> +			      struct bpf_iter_path_ancestors_rcu *rcu_it__iter)
> +{
> +	struct bpf_path_ancestors_kern *rcu_kit = (void *)rcu_it__iter;
> +	struct bpf_path_ancestors_kern *kit = (void *)it__uninit;
> +
> +	/* A zeroed walk makes destroying the iterator a no-op. */
> +	memset(kit, 0, sizeof(*kit));
> +	kit->step = 1;
> +
> +	/* Drained, or already failed: nothing to hand over. */
> +	if (rcu_kit->step)
> +		return -ENOENT;
> +	if (!vfs_walk_handover(&kit->aw, &rcu_kit->aw)) {
> +		rcu_kit->step = -ECHILD;
> +		return -ECHILD;
> +	}
> +	kit->step = 0;
> +	return 0;
> +}

When the lockless iteration had already failed with -ECHILD, or when
vfs_walk_handover() fails, @it__uninit is left with step == 1.  That is the
same state as a referenced walk that has passed the real root, so
bpf_iter_path_ancestors_next() returns NULL and bpf_path_ancestors_pos_flags()
on that iterator returns 0.

The documented contract of bpf_iter_path_ancestors_next() says NULL comes
"once the walk has passed the real root - or on an allocation failure ...
reported as NOMEM", so a lost race reads as a completed walk.  The kernel-doc
above also tells programs they "need not branch on the result".

Can a policy program that follows that advice and only inspects the resumed
iterator treat a lost race as a finished ancestry walk with no match?
BPF_PATH_ANCESTORS_RETRY is flagged only on the lockless iterator, which has
to be destroyed before the resumed iteration can run.

Would setting kit->step = -ECHILD on the destination in both failure paths
work?  Stepping would still stop (step != 0), the zeroed walk would keep
destroy a no-op, and bpf_path_ancestors_pos_flags() on the resumed iterator
would report BPF_PATH_ANCESTORS_RETRY.

Separately, when the source had already lost a race, the first check returns
-ENOENT rather than -ECHILD, which conflicts with the kernel-doc describing
-ENOENT as "not on a position".


---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md

CI run summary: https://github.com/kernel-patches/bpf/actions/runs/37395354107


More information about the Linux-security-module-archive mailing list