[RFC PATCH bpf-next 00/12] fs: unified VFS ancestor walk for Landlock and BPF

Justin Suess utilityemal77 at gmail.com
Tue Oct 6 00:20:07 UTC 2026


Howdy folks,

This series adds a VFS API for walking a path's ancestors that covers
both path-walk modes (refwalk and rcu-walk, plus the hybrid escalation
between them) through a unified interface, converts Landlock to it, and
exposes it to BPF LSM programs.  It picks up Song Liu's bpf path
iterator series [1] where the discussion left it, and tries to be the
synthesis of what came out of that thread: a combination of the
callback and open-coded approaches proposed there.

It's a bigger diff, because it tries to cover the whole three-cheese
enchilada of VFS path walk: refwalk, rcu and hybrid. I figure doing all
modes at once is best, even if it hurts now, as bolting on an RCU api as
an afterthought would make for a miserable experience for in-kernel and
BPF users alike dealing with breakage / migration.

I'm RFC'ing this for now, as I'm still not 100% confident on the design
choices, especially on some of the weirder corners of VFS (disconnected
directories, MNT_DOOMED).

Much credit to Song Liu for the inspiration for the initial design.

Background
==========

Landlock evaluates filesystem access by walking from a path to the real
root with an open-coded dget_parent()/follow_up() loop.  This works,
but the semantics are difficult to get right. Al noted that follow_up()
ignores the mount's disappearance, so a concurrent umount can leave
that walk operating on an unaccounted mount [2].  BPF LSMs (Tetragon et
al.) want the same walk and today approximate it with racy probe reads.
Song's series added a step-up helper and a referenced-only BPF
iterator.

This attempts to bring some of the lessons from Landlock's VFS walk to
VFS core code, and exposes the same API to BPF. Other callers can be
converted later.

>From what I understand from the feedback, Christian rejected merging a
referenced-mode API now and bolting rcu-walk on later, and asked for
one unified API serving both use-cases, with the callback design Neil
sketched as the accepted shape for in-kernel callers [3].  Mickaël
additionally required that disconnected directories be well-defined
before anything lands, and suggested the layering used here:
"the best approach would be to have a VFS API with a callback,
and a BPF helper (leveraging this VFS API) with an iterator state" [4].

What this series does
=====================

It's a path walking API for VFS, with BPF and Landlock consumers.

The core is a struct vfs_ancestor_walk driven by vfs_walk_start() /
vfs_walk_next() / vfs_walk_end(), sharing its step-to-parent
cores with follow_dotdot() and follow_dotdot_rcu().  The engine
owns every invariant: references in refwalk mode, d_seq/mount_lock
validation in rcu mode, mount crossings via choose_mountpoint(),
which is what fixes the Landlock race.  Both modes live behind the
same calls; a VFS_WALK_RCU flag picks the mode and nothing outside
namei.c sees seqcounts or nameidata.

On top of the engine sit two front-ends matched to their consumers.
In-kernel callers get the callback design, which must not sleep.

  int vfs_walk_ancestors(const struct path *path,
			 int (*cb)(const struct path *ancestor,
				   unsigned int pos_flags, void *data),
			 void *data, unsigned int flags);

@cb is invoked on @path, then on each ancestor up to the real root,
and returns VFS_WALK_CONTINUE, VFS_WALK_STOP or a negative errno;
@flags takes VFS_WALK_RCU when the callback copes with unreferenced
positions.  Landlock's walk converts to this, preserving its evaluation
order exactly (including the disconnected-directory semantics of its
recent fixes) plus the mount_lock-validated crossings.

The engine itself is declared in fs/internal.h:

  void vfs_walk_start(struct vfs_ancestor_walk *aw,
		      const struct path *path, unsigned int flags);
  int  vfs_walk_next(struct vfs_ancestor_walk *aw);
  void vfs_walk_end(struct vfs_ancestor_walk *aw);
  bool vfs_walk_handover(struct vfs_ancestor_walk *to,
			 struct vfs_ancestor_walk *from);

So the one consumer that cannot be a callback (BPF) can drive it one
step per call, from fs/bpf_fs_kfuncs.c. Here are the (numerous) kfuncs,
many of which are thin wrappers around the engine:

  /* referenced, sleepable: each position comes acquired */
  int  bpf_iter_path_ancestors_new(struct bpf_iter_path_ancestors *it,
				   struct path *path, u64 flags);
  struct path *
       bpf_iter_path_ancestors_next(struct bpf_iter_path_ancestors *it);
  void bpf_iter_path_ancestors_destroy(struct bpf_iter_path_ancestors *it);
  u32  bpf_path_ancestors_pos_flags(struct bpf_iter_path_ancestors *it);
  void bpf_path_put(struct path *path);

  /* lockless: positions are borrowed, valid until the next step */
  int  bpf_iter_path_ancestors_rcu_new(struct bpf_iter_path_ancestors_rcu *it,
				       struct path *path, u64 flags);
  struct path *
       bpf_iter_path_ancestors_rcu_next(struct bpf_iter_path_ancestors_rcu *it);
  void bpf_iter_path_ancestors_rcu_destroy(struct bpf_iter_path_ancestors_rcu *it);
  u32  bpf_path_ancestors_rcu_pos_flags(struct bpf_iter_path_ancestors_rcu *it);

  /* escalation for hybrid walk. Mirroring unlazy_walk(): acquire the
   * lockless walk's current position straight into a referenced iterator
   */
  int  bpf_path_ancestors_legitimize(struct bpf_iter_path_ancestors *it,
				     struct bpf_iter_path_ancestors_rcu *rcu_it);

The lockless constructor is KF_RCU_PROTECTED, so the verifier forces
the whole iteration into one RCU read-side critical section, within
which everything sleepable is already rejected. Each next() is still
namei.c code, so the walk invariants stay in the VFS even though the
loop is in the program.  A hybrid walk runs lockless until it needs
sleepable work, then it escalates. Here's what hybrid walk looks like
concretely:

	bpf_rcu_read_lock();
	bpf_iter_path_ancestors_rcu_new(&rit, dir, 0);
	while ((pos = bpf_iter_path_ancestors_rcu_next(&rit))) {
		/* borrowed position: evaluate, but no sleeping allowed */
		if (foobar(pos))
			break;
	}
  /* converts the iterator from rcu -> refwalk */
	bpf_path_ancestors_legitimize(&it, &rit);
	bpf_iter_path_ancestors_rcu_destroy(&rit);
	bpf_rcu_read_unlock();

	/* sleepable kfuncs are legal again. */
	while ((pos = bpf_iter_path_ancestors_next(&it))) {
		bpf_path_d_path(pos, buf, sizeof(buf));
    /* refwalk gives you path references that you gotta free to appease
     * the verifier
     */
		bpf_path_put(pos);
	}
	bpf_iter_path_ancestors_destroy(&it);

Nothing is allocated during the escalation, so it cannot fail for want
of memory inside the critical section, and the sleepable work lands
after bpf_rcu_read_unlock() by construction: the resuming next() is a
sleepable kfunc, an ordering the verifier enforces.

A lockless walk that loses a race does not restart transparently: it
dies with -ECHILD (BPF: BPF_PATH_ANCESTORS_RETRY), the caller discards
what it accumulated and retries in referenced mode, so one lockless
attempt bounds the retries.  This is simpler to reason about than the
restart-signal contract discussed in the thread.

Disconnected directories are exposed first-class rather than papered over:
positions whose dentry is a disconnected root are flagged
VFS_WALK_POS_DISCONNECTED (plus VFS_WALK_POS_MOUNTPOINT for the
mountpoint a crossing landed on, which the old Landlock loop never
visited), and continuing over one resumes at the root of its mount.
The mechanism lives in the walker; the MNT_INTERNAL allow-and-stop
policy stays in Landlock.

Supporting pieces: mnt_undo_legitimize() gives a failed
__legitimize_mnt() an undo callable inside the RCU read-side critical
section, deferring a final mntput to delayed_mntput().

Secondly, for the bpf_path_ancestors_legitimize, a patch is added to allow
kfuncs to accept an "__uninit"-suffixed iterator argument so a generic
kfunc (the handover) can initialize an iterator from another one, since
KF_ITER_NEW allows only one constructor per type; and an acquiring
KF_ITER_NEXT's drained branch now releases the reference its acquire
bookkeeping created, which no in-tree iterator needed before.

Patches 1 and 4 are carried from Song's series; patch 3 is derived from
his Landlock conversion and carries the Fixes: tag for the follow_up()
race, per Mickaël's request on v5.

Open questions
==============

- Whether an open-coded iterator is acceptable to the VFS as the BPF
  front-end, given the engine and invariants stay in namei.c and the
  verifier provides the discipline a kernel-owned loop would.  The
  callback shape was the thread's endorsed design for in-kernel
  callers; I believe this split is what Mickaël proposed in [4], and
  the alternative (BPF programs as the callback) costs considerably
  more verifier machinery for a worse programming model.

- Whether VFS_WALK_POS_MOUNTPOINT should exist at all.  It is there so
  Landlock's rule evaluation stays bit-for-bit what it was before the
  conversion; if matching rules on crossed-onto disconnected
  mountpoints is acceptable as a behavior change, the flag disappears.

- The retry-on-ECHILD contract versus a transparent restart signal.

[1] https://lore.kernel.org/bpf/20250617061116.3681325-1-song@kernel.org/
[2] https://lore.kernel.org/r/20250529231018.GP2023217@ZenIV
[3] https://lore.kernel.org/all/20250707-netto-campieren-501525a7d10a@brauner/
[4] https://lore.kernel.org/all/20250704.quio1ceil4Xi@digikod.net/

Justin Suess (11):
  namei: add vfs_walk_ancestors()
  landlock: convert ancestor walk to vfs_walk_ancestors()
  bpf: mark struct path trusted
  namei: make vfs_walk_ancestors() stepwise
  bpf: add a path ancestor iterator
  selftests/bpf: exercise the path ancestor iterator
  fs: add mnt_undo_legitimize()
  namei: add an rcu-walk mode to the ancestor walk
  bpf: support "__uninit" iterator arguments in generic kfuncs
  bpf: add a lockless path ancestor iterator
  selftests/bpf: exercise the lockless path ancestor iterator

Song Liu (1):
  namei: introduce __path_walk_parent()

 fs/bpf_fs_kfuncs.c                            | 248 ++++++++++++
 fs/internal.h                                 |  22 +
 fs/mount.h                                    |   1 +
 fs/namei.c                                    | 377 ++++++++++++++++--
 fs/namespace.c                                |  63 ++-
 include/linux/namei.h                         |  17 +
 kernel/bpf/verifier.c                         |  49 ++-
 security/landlock/fs.c                        | 265 ++++++------
 .../selftests/bpf/prog_tests/path_ancestors.c |  77 ++++
 .../selftests/bpf/progs/path_ancestors.c      | 134 +++++++
 10 files changed, 1076 insertions(+), 177 deletions(-)
 create mode 100644 tools/testing/selftests/bpf/prog_tests/path_ancestors.c
 create mode 100644 tools/testing/selftests/bpf/progs/path_ancestors.c


base-commit: 99dc1ba542420db6b8df209744f55cc52466ad91
-- 
2.55.0




More information about the Linux-security-module-archive mailing list