[RFC PATCH v2 2/3] security: Add PR_CAPBSET_DROP_MASK for process-wide bounding-set drops
sashiko-bot at kernel.org
sashiko-bot at kernel.org
Tue Sep 29 13:19:07 UTC 2026
Thank you for your contribution! Sashiko AI review found 2 potential issue(s) to consider:
- [High] userns: PR_CAPBSET_DROP_MASK leaks into new user namespaces
- [High] cred: PR_CAPBSET_DROP_MASK breaks self-containment of struct cred
--
commit fcb774e0491f03726a243dd79d9b587551550d6e
Author: Jinjie Ruan <ruanjinjie at huawei.com>
security: Add PR_CAPBSET_DROP_MASK for process-wide bounding-set drops
This commit introduces a new prctl, PR_CAPBSET_DROP_MASK, to drop bounding set
capabilities across an entire thread group simultaneously. It achieves this by
storing a pending drop mask in the signal_struct and dynamically subtracting
it during capability checks.
> diff --git a/security/commoncap.c b/security/commoncap.c
> index 3399535808fec..3bec15d437312 100644
> --- a/security/commoncap.c
> +++ b/security/commoncap.c
[ ... ]
> @@ -30,6 +32,25 @@
> #define CREATE_TRACE_POINTS
> #include <trace/events/capability.h>
>
> +/**
> + * Effective bounding set of @cred in @task's thread group
> + * @task: task whose thread group's pending drop applies
> + * @cred: credentials to read the bounding set from
> + *
> + * A drop recorded by PR_CAPBSET_DROP_MASK is authoritative on the thread group
> + * and may not have been materialized into every thread's cred yet, so the
> + * effective bounding set is the cred's own set minus the group's pending drop.
> + */
> +kernel_cap_t cap_bset_effective(const struct task_struct *task,
> + const struct cred *cred)
> +{
> + kernel_cap_t pending = {
> + .val = atomic64_read(&task->signal->cap_bset_pending),
> + };
> +
> + return cap_drop(cred->cap_bset, pending);
> +}
[Severity: High]
Does this dynamic subtraction inadvertently leak parent bounding set drops into
new user namespaces?
When a new user namespace is created via clone(CLONE_NEWUSER) or unshare(),
set_cred_user_ns() resets cred->cap_bset to CAP_FULL_SET to ensure the new
namespace starts with full capabilities relative to itself.
However, because cap_bset_effective() dynamically subtracts the pending mask
from task->signal, any pending drops from the parent namespace will still be
applied to the new namespace's credentials. This could permanently restrict
them.
[ ... ]
> @@ -1283,6 +1306,63 @@ static int cap_prctl_drop(unsigned long cap)
> return commit_creds(new);
> }
>
> +static int cap_bset_drop_process(kernel_cap_t mask)
> +{
> + kernel_cap_t pending;
> + struct cred *new;
> +
> + new = prepare_creds();
> + if (!new)
> + return -ENOMEM;
> +
> + /*
> + * Record the drop before committing the caller's cred, so that a
> + * thread created from now on is guaranteed to observe it. Apply the
> + * current union, not just this call's mask, to the caller's cred.
> + */
> + spin_lock_irq(¤t->sighand->siglock);
> + atomic64_or(mask.val, ¤t->signal->cap_bset_pending);
[Severity: High]
Does storing the capability drop outside of struct cred break subsystems that
rely on the immutability and self-containment of saved credentials?
If a thread calls PR_CAPBSET_DROP_MASK, while another thread in the same
group opens an AppArmor-mediated file, that file saves the thread's stale
credential (e.g., in file->f_cred).
AppArmor might then evaluate the bounding set directly from this standalone
credential snapshot (like in is_subset_of_obj_privilege()). By placing the
drop state in signal_struct rather than struct cred, these asynchronous
evaluations will miss the pending drops.
Could this lead to false denials or security bypasses because the capability
evaluations are using stale capabilities?
> + pending.val = atomic64_read(¤t->signal->cap_bset_pending);
> + spin_unlock_irq(¤t->sighand->siglock);
> +
> + new->cap_bset = cap_drop(new->cap_bset, pending);
> + commit_creds(new);
> +
> + return 0;
> +}
> +
> +/*
> + * Propagate a pending process-wide bounding-set drop to @p, a task being
> + * created by the current thread. Threads sharing the group read
> + * current->signal->cap_bset_pending directly; a forked child gets its own
> + * signal_struct and must carry the mask itself. Called under
> + * current->sighand->siglock, which serializes it with cap_bset_drop_process().
> + */
> +void cap_bset_drop_fork(struct task_struct *p)
> +{
> + kernel_cap_t mask = {
> + .val = atomic64_read(¤t->signal->cap_bset_pending),
> + };
> +
> + if (cap_isclear(mask) || p->signal == current->signal)
> + return;
> +
> + atomic64_or(mask.val, &p->signal->cap_bset_pending);
> +}
[Severity: High]
Does actively copying the parent's dropped capabilities mask into a newly
created process erase namespace isolation boundaries?
When CLONE_NEWUSER is requested, the child is meant to receive a full
bounding set within its new namespace.
Because cap_bset_drop_fork() copies the parent's pending mask into the
child's signal_struct without clearing it for new user namespaces, those
capability drops will bleed into the child's effective bounding set. This
breaks the isolation guarantees of user namespaces.
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260929130200.1638343-1-ruanjinjie@huawei.com?part=2
More information about the Linux-security-module-archive
mailing list