Operating-system confinement
Every bash command crosses a host-owned sandbox service before any command
process starts. Tool arguments describe the command, not the backend. The host
resolves one immutable filesystem, network, environment, command and resource
plan; probes the backend; refuses missing hard capabilities; materializes inert
inputs; and only then launches the process.
OS confinement is disabled by default. Enable it with
{"sandbox":{"enabled":true}}; enabled: false in your home configuration
selects unconfined execution. Approval, minimal environment, command deadlines,
output bounds and audit remain active whether confinement is enabled or disabled.
Configure filesystem grants and restrictions, domain mediation, local sockets
and command limits in the same block. Project settings can require confinement
and narrow inherited authority. See configuration.
Turning it on
- Give the platform what it needs. Confinement runs only where this
machine can enforce it; while it is on and this machine cannot, a command
is refused rather than run unconfined.
- Linux: Bubblewrap 0.11.0 or newer from the system's packages (
bwrap --versionsays which you have), andcrucible-sandbox-brokerbesidecrucible, whichinstall.shputs there. - macOS: nothing to install. The built-in
/usr/bin/sandbox-execdoes the confining, withcrucible-sandbox-brokerbesidecrucible. - Windows:
crucible-sandbox-broker.exebesidecrucible.exe, and once, from an Administrator PowerShell in that directory,.\crucible.exe sandbox setup, adding--ownerwith your account when that PowerShell runs as another administrator. It creates the account and network rules described under Windows setup maintenance.
- Linux: Bubblewrap 0.11.0 or newer from the system's packages (
- Turn it on with
/sandbox enablein a conversation, or with{"sandbox":{"enabled":true}}in~/.crucible/config.json. The command writes that same setting, but only once it has checked that this machine can enforce it; when it cannot, it says why and changes nothing. - Check it.
crucible --sandbox, run in the project, prints the backend and what a command there would run under, without running one (reading the report). In a conversation,/sandboxshows the settings, and its Dependencies tab says whether this machine's backend is available and, when it is not, why. - Turn it off with
/sandbox disable, unless the project requires it. On Windows,.\crucible.exe sandbox uninstall, from an Administrator PowerShell in the same directory and with the same--ownersetup was given, removes what setup created.
How Linux and macOS confine a command
On Linux, the production backend uses a
canonical, root-owned, non-writable system Bubblewrap executable reached only
through root-owned, non-writable parent directories. Every bounded PATH
candidate is tried until one passes its exact command-surface, numeric version,
namespace and SHA-256 identity probes. This release records that the bundled
choice is unavailable; it does not silently substitute a plain subprocess,
Landlock, a container, or a worktree. If no system candidate satisfies the
effective policy, preparation fails before materialization or spawn.
Bubblewrap 0.11.0 or newer is needed: the view is built from descriptor binds and temporary overlays, which older releases do not offer. The probe checks the options themselves rather than the version, because distributions backport some of them. Ubuntu 24.04's 0.9.0 has the descriptor binds but not the overlays, so there enabled confinement reports an unavailable backend until a newer Bubblewrap is installed. Should a launch still be refused by the system Bubblewrap, its own message is quoted in the error.
Inside the namespace, PID 1 is Crucible's own crucible-sandbox-broker
executable, shipped in every Linux release archive and installed beside the
Crucible binary, where it is found and pinned by open descriptor before the
namespace starts. It is accepted only when it and every directory
above it belong to root or to the user running Crucible and are writable by
neither group nor others; a copy under /tmp, in another user's directory or
below a group-writable directory is ignored, and chmod g-w on the offending
directory is the remedy. When no copy qualifies, the error names each place it
looked, beside the Crucible binary or in the directory above it, and what
turned it down, so a broker that was never built reads differently from one
below a directory others can rewrite.
The Linux view starts from an empty temporary root. It exposes only the minimal
read-only runtime needed to execute the selected absolute program, the exact
workspace/reached roots at their granted access, protected repository control
metadata and recognized agent configuration directories, a minimal /proc and
/dev, and a transactionally staged manifest. Bounded unreadable patterns use a
deliberately small */single-**
grammar and one deterministic, no-symlink, no-mount-crossing tree scan. It
creates isolated user, PID, IPC, UTS and network namespaces, drops capabilities,
sets no-new-privileges through Bubblewrap, disables nested user namespaces,
clears the environment, and closes every undeclared file descriptor. Closed
networking has no usable host, loopback, Unix-socket, DNS, metadata-service or
inherited-socket route. Killing the namespace owner also kills descendants that
deliberately leave the original process group or session.
On macOS, the production backend uses the built-in Seatbelt framework through
the fixed /usr/bin/sandbox-exec executable. The launcher must be root-owned
and non-writable. Crucible's separately packaged sandbox broker must be beside
the main executable on a path writable only by root or the current user; it
closes every inherited descriptor above standard input, output and error,
applies the requested open-file limit, and then replaces itself with the system
launcher. Darwin's CPU rlimit delivers a catchable signal and is therefore not
advertised as a hard CPU ceiling.
The generated Seatbelt profile is closed by default. It reads only fixed
system-runtime directories plus the declared filesystem roots,
writes only declared read-write roots, carves protected and unreadable paths
back out, and denies networking. Declared grants and exact carve-outs are
supplied as Seatbelt parameters. Case-insensitive protected-name and unreadable
pattern predicates are escaped and anchored to validated roots in the generated
profile so an APFS spelling alias cannot reopen them. On a case-insensitive,
case-preserving APFS volume, macOS 26 may accept a case-only rename such as
.git to .GIT even though Seatbelt denies the protected object and both
rename operation classes. The spelling can change, but it remains the same
single protected object: every case-equivalent spelling stays readable and
non-writable, cannot be moved to an unprotected name, and cannot be hard-linked
into writable space. Each command receives one private mode-0700 temporary
directory through TMPDIR; Crucible owns and removes it with the command
lifecycle. Empty manifests are supported, while a request to materialize files
or mounts is refused before the temporary directory is created.
Seatbelt policy inheritance keeps descendants confined after fork and exec. macOS has no PID namespace or cgroup-equivalent process census here, so cleanup can prove termination of Crucible's owned process group but cannot prove that a hostile descendant which deliberately starts a new session has exited. Such a descendant retains the same Seatbelt filesystem and network restrictions and loses the command's removed private temporary path.
Seatbelt also relies on trusted macOS code-validation services. Native
feasibility testing observed that a confined non-root caller could ask the
system taskgated service to attach code-validation metadata to a controlled
foreign process when the existing signing policy allowed that metadata. The
target's executable bytes and signing flags did not change, and the caller did
not acquire its task port, credentials, entitlements, filesystem access, or
network access. Crucible treats that validation bookkeeping as a trusted-OS
effect rather than guest authority.
Windows setup maintenance
Windows confinement uses a dedicated local account and machine firewall
policy, so an administrator must provision them once before ordinary Crucible
runs can use the native backend. Keep the release's crucible-sandbox-broker.exe
beside crucible.exe. From an Administrator PowerShell in that directory, run:
.\crucible.exe sandbox setupThe command does not auto-elevate and an ordinary Crucible run remains unelevated. If PowerShell was elevated with a different administrator account, name the developer account explicitly:
.\crucible.exe sandbox setup --owner 'MACHINE\person'Setup creates one deterministic local sandbox account for that owner, stores its random password under machine-scope DPAPI in a protected HKLM record, and installs persistent Windows Filtering Platform rules that deny outbound IPv4 and IPv6 connections and socket binding for that account. Re-running the command repairs the exact account, record, and filters. Concurrent maintenance for the same owner is serialized; an interrupted setup leaves bounded state that a later run can repair.
To remove that state, use the same owner choice from an Administrator PowerShell:
.\crucible.exe sandbox uninstallRemoval disables the account before changing its firewall rules, then deletes the account and record. If cleanup fails, the disabled account and protected record remain so the command can be retried safely.
.\crucible-sandbox-broker.exe --windows-sandbox-setup and
--windows-sandbox-uninstall, with the same --owner, run the same setup and
removal from the broker alone.
An enabled command starts through the packaged broker, which logs on that
dedicated account and then creates a WRITE_RESTRICTED token with all ordinary
privileges disabled except the standard directory-traversal privilege required
by ordinary Windows programs. Crucible adds account and path-capability grants
only to declared roots; protected repository and Crucible metadata receive
write-deny capabilities even beneath a writable workspace. Each command gets a
private temporary directory through TEMP and TMP, and Crucible removes it
with the process. The target starts on a private desktop with exactly standard
input, output and error inherited, and the outer broker's kill-on-close Job
Object owns every descendant. The Job also applies the requested aggregate
processor-time limit. WFP denies outbound connections and socket binding for
both the broker account and its descendants.
Windows uses WRITE_RESTRICTED deliberately. A fully read-restricted token can
pass filesystem access probes but cannot start ordinary Win32 tools: the system
loader needs protected KnownDlls object-manager resources whose ACL cannot be
extended by setup, including under SYSTEM. Reads therefore follow the dedicated
low-privilege account's normal Windows access. Files in another user's private
profile remain protected by their own ACL, while system and shared files that
grant ordinary users access retain that access. This also means an existing ACL
for the sandbox account or Everyone can permit writes outside the roots in the
current request, including a root granted to an earlier command. Crucible never
adds such a grant for an undeclared root, and its protected-path deny
capabilities still win for the current request. Requests containing explicit
unreadable roots or unreadable wildcard patterns are refused before the broker
or workspace is touched. Linux and macOS retain their narrower filesystem
views.
The same limitation applies to configured sandbox.filesystem.unreadable
paths. Native Windows also refuses domain policies, allowLocalBinding: true
and nonempty allowUnixSockets before starting the command. It supports the
enabled switch, writable/read-only/protected path rules, all three configured
command limits and the /sandbox panel. Empty network lists with binding
disabled keep WFP network denial. Run Crucible inside WSL2 to use Linux's richer
filesystem and network policies on a Windows machine.
ACL projection adds deterministic capability and account entries to the roots a command uses. They can remain after a command or uninstall because Windows provides no namespace-like per-process filesystem view to remove. A later command using the same sandbox account can therefore retain the access those account entries grant. Uninstalling and reinstalling creates a new account SID, so entries for the deleted account become inert.
Linux and macOS mediate configured domains through a host-owned authenticated
HTTP/CONNECT proxy. Linux exposes a pinned Unix endpoint only to its namespace
broker, which relays it to private loopback. macOS permits only the host proxy's
IPv4 loopback port. Direct connections remain denied; enabling local binding
or selected Unix sockets adds only those separately declared operations.
Proxy credentials are unique to a command, replace inherited proxy settings,
and are masked in captured stdout and stderr without changing byte counts or
MCP framing. A permitted connection leaves through the http:// proxy named
in the environment crucible was started in, unless NO_PROXY lists its host,
as a CONNECT to the address the host proxy checked; the command sees neither
that proxy's address nor its credential. Under an https:// proxy, or a
socks4a:// or socks5h:// one, permitted connections fail with 502 rather
than bypassing it; a socks://, socks4:// or socks5:// proxy is passed by,
as crucible's own requests pass it (commands connect on their
own). Output cut by
the output limit, or from a command crucible stopped, also masks any last few
bytes that could begin the credential.
Listener and relay cleanup belongs to the command lifecycle; failed cleanup
retains its resources and admission slot for recovery.
Platform support
| Platform | Enforcing backend in this version | When enabled |
|---|---|---|
| Linux | Verified system Bubblewrap | Runs only when the backend satisfies the requested policy |
| macOS | Built-in Seatbelt through /usr/bin/sandbox-exec | Runs only when the system launcher and packaged broker pass provenance and functional probes |
| Windows | Dedicated account, ACL capabilities, WFP, restricted token, private desktop and Job Object | Runs only after explicit Administrator setup passes account and network-policy verification |
An unavailable backend never turns enabled: true into permission to run
outside confinement.
SDK callers constructing SandboxPolicy::standard or Bash::new still receive
an enabled policy. The application applies its resolved configuration explicitly;
changing the application's default does not weaken independently constructed
SDK policies or inherited child authority.
Capability matrix
enforced is a hard boundary, observed is bounded measurement only, and
unsupported is rejected whenever the effective policy explicitly requires
that feature. Disabling confinement with SandboxPolicy::with_enabled(false)
removes the baseline kernel isolation. Explicit limits, manifests, network rules,
persistence and snapshot requests still require an enforcing implementation.
| Capability | Linux Bubblewrap | macOS Seatbelt | Windows native | Compatibility |
|---|---|---|---|---|
filesystem | enforced | enforced | enforced | unsupported |
network_deny | enforced | enforced | enforced | unsupported |
network_allowlist | enforced | enforced | unsupported | unsupported |
descriptor_isolation | enforced | enforced | enforced | unsupported |
process_isolation | enforced | enforced | enforced | unsupported |
kernel_surface | enforced | enforced | enforced | unsupported |
privilege_isolation | enforced | enforced | enforced | unsupported |
materialization | enforced | unsupported | unsupported | unsupported |
cpu_limit | enforced | unsupported | enforced | unsupported |
memory_limit | enforced | unsupported | unsupported | unsupported |
disk_limit | unsupported | unsupported | unsupported | unsupported |
process_limit | enforced on Linux 5.14 or newer | unsupported | unsupported | unsupported |
open_file_limit | enforced | enforced | unsupported | unsupported |
command_time_limit | enforced | enforced | enforced | enforced |
session_time_limit | unsupported | unsupported | unsupported | unsupported |
outbound_byte_limit | unsupported | unsupported | unsupported | unsupported |
output_limit | enforced | enforced | enforced | enforced |
concurrency_limit | enforced | enforced | enforced | enforced |
cost_limit | unsupported | unsupported | unsupported | unsupported |
pty | unsupported | unsupported | unsupported | unsupported |
file_operations | unsupported | unsupported | unsupported | unsupported |
persistence | unsupported | unsupported | unsupported | unsupported |
snapshot | unsupported | unsupported | unsupported | unsupported |
resume | unsupported | unsupported | unsupported | unsupported |
audit | enforced | enforced | enforced | enforced |
usage | observed | observed | observed | observed |
A capability being enforced says what a backend can apply, not what the
default policy asks for. With confinement enabled, Linux's standard policy states
cpu_limit at one hour per process and open_file_limit at 4096. Windows
states one hour of aggregate Job processor time and no handle-count limit.
macOS states only the open-file ceiling because its catchable CPU signal is not
a hard boundary. bash adds
command_time_limit, output_limit and concurrency_limit for the one command
it is running. memory_limit is enforced but not asked for: the knob is the
address space a process may map rather than the memory it uses, and runtimes
that reserve enormously and touch little would be refused by any ceiling low
enough to catch a real runaway. disk_limit is unsupported above, and a policy
may not ask for a ceiling its backend cannot apply, which is also why disabling
confinement takes the two confining ceilings off with it, rather than carrying
numbers the compatibility backend would have to refuse.
process_limit is the one ceiling a policy does not have to ask for. The broker
is PID 1 of the namespace, and it caps the processes beneath it at 1024 whether
or not anything states a number, the way it zeroes the core-dump ceiling: a
workload that forks in a loop is otherwise bounded by nothing the sandbox owns,
and the processor and descriptor ceilings do not help, because each new process
gets its own. A policy may state fewer and the broker takes the lower of the
two; it may not state more.
Stating one is what the table's row is about, and it needs a kernel that counts processes per user namespace (Linux 5.14 and newer). Below that the kernel counts them for the real user across the whole machine, so a stated ceiling would bound the host's other work rather than the sandbox's, and enabled confinement refuses the policy instead of applying a number that means something else. The broker's own 1024 still applies there, and under the older counting it can bind before the sandbox has reached 1024 of its own, but the only thing it ever stops is the sandbox forking. Nothing outside the namespace is ended by it.
The currently declared session surface is prepare, materialize, start, inspect, read bounded output, observe usage/violations, stop and dispose. PTY, direct file operations, persistence, snapshots and resume are absent rather than stubs that run outside policy.
Conformance
The table above is a statement, and nothing in a type system can check it. A row that claims more than the backend keeps would put a command behind a fence that is not there; a row that claims less would refuse work for no reason. So the claims are asked about rather than trusted, by a suite Crucible publishes rather than keeps to itself:
use crucible_sandbox_local::conformance::{Conformance, SandboxClaim};
let runtime = tokio::runtime::Builder::new_current_thread()
.enable_all()
.build()?;
let audited = runtime.block_on(Conformance::audit(&backend, workspace_root))?;
assert!(audited.faults().next().is_none());
assert!(audited.holds(SandboxClaim::Isolation));
println!("{}", audited.report());The audit is async and never bounds the wait itself, so the harness drives
it on its own runtime and bounds an unanswering backend there.
For every feature a policy can name, the suite writes the smallest policy that requires exactly that feature and offers it to the backend, in the mode that reaches the backend that was probed. It then reads the claim and the answer as one statement. A claimed feature must be accepted; a disclaimed one must be refused by that name, before any session exists. Two answers are faults and nothing else is: overclaimed, where a backend that says it cannot do something takes the policy anyway, and withheld, where a backend refuses what it says it can do. A backend that fails for its own reasons (no executable, a kernel that will not give it namespaces) leaves the claim unreached, because a host without a sandbox must not read as a backend that lies.
Where there is no backend to probe at all, audit returns the probe's error
instead of a table. A Windows host without its administrator setup, a macOS host
whose Seatbelt or packaged-broker probe fails, or a Linux host without usable
Bubblewrap has no enforcing backend to audit, so a suite that reported the
absence as a fault would be accusing a backend that never claimed readiness.
Whether that error is a skip or a failure belongs to the caller: Crucible's own
native CI jobs require the applicable backend, while an unprovisioned developer
machine may treat it as a skip.
Some features cannot be asked for in a policy at all. A terminal, direct file operations and resuming somebody else's session are asked of a session, and a bare policy offered in their name would be accepted by everyone. Those are reported stated or absent: read, not tested. Saying a claim was exercised when nothing could exercise it is the same failure the suite exists to catch.
The answers are grouped into the families a backend is chosen by: isolation,
materialization, network, resources, terminal, persistence, accounting and
cost. A backend that fences a filesystem perfectly and cannot bound one byte of
egress is not partly conformant. It holds one family and not another, and a
policy is matched against the family it needs. holds says a family is exact,
not that it is supported: a backend that disclaims every network feature and
refuses every network policy holds that family, and the claims are what tell a
caller it can do nothing there.
This is published because the backends that must pass it are not all in this
repository. A container, a Kubernetes executor, a hosted session or another
operating system's adapter can depend on crucible-sandbox-local, run the
suite over a directory it owns and get the same verdicts from the same table,
before it is wired to anything. Nothing is materialized and no command is
started: each session is prepared to see whether it can be, and dropped.
Unconfined execution
Only home/user configuration may explicitly disable confinement. Project and
descendant policy may preserve or strengthen that choice but never weaken it.
When confinement is disabled, inspection and audit records say confined: false,
name the disabled boundary and retain the exact compatibility capability snapshot.
There is no enforcing backend on FreeBSD. Commands run through compatibility by default and are reported as unconfined. Enabling the sandbox there refuses a command before it starts; it does not fall back silently. Linux selects Bubblewrap, macOS selects Seatbelt, and Windows selects its native backend after the administrator setup passes verification.
The SDK uses the same boolean choice: SandboxPolicy::enabled() and
SandboxPolicy::with_enabled(bool). There are no sandbox modes or degraded
fallbacks. Checkpoints use format 3 with a boolean enabled field; obsolete
checkpoint formats are rejected. Sandbox audit plans also record enabled,
and deliberately unconfined execution includes a bounded disabled_reason.
Compatibility still clears and explicitly rebuilds the command environment, checks requested and transformed command guardrails, enforces command deadlines, captured-output and concurrency ceilings, supervises its owned process scope, records bounded usage and emits lifecycle audit facts. It does not restrict filesystem or network reach and must not be described as a sandbox. Its process-isolation capability is explicitly unsupported; unlike the enforcing PID-namespace backend, it cannot promise containment of a hostile process that deliberately escapes the owned process group.
Environment and credentials
The command environment is an explicit, bounded map rather than a copy of the
host environment. Linux supplies only a private HOME and TMPDIR; macOS
supplies a private TMPDIR; Windows supplies only the variables selected by the
host and replaces TEMP and TMP with one private command directory. SSH/GPG
agent sockets, inherited descriptors or handles, provider keys, cloud
configuration and arbitrary host variables do not cross the boundary
automatically. Values reach the command through the backend's cleared process
environment, never through its argument list. The helper each backend starts
the command through is started from a cleared environment too: on Windows,
crucible-sandbox-broker.exe keeps only SystemRoot from crucible's own.
A secret projection carries a bounded opaque credential handle and user/account
provenance alongside the host-resolved value. Handles and values are redacted
from debugging, inspection, audit, JSONL and diagnostics. Credential variables
share the ordinary environment count, name, uniqueness, NUL and aggregate-byte
bounds; a credential cannot silently replace a literal variable with the same
name. Every non-empty credential value is masked the way proxy credentials
are, one * per byte wherever its exact bytes appear, without changing byte
counts: in captured stdout and stderr of an ordinary command, and in stderr
alone of a command crucible speaks a protocol to, such as an MCP server. On
that command's stdout no credential value is masked, so no frame is rewritten
by one; only crucible's own proxy credential is masked there, when the command
has network domains. The MCP client hides each value in the words of every
reply it keeps after decoding it, which also finds a value the reply escaped;
a number it keeps, such as an error code, is shown as sent. A value the
command re-encodes, splits or transforms before printing it is not matched. On
Windows the output masker matches a value's bytes as crucible holds them, which
are UTF-8 for any valid Unicode value, so a value printed in UTF-16 or an ANSI
code page is not matched there; the decoded hiding in MCP replies applies on
every platform.
Lifecycle and inspection
Policy resolution, capability negotiation, materialization, both guardrail decisions, command start/finish, violations, usage and cleanup are bounded typed facts carrying the original run ancestry, tool call and sandbox ID. They are written to the framework journal before their live events. Detached commands retain the same fixed attribution; facts produced after the starting tool call returns are drained at the next runner boundary.
Selected MCP servers use the same audit registry. Their preparation, restarts
and cleanup keep the lifecycle's run ancestry and mcp:server identity, with a
new sandbox ID for each preparation. The runner delivers lifecycle facts before
the next provider request and after disposal, including when preparation,
snapshot, refresh or disposal fails. An audit delivery failure ends the turn
without hiding an operation or cleanup failure that also occurred. MCP restart
requires confirmed cleanup of the previous process scope. Unconfirmed cleanup
remains a disposal failure and blocks another preparation, including when a
later server failed during partial preparation. Missing pipes and failed
handshake or catalogue exchanges also require explicit cleanup. An optional
server can be skipped only when its cleanup is confirmed; otherwise preparation
stops and later disposal retains the failure. A confined server's writes stay
private while it runs. They are published when it exits on its own after
crucible closes its input, at a restart or at the end of the run, and
discarded when it has to be stopped instead. A server that has exited is not
stopped while its writes wait for another command's publication. Where they
cannot be published, because a root it wrote into changed while it ran, the
restart or the disposal fails with that reason. No replacement is started behind
that call, because what it wrote is in a state only a fresh start should settle;
the next turn starts the server as usual.
Inspection retains backend ID/version/provenance, capability claims, separate hashed requested and effective policies and redacted plans, manifest, working-directory and root identities, root access/provenance, network shape, requested limits, unreadable-pattern counts, command-policy digest, degradation and cleanup state. Parent filesystem carve-outs, command filters, network authority, resource ceilings, session grants and unreadable patterns cannot be dropped or relabelled by a descendant. It does not retain command arguments, environment values, endpoint names, raw approval proofs, credential handles or values, proxy material or raw out-of-scope paths.
The local supervisor attempts to stop the owned process scope and reap its leader on normal exit, deadline, output violation, cancellation, refusal, launch failure, panic, explicit stop and ordinary host shutdown. On enforcing Linux a deadline or output ceiling first tells the broker to end the workload, so the workload's own wait status is still reported, and kills the launcher once the broker has exited or after a short budget. The local process owner used by compatibility mode retries unfinished cleanup within bounded waits. The enforcing Linux projection owner retains its separate terminal failure and quarantine outcome. Successful cleanup is idempotent; a failed attempt does not become successful just because it is called again. Staging data and the command's admission slot remain held until process cleanup is confirmed. If the process owner is dropped while cleanup is still uncertain, staging data is retained and the slot remains consumed for the lifetime of that sandbox service. A recorded supervisor failure stays visible even if later cleanup releases those resources.
These ownership rules also apply when startup fails after a process has been spawned. Startup cleanup uses the same bounded stop and reap path. An uncertain result reports failed cleanup and retains its staging data and command slot; the enforcing Linux backend also retains its separate projection and journal for recovery. A failure to record another audit fact cannot replace the original startup or cleanup error.
An uncatchable host/process kill cannot run user-space destructors. On enforcing Linux, loss of the broker status channel and Bubblewrap's parent-death boundary still terminate the PID-namespace workload. The next preparation replays the checksummed lifecycle WAL, rolls back an unambiguous interrupted transaction, and quarantines ambiguous publication instead of inventing cleanup or success.
Reading the report
crucible sandbox inspect prints that inspection report for the directory you
are standing in and stops; crucible --sandbox prints the same report. Nothing
is started to produce it: not the backend, not its broker, not a command. No
sandbox is prepared, nothing is materialized and nothing is written. The
backend reported is the first one in the places a command's preparation looks
that passes the same trust checks on its owner, on who can write it and on
whether you may run it, and build is the SHA-256 of that file. Preparation
also starts each candidate to check it, and passes over one that fails for the
next, so where a trusted file will not start, the backend a command gets can be
a later one than the report names, or none, and then no command is run.
sandbox enabled in <root>
mode required by project configuration
backend linux-bubblewrap, system
version unverified; reading it would start Bubblewrap, and inspecting starts nothing
build sha256:523da3e7399044be5163aee6f57a77a6bef7454376e28f0a0627920bae1b76b6
what this backend can hold:
filesystem enforced
network_deny enforced
network_allowlist enforced
descriptor_isolation enforced
process_isolation enforced
kernel_surface enforced
privilege_isolation enforced
materialization enforced
cpu_limit enforced
memory_limit enforced
disk_limit unsupported
process_limit enforced
open_file_limit enforced
command_time_limit enforced
session_time_limit unsupported
outbound_byte_limit unsupported
output_limit enforced
concurrency_limit enforced
cost_limit unsupported
pty unsupported
file_operations unsupported
persistence unsupported
snapshot unsupported
resume unsupported
audit enforced
usage observed
what a command would run under:
enabled true
cwd sha256:de761e35c85ea9a7615fe3e45e1fef07d6631d0f02c4de318c94d05f30813ca2
reach 2 places, named by digest
read_write workspace sha256:cd490a63f17abf417975914cb065fde6035ea5875cc4e86e6c338842c0cd6262
protected protected_metadata sha256:1c6e26020fe934c29b865b19537f043598e19249d3a277492a790690551c8955
hidden 0 patterns
network closed
ceilings
cpu 60m enforced
4096 open files enforced
10 MiB captured enforced
4 at once enforced
20m per command enforced
staged nothing
outlives no
snapshots no
confined yes
policy sha256:ad76ba3f9182bb58bad76fc0c8167c7d7170c2c662ad7521ac6a90ed837c90ef
manifest sha256:72c169df3655f6857b163163b2618631ec65982557eea70b57ea1cb983ae0890
not checked, since checking would start something:
whether Bubblewrap starts and can make its namespaces on this host, and whether each root passes the checks a command's preparation makes
What only starting the backend could tell is reported as unverified, with the
reason, rather than guessed: the version line says so, and the last section
lists what was left unchecked. Those are the questions a command's preparation
answers when it starts the backend, so a report that says confined yes says
that the backend found claims this policy, not that it has been seen to start
on this host. In the sample the workspace root is shown as <root>.
Every ceiling is printed beside the claim it rests on, because a number the
backend only observes is recorded after the fact rather than imposed, and the
two read identically until they are on one line. The capability matrix lists
every feature, including the ones this policy never asks for; it is what makes
an enforced elsewhere in the report mean something.
Time ceilings retain their exact value, including fractional seconds: a 1.5-second
limit is shown as 1.5s, and a limit below one second is never rounded down to 0s.
The workspace root is the only path printed. Roots, the working directory and the policies are named by the digests the record keeps, so a report can be pasted into an issue without pasting your tree with it. Two reports over the same directory produce the same digests, which is what makes them comparable between machines.
Where a backend is found but its matrix will not hold this workspace's policy,
the matrix is printed, then what was asked for, then no command could be run here with the reason. An unsupported row in the matrix is usually the whole
explanation. Where no backend is found at all, the report says no sandbox backend was found with the reason, rather than printing a matrix of claims
belonging to nobody. Neither is an error: both are the answer, and the run
ends successfully.
crucible sandbox inspect --json writes the same report as one JSON document
on one line, with format_version 1, kind sandbox-inspection, and a
status of ready, refused, unavailable or failed. It carries no path
at all, not even the workspace root. Its fields and exit statuses are listed
under Command line.
Writable roots and publication
On enforcing Linux a writable root is not handed to the command directly. The command writes into a private projection of that root, and the host publishes the changed paths back only after the command has ended. Until then nothing the command wrote is visible outside the sandbox, and a reader of the workspace sees the root exactly as it was when the command started.
That projection is an overlay, and Bubblewrap 0.11 takes no mount options for
one, so it is the single mount inside the sandbox that would honour a setuid
bit or a device node. Neither is a way out. The namespace maps one unprivileged
identity and nothing above it, so a file the command marks setuid can only ever
name the identity it already runs as, and a mknod for a host disc is refused
because the command holds no capability to make a device with. Every other
mount, including the read-only runtime and the whole of /dev apart from the
device nodes themselves, is mounted nosuid and nodev. The tests assert the
behaviour rather than the flag, because the flag is the one thing here that
cannot be set.
Because the projection is an overlay over the root itself, a publication into that root while the command runs changes what the command sees beneath its own writes. The command's own publication is then refused, as described below, so such a change can confuse a running command but cannot reach the root through it.
Publication is decided by how the command ended:
- An ordinary exit, zero or nonzero, publishes the changed paths. A failing build still leaves the files it wrote, as it would have without confinement.
- Termination by a signal, a deadline, Esc, an output ceiling or a refusal discards the projection. Nothing partial reaches the workspace. A command that has already ended is not stopped by a deadline or Esc while its writes wait their turn to publish, until the ceiling below passes.
- A command that wrote into a root which changed after it started, whether another command published into it or something outside the sandbox wrote to it, publishes nothing. The delta is discarded rather than merged, and the result says so: the call's own result for a command that was waited for, the note about its ending for a command left running, and the restart or disposal for a confined MCP server. A root the command left unchanged is not checked, so a command that wrote nothing into a changed root is not refused for it.
Publication itself is transactional. The changed paths are staged in this
user's private sandbox state directory under /var/tmp, which no other user can
read or enter, journaled in a checksummed write-ahead log, applied, and
verified. A projected file larger than 8 GiB is refused rather than digested, so
a sparse file cannot make publication read through the whole of it. A
failure between those steps rolls the root back to its pinned baseline. Where a
rollback cannot itself be proved, the staged content is retained as quarantine
evidence and the cleanup outcome reports it, rather than deleting what cannot be
accounted for. That evidence records the root as its own transaction last saw
it, and another command may have published into the root since, so compare it
with the root's current content before restoring anything from it. The next
preparation recovers any transaction an earlier process abandoned.
The state directory's shipped name,
/var/tmp/crucible-code-sandbox-<uid>-v1 where <uid> is the user's numeric
id, is predictable, so another local user can create it first, whether a
directory, a symlink or a plain file. crucible refuses to use what it finds
there and reports that the sandbox backend is unavailable, naming which of
those it is without the path or the other user's numeric id; removing it,
which may need an administrator, is what lets the sandbox be used. The
directory can instead already be this user's own but carry the wrong group or
permissions, for example after running crucible under a different group;
crucible refuses that too, but says to restore its group and mode instead of
removing it, since the directory holds this user's own unrecovered
transaction journals and quarantine evidence.
Commands that can write run side by side, including a command left running in the background and a confined MCP server, and none of them holds up another while it runs. Publication is what has to happen alone. A host-owned lock, one for each user, is held while a command's publication is checked against its baseline, applied and verified, and while a starting command takes its baselines, so no two publications overlap and no baseline is taken halfway through one. A command that ends cleanly while another is publishing waits for that publication to finish before its own begins; one killed, or stopped by a limit, has nothing to publish and does not wait. Waiting commands are not served in order, and each keeps its place among the commands allowed to run at once until its own publication finishes. Every wait for the lock has a ceiling: a command being prepared is refused if the lock does not come free within a minute, a command whose own deadline has passed is stopped after a minute, and five seconds is the bound everywhere somebody is waiting: a cancelled turn, a command left running whose report is overdue, a confined server being restarted or disposed of at a turn's end, and the run itself ending.
A command that ran while another published into one of its roots publishes nothing, whatever the root looks like afterwards: this user's state directory remembers how many publications have touched each root, and a command compares that count with the one it recorded when it took its baseline. A crucible from before this release takes the same lock, so the two still publish one at a time, but it does not keep that count. Against such a peer the comparison has nothing to see, and only the check against the root's own content remains. The holder may be another crucible of this user, including one from before this release, which keeps the lock for as long as its commands run; a wait with no end would hold up the turn, the cancel and the exit instead.
A detached command follows the same rules when it ends later. Its start result is accepted only after it is durably stored, and its terminal publication is journaled under the same call identity, so a host restart between the two neither loses the command nor publishes it twice.