Augments LabsCrucible Code

Operating-system confinement

Every bash command crosses a host-owned sandbox service before any command process starts. Tool arguments describe the command, not the backend. The host resolves one immutable filesystem, network, environment, command and resource plan; probes the backend; refuses missing hard capabilities; materializes inert inputs; and only then launches the process.

OS confinement is disabled by default. Enable it with {"sandbox":{"enabled":true}}; enabled: false in your home configuration selects unconfined execution. Approval, minimal environment, command deadlines, output bounds and audit remain active whether confinement is enabled or disabled. Configure filesystem grants and restrictions, domain mediation, local sockets and command limits in the same block. Project settings can require confinement and narrow inherited authority. See configuration.

Turning it on

  1. Give the platform what it needs. Confinement runs only where this machine can enforce it; while it is on and this machine cannot, a command is refused rather than run unconfined.
    • Linux: Bubblewrap 0.11.0 or newer from the system's packages (bwrap --version says which you have), and crucible-sandbox-broker beside crucible, which install.sh puts there.
    • macOS: nothing to install. The built-in /usr/bin/sandbox-exec does the confining, with crucible-sandbox-broker beside crucible.
    • Windows: crucible-sandbox-broker.exe beside crucible.exe, and once, from an Administrator PowerShell in that directory, .\crucible.exe sandbox setup, adding --owner with your account when that PowerShell runs as another administrator. It creates the account and network rules described under Windows setup maintenance.
  2. Turn it on with /sandbox enable in a conversation, or with {"sandbox":{"enabled":true}} in ~/.crucible/config.json. The command writes that same setting, but only once it has checked that this machine can enforce it; when it cannot, it says why and changes nothing.
  3. Check it. crucible --sandbox, run in the project, prints the backend and what a command there would run under, without running one (reading the report). In a conversation, /sandbox shows the settings, and its Dependencies tab says whether this machine's backend is available and, when it is not, why.
  4. Turn it off with /sandbox disable, unless the project requires it. On Windows, .\crucible.exe sandbox uninstall, from an Administrator PowerShell in the same directory and with the same --owner setup was given, removes what setup created.

How Linux and macOS confine a command

On Linux, the production backend uses a canonical, root-owned, non-writable system Bubblewrap executable reached only through root-owned, non-writable parent directories. Every bounded PATH candidate is tried until one passes its exact command-surface, numeric version, namespace and SHA-256 identity probes. This release records that the bundled choice is unavailable; it does not silently substitute a plain subprocess, Landlock, a container, or a worktree. If no system candidate satisfies the effective policy, preparation fails before materialization or spawn.

Bubblewrap 0.11.0 or newer is needed: the view is built from descriptor binds and temporary overlays, which older releases do not offer. The probe checks the options themselves rather than the version, because distributions backport some of them. Ubuntu 24.04's 0.9.0 has the descriptor binds but not the overlays, so there enabled confinement reports an unavailable backend until a newer Bubblewrap is installed. Should a launch still be refused by the system Bubblewrap, its own message is quoted in the error.

Inside the namespace, PID 1 is Crucible's own crucible-sandbox-broker executable, shipped in every Linux release archive and installed beside the Crucible binary, where it is found and pinned by open descriptor before the namespace starts. It is accepted only when it and every directory above it belong to root or to the user running Crucible and are writable by neither group nor others; a copy under /tmp, in another user's directory or below a group-writable directory is ignored, and chmod g-w on the offending directory is the remedy. When no copy qualifies, the error names each place it looked, beside the Crucible binary or in the directory above it, and what turned it down, so a broker that was never built reads differently from one below a directory others can rewrite.

The Linux view starts from an empty temporary root. It exposes only the minimal read-only runtime needed to execute the selected absolute program, the exact workspace/reached roots at their granted access, protected repository control metadata and recognized agent configuration directories, a minimal /proc and /dev, and a transactionally staged manifest. Bounded unreadable patterns use a deliberately small */single-** grammar and one deterministic, no-symlink, no-mount-crossing tree scan. It creates isolated user, PID, IPC, UTS and network namespaces, drops capabilities, sets no-new-privileges through Bubblewrap, disables nested user namespaces, clears the environment, and closes every undeclared file descriptor. Closed networking has no usable host, loopback, Unix-socket, DNS, metadata-service or inherited-socket route. Killing the namespace owner also kills descendants that deliberately leave the original process group or session.

On macOS, the production backend uses the built-in Seatbelt framework through the fixed /usr/bin/sandbox-exec executable. The launcher must be root-owned and non-writable. Crucible's separately packaged sandbox broker must be beside the main executable on a path writable only by root or the current user; it closes every inherited descriptor above standard input, output and error, applies the requested open-file limit, and then replaces itself with the system launcher. Darwin's CPU rlimit delivers a catchable signal and is therefore not advertised as a hard CPU ceiling.

The generated Seatbelt profile is closed by default. It reads only fixed system-runtime directories plus the declared filesystem roots, writes only declared read-write roots, carves protected and unreadable paths back out, and denies networking. Declared grants and exact carve-outs are supplied as Seatbelt parameters. Case-insensitive protected-name and unreadable pattern predicates are escaped and anchored to validated roots in the generated profile so an APFS spelling alias cannot reopen them. On a case-insensitive, case-preserving APFS volume, macOS 26 may accept a case-only rename such as .git to .GIT even though Seatbelt denies the protected object and both rename operation classes. The spelling can change, but it remains the same single protected object: every case-equivalent spelling stays readable and non-writable, cannot be moved to an unprotected name, and cannot be hard-linked into writable space. Each command receives one private mode-0700 temporary directory through TMPDIR; Crucible owns and removes it with the command lifecycle. Empty manifests are supported, while a request to materialize files or mounts is refused before the temporary directory is created.

Seatbelt policy inheritance keeps descendants confined after fork and exec. macOS has no PID namespace or cgroup-equivalent process census here, so cleanup can prove termination of Crucible's owned process group but cannot prove that a hostile descendant which deliberately starts a new session has exited. Such a descendant retains the same Seatbelt filesystem and network restrictions and loses the command's removed private temporary path.

Seatbelt also relies on trusted macOS code-validation services. Native feasibility testing observed that a confined non-root caller could ask the system taskgated service to attach code-validation metadata to a controlled foreign process when the existing signing policy allowed that metadata. The target's executable bytes and signing flags did not change, and the caller did not acquire its task port, credentials, entitlements, filesystem access, or network access. Crucible treats that validation bookkeeping as a trusted-OS effect rather than guest authority.

Windows setup maintenance

Windows confinement uses a dedicated local account and machine firewall policy, so an administrator must provision them once before ordinary Crucible runs can use the native backend. Keep the release's crucible-sandbox-broker.exe beside crucible.exe. From an Administrator PowerShell in that directory, run:

.\crucible.exe sandbox setup

The command does not auto-elevate and an ordinary Crucible run remains unelevated. If PowerShell was elevated with a different administrator account, name the developer account explicitly:

.\crucible.exe sandbox setup --owner 'MACHINE\person'

Setup creates one deterministic local sandbox account for that owner, stores its random password under machine-scope DPAPI in a protected HKLM record, and installs persistent Windows Filtering Platform rules that deny outbound IPv4 and IPv6 connections and socket binding for that account. Re-running the command repairs the exact account, record, and filters. Concurrent maintenance for the same owner is serialized; an interrupted setup leaves bounded state that a later run can repair.

To remove that state, use the same owner choice from an Administrator PowerShell:

.\crucible.exe sandbox uninstall

Removal disables the account before changing its firewall rules, then deletes the account and record. If cleanup fails, the disabled account and protected record remain so the command can be retried safely.

.\crucible-sandbox-broker.exe --windows-sandbox-setup and --windows-sandbox-uninstall, with the same --owner, run the same setup and removal from the broker alone.

An enabled command starts through the packaged broker, which logs on that dedicated account and then creates a WRITE_RESTRICTED token with all ordinary privileges disabled except the standard directory-traversal privilege required by ordinary Windows programs. Crucible adds account and path-capability grants only to declared roots; protected repository and Crucible metadata receive write-deny capabilities even beneath a writable workspace. Each command gets a private temporary directory through TEMP and TMP, and Crucible removes it with the process. The target starts on a private desktop with exactly standard input, output and error inherited, and the outer broker's kill-on-close Job Object owns every descendant. The Job also applies the requested aggregate processor-time limit. WFP denies outbound connections and socket binding for both the broker account and its descendants.

Windows uses WRITE_RESTRICTED deliberately. A fully read-restricted token can pass filesystem access probes but cannot start ordinary Win32 tools: the system loader needs protected KnownDlls object-manager resources whose ACL cannot be extended by setup, including under SYSTEM. Reads therefore follow the dedicated low-privilege account's normal Windows access. Files in another user's private profile remain protected by their own ACL, while system and shared files that grant ordinary users access retain that access. This also means an existing ACL for the sandbox account or Everyone can permit writes outside the roots in the current request, including a root granted to an earlier command. Crucible never adds such a grant for an undeclared root, and its protected-path deny capabilities still win for the current request. Requests containing explicit unreadable roots or unreadable wildcard patterns are refused before the broker or workspace is touched. Linux and macOS retain their narrower filesystem views.

The same limitation applies to configured sandbox.filesystem.unreadable paths. Native Windows also refuses domain policies, allowLocalBinding: true and nonempty allowUnixSockets before starting the command. It supports the enabled switch, writable/read-only/protected path rules, all three configured command limits and the /sandbox panel. Empty network lists with binding disabled keep WFP network denial. Run Crucible inside WSL2 to use Linux's richer filesystem and network policies on a Windows machine.

ACL projection adds deterministic capability and account entries to the roots a command uses. They can remain after a command or uninstall because Windows provides no namespace-like per-process filesystem view to remove. A later command using the same sandbox account can therefore retain the access those account entries grant. Uninstalling and reinstalling creates a new account SID, so entries for the deleted account become inert.

Linux and macOS mediate configured domains through a host-owned authenticated HTTP/CONNECT proxy. Linux exposes a pinned Unix endpoint only to its namespace broker, which relays it to private loopback. macOS permits only the host proxy's IPv4 loopback port. Direct connections remain denied; enabling local binding or selected Unix sockets adds only those separately declared operations. Proxy credentials are unique to a command, replace inherited proxy settings, and are masked in captured stdout and stderr without changing byte counts or MCP framing. A permitted connection leaves through the http:// proxy named in the environment crucible was started in, unless NO_PROXY lists its host, as a CONNECT to the address the host proxy checked; the command sees neither that proxy's address nor its credential. Under an https:// proxy, or a socks4a:// or socks5h:// one, permitted connections fail with 502 rather than bypassing it; a socks://, socks4:// or socks5:// proxy is passed by, as crucible's own requests pass it (commands connect on their own). Output cut by the output limit, or from a command crucible stopped, also masks any last few bytes that could begin the credential. Listener and relay cleanup belongs to the command lifecycle; failed cleanup retains its resources and admission slot for recovery.

Platform support

PlatformEnforcing backend in this versionWhen enabled
LinuxVerified system BubblewrapRuns only when the backend satisfies the requested policy
macOSBuilt-in Seatbelt through /usr/bin/sandbox-execRuns only when the system launcher and packaged broker pass provenance and functional probes
WindowsDedicated account, ACL capabilities, WFP, restricted token, private desktop and Job ObjectRuns only after explicit Administrator setup passes account and network-policy verification

An unavailable backend never turns enabled: true into permission to run outside confinement.

SDK callers constructing SandboxPolicy::standard or Bash::new still receive an enabled policy. The application applies its resolved configuration explicitly; changing the application's default does not weaken independently constructed SDK policies or inherited child authority.

Capability matrix

enforced is a hard boundary, observed is bounded measurement only, and unsupported is rejected whenever the effective policy explicitly requires that feature. Disabling confinement with SandboxPolicy::with_enabled(false) removes the baseline kernel isolation. Explicit limits, manifests, network rules, persistence and snapshot requests still require an enforcing implementation.

CapabilityLinux BubblewrapmacOS SeatbeltWindows nativeCompatibility
filesystemenforcedenforcedenforcedunsupported
network_denyenforcedenforcedenforcedunsupported
network_allowlistenforcedenforcedunsupportedunsupported
descriptor_isolationenforcedenforcedenforcedunsupported
process_isolationenforcedenforcedenforcedunsupported
kernel_surfaceenforcedenforcedenforcedunsupported
privilege_isolationenforcedenforcedenforcedunsupported
materializationenforcedunsupportedunsupportedunsupported
cpu_limitenforcedunsupportedenforcedunsupported
memory_limitenforcedunsupportedunsupportedunsupported
disk_limitunsupportedunsupportedunsupportedunsupported
process_limitenforced on Linux 5.14 or newerunsupportedunsupportedunsupported
open_file_limitenforcedenforcedunsupportedunsupported
command_time_limitenforcedenforcedenforcedenforced
session_time_limitunsupportedunsupportedunsupportedunsupported
outbound_byte_limitunsupportedunsupportedunsupportedunsupported
output_limitenforcedenforcedenforcedenforced
concurrency_limitenforcedenforcedenforcedenforced
cost_limitunsupportedunsupportedunsupportedunsupported
ptyunsupportedunsupportedunsupportedunsupported
file_operationsunsupportedunsupportedunsupportedunsupported
persistenceunsupportedunsupportedunsupportedunsupported
snapshotunsupportedunsupportedunsupportedunsupported
resumeunsupportedunsupportedunsupportedunsupported
auditenforcedenforcedenforcedenforced
usageobservedobservedobservedobserved

A capability being enforced says what a backend can apply, not what the default policy asks for. With confinement enabled, Linux's standard policy states cpu_limit at one hour per process and open_file_limit at 4096. Windows states one hour of aggregate Job processor time and no handle-count limit. macOS states only the open-file ceiling because its catchable CPU signal is not a hard boundary. bash adds command_time_limit, output_limit and concurrency_limit for the one command it is running. memory_limit is enforced but not asked for: the knob is the address space a process may map rather than the memory it uses, and runtimes that reserve enormously and touch little would be refused by any ceiling low enough to catch a real runaway. disk_limit is unsupported above, and a policy may not ask for a ceiling its backend cannot apply, which is also why disabling confinement takes the two confining ceilings off with it, rather than carrying numbers the compatibility backend would have to refuse.

process_limit is the one ceiling a policy does not have to ask for. The broker is PID 1 of the namespace, and it caps the processes beneath it at 1024 whether or not anything states a number, the way it zeroes the core-dump ceiling: a workload that forks in a loop is otherwise bounded by nothing the sandbox owns, and the processor and descriptor ceilings do not help, because each new process gets its own. A policy may state fewer and the broker takes the lower of the two; it may not state more.

Stating one is what the table's row is about, and it needs a kernel that counts processes per user namespace (Linux 5.14 and newer). Below that the kernel counts them for the real user across the whole machine, so a stated ceiling would bound the host's other work rather than the sandbox's, and enabled confinement refuses the policy instead of applying a number that means something else. The broker's own 1024 still applies there, and under the older counting it can bind before the sandbox has reached 1024 of its own, but the only thing it ever stops is the sandbox forking. Nothing outside the namespace is ended by it.

The currently declared session surface is prepare, materialize, start, inspect, read bounded output, observe usage/violations, stop and dispose. PTY, direct file operations, persistence, snapshots and resume are absent rather than stubs that run outside policy.

Conformance

The table above is a statement, and nothing in a type system can check it. A row that claims more than the backend keeps would put a command behind a fence that is not there; a row that claims less would refuse work for no reason. So the claims are asked about rather than trusted, by a suite Crucible publishes rather than keeps to itself:

use crucible_sandbox_local::conformance::{Conformance, SandboxClaim};
 
let runtime = tokio::runtime::Builder::new_current_thread()
    .enable_all()
    .build()?;
let audited = runtime.block_on(Conformance::audit(&backend, workspace_root))?;
assert!(audited.faults().next().is_none());
assert!(audited.holds(SandboxClaim::Isolation));
println!("{}", audited.report());

The audit is async and never bounds the wait itself, so the harness drives it on its own runtime and bounds an unanswering backend there.

For every feature a policy can name, the suite writes the smallest policy that requires exactly that feature and offers it to the backend, in the mode that reaches the backend that was probed. It then reads the claim and the answer as one statement. A claimed feature must be accepted; a disclaimed one must be refused by that name, before any session exists. Two answers are faults and nothing else is: overclaimed, where a backend that says it cannot do something takes the policy anyway, and withheld, where a backend refuses what it says it can do. A backend that fails for its own reasons (no executable, a kernel that will not give it namespaces) leaves the claim unreached, because a host without a sandbox must not read as a backend that lies.

Where there is no backend to probe at all, audit returns the probe's error instead of a table. A Windows host without its administrator setup, a macOS host whose Seatbelt or packaged-broker probe fails, or a Linux host without usable Bubblewrap has no enforcing backend to audit, so a suite that reported the absence as a fault would be accusing a backend that never claimed readiness. Whether that error is a skip or a failure belongs to the caller: Crucible's own native CI jobs require the applicable backend, while an unprovisioned developer machine may treat it as a skip.

Some features cannot be asked for in a policy at all. A terminal, direct file operations and resuming somebody else's session are asked of a session, and a bare policy offered in their name would be accepted by everyone. Those are reported stated or absent: read, not tested. Saying a claim was exercised when nothing could exercise it is the same failure the suite exists to catch.

The answers are grouped into the families a backend is chosen by: isolation, materialization, network, resources, terminal, persistence, accounting and cost. A backend that fences a filesystem perfectly and cannot bound one byte of egress is not partly conformant. It holds one family and not another, and a policy is matched against the family it needs. holds says a family is exact, not that it is supported: a backend that disclaims every network feature and refuses every network policy holds that family, and the claims are what tell a caller it can do nothing there.

This is published because the backends that must pass it are not all in this repository. A container, a Kubernetes executor, a hosted session or another operating system's adapter can depend on crucible-sandbox-local, run the suite over a directory it owns and get the same verdicts from the same table, before it is wired to anything. Nothing is materialized and no command is started: each session is prepared to see whether it can be, and dropped.

Unconfined execution

Only home/user configuration may explicitly disable confinement. Project and descendant policy may preserve or strengthen that choice but never weaken it. When confinement is disabled, inspection and audit records say confined: false, name the disabled boundary and retain the exact compatibility capability snapshot.

There is no enforcing backend on FreeBSD. Commands run through compatibility by default and are reported as unconfined. Enabling the sandbox there refuses a command before it starts; it does not fall back silently. Linux selects Bubblewrap, macOS selects Seatbelt, and Windows selects its native backend after the administrator setup passes verification.

The SDK uses the same boolean choice: SandboxPolicy::enabled() and SandboxPolicy::with_enabled(bool). There are no sandbox modes or degraded fallbacks. Checkpoints use format 3 with a boolean enabled field; obsolete checkpoint formats are rejected. Sandbox audit plans also record enabled, and deliberately unconfined execution includes a bounded disabled_reason.

Compatibility still clears and explicitly rebuilds the command environment, checks requested and transformed command guardrails, enforces command deadlines, captured-output and concurrency ceilings, supervises its owned process scope, records bounded usage and emits lifecycle audit facts. It does not restrict filesystem or network reach and must not be described as a sandbox. Its process-isolation capability is explicitly unsupported; unlike the enforcing PID-namespace backend, it cannot promise containment of a hostile process that deliberately escapes the owned process group.

Environment and credentials

The command environment is an explicit, bounded map rather than a copy of the host environment. Linux supplies only a private HOME and TMPDIR; macOS supplies a private TMPDIR; Windows supplies only the variables selected by the host and replaces TEMP and TMP with one private command directory. SSH/GPG agent sockets, inherited descriptors or handles, provider keys, cloud configuration and arbitrary host variables do not cross the boundary automatically. Values reach the command through the backend's cleared process environment, never through its argument list. The helper each backend starts the command through is started from a cleared environment too: on Windows, crucible-sandbox-broker.exe keeps only SystemRoot from crucible's own.

A secret projection carries a bounded opaque credential handle and user/account provenance alongside the host-resolved value. Handles and values are redacted from debugging, inspection, audit, JSONL and diagnostics. Credential variables share the ordinary environment count, name, uniqueness, NUL and aggregate-byte bounds; a credential cannot silently replace a literal variable with the same name. Every non-empty credential value is masked the way proxy credentials are, one * per byte wherever its exact bytes appear, without changing byte counts: in captured stdout and stderr of an ordinary command, and in stderr alone of a command crucible speaks a protocol to, such as an MCP server. On that command's stdout no credential value is masked, so no frame is rewritten by one; only crucible's own proxy credential is masked there, when the command has network domains. The MCP client hides each value in the words of every reply it keeps after decoding it, which also finds a value the reply escaped; a number it keeps, such as an error code, is shown as sent. A value the command re-encodes, splits or transforms before printing it is not matched. On Windows the output masker matches a value's bytes as crucible holds them, which are UTF-8 for any valid Unicode value, so a value printed in UTF-16 or an ANSI code page is not matched there; the decoded hiding in MCP replies applies on every platform.

Lifecycle and inspection

Policy resolution, capability negotiation, materialization, both guardrail decisions, command start/finish, violations, usage and cleanup are bounded typed facts carrying the original run ancestry, tool call and sandbox ID. They are written to the framework journal before their live events. Detached commands retain the same fixed attribution; facts produced after the starting tool call returns are drained at the next runner boundary.

Selected MCP servers use the same audit registry. Their preparation, restarts and cleanup keep the lifecycle's run ancestry and mcp:server identity, with a new sandbox ID for each preparation. The runner delivers lifecycle facts before the next provider request and after disposal, including when preparation, snapshot, refresh or disposal fails. An audit delivery failure ends the turn without hiding an operation or cleanup failure that also occurred. MCP restart requires confirmed cleanup of the previous process scope. Unconfirmed cleanup remains a disposal failure and blocks another preparation, including when a later server failed during partial preparation. Missing pipes and failed handshake or catalogue exchanges also require explicit cleanup. An optional server can be skipped only when its cleanup is confirmed; otherwise preparation stops and later disposal retains the failure. A confined server's writes stay private while it runs. They are published when it exits on its own after crucible closes its input, at a restart or at the end of the run, and discarded when it has to be stopped instead. A server that has exited is not stopped while its writes wait for another command's publication. Where they cannot be published, because a root it wrote into changed while it ran, the restart or the disposal fails with that reason. No replacement is started behind that call, because what it wrote is in a state only a fresh start should settle; the next turn starts the server as usual.

Inspection retains backend ID/version/provenance, capability claims, separate hashed requested and effective policies and redacted plans, manifest, working-directory and root identities, root access/provenance, network shape, requested limits, unreadable-pattern counts, command-policy digest, degradation and cleanup state. Parent filesystem carve-outs, command filters, network authority, resource ceilings, session grants and unreadable patterns cannot be dropped or relabelled by a descendant. It does not retain command arguments, environment values, endpoint names, raw approval proofs, credential handles or values, proxy material or raw out-of-scope paths.

The local supervisor attempts to stop the owned process scope and reap its leader on normal exit, deadline, output violation, cancellation, refusal, launch failure, panic, explicit stop and ordinary host shutdown. On enforcing Linux a deadline or output ceiling first tells the broker to end the workload, so the workload's own wait status is still reported, and kills the launcher once the broker has exited or after a short budget. The local process owner used by compatibility mode retries unfinished cleanup within bounded waits. The enforcing Linux projection owner retains its separate terminal failure and quarantine outcome. Successful cleanup is idempotent; a failed attempt does not become successful just because it is called again. Staging data and the command's admission slot remain held until process cleanup is confirmed. If the process owner is dropped while cleanup is still uncertain, staging data is retained and the slot remains consumed for the lifetime of that sandbox service. A recorded supervisor failure stays visible even if later cleanup releases those resources.

These ownership rules also apply when startup fails after a process has been spawned. Startup cleanup uses the same bounded stop and reap path. An uncertain result reports failed cleanup and retains its staging data and command slot; the enforcing Linux backend also retains its separate projection and journal for recovery. A failure to record another audit fact cannot replace the original startup or cleanup error.

An uncatchable host/process kill cannot run user-space destructors. On enforcing Linux, loss of the broker status channel and Bubblewrap's parent-death boundary still terminate the PID-namespace workload. The next preparation replays the checksummed lifecycle WAL, rolls back an unambiguous interrupted transaction, and quarantines ambiguous publication instead of inventing cleanup or success.

Reading the report

crucible sandbox inspect prints that inspection report for the directory you are standing in and stops; crucible --sandbox prints the same report. Nothing is started to produce it: not the backend, not its broker, not a command. No sandbox is prepared, nothing is materialized and nothing is written. The backend reported is the first one in the places a command's preparation looks that passes the same trust checks on its owner, on who can write it and on whether you may run it, and build is the SHA-256 of that file. Preparation also starts each candidate to check it, and passes over one that fails for the next, so where a trusted file will not start, the backend a command gets can be a later one than the report names, or none, and then no command is run.

sandbox enabled in <root>
  mode      required by project configuration
  backend   linux-bubblewrap, system
  version   unverified; reading it would start Bubblewrap, and inspecting starts nothing
  build     sha256:523da3e7399044be5163aee6f57a77a6bef7454376e28f0a0627920bae1b76b6

what this backend can hold:
  filesystem            enforced
  network_deny          enforced
  network_allowlist     enforced
  descriptor_isolation  enforced
  process_isolation     enforced
  kernel_surface        enforced
  privilege_isolation   enforced
  materialization       enforced
  cpu_limit             enforced
  memory_limit          enforced
  disk_limit            unsupported
  process_limit         enforced
  open_file_limit       enforced
  command_time_limit    enforced
  session_time_limit    unsupported
  outbound_byte_limit   unsupported
  output_limit          enforced
  concurrency_limit     enforced
  cost_limit            unsupported
  pty                   unsupported
  file_operations       unsupported
  persistence           unsupported
  snapshot              unsupported
  resume                unsupported
  audit                 enforced
  usage                 observed

what a command would run under:
  enabled   true
  cwd       sha256:de761e35c85ea9a7615fe3e45e1fef07d6631d0f02c4de318c94d05f30813ca2
  reach     2 places, named by digest
    read_write  workspace           sha256:cd490a63f17abf417975914cb065fde6035ea5875cc4e86e6c338842c0cd6262
    protected   protected_metadata  sha256:1c6e26020fe934c29b865b19537f043598e19249d3a277492a790690551c8955
  hidden    0 patterns
  network   closed
  ceilings
    cpu 60m                 enforced
    4096 open files         enforced
    10 MiB captured         enforced
    4 at once               enforced
    20m per command         enforced
  staged    nothing
  outlives  no
  snapshots no
  confined  yes
  policy    sha256:ad76ba3f9182bb58bad76fc0c8167c7d7170c2c662ad7521ac6a90ed837c90ef
  manifest  sha256:72c169df3655f6857b163163b2618631ec65982557eea70b57ea1cb983ae0890

not checked, since checking would start something:
  whether Bubblewrap starts and can make its namespaces on this host, and whether each root passes the checks a command's preparation makes

What only starting the backend could tell is reported as unverified, with the reason, rather than guessed: the version line says so, and the last section lists what was left unchecked. Those are the questions a command's preparation answers when it starts the backend, so a report that says confined yes says that the backend found claims this policy, not that it has been seen to start on this host. In the sample the workspace root is shown as <root>.

Every ceiling is printed beside the claim it rests on, because a number the backend only observes is recorded after the fact rather than imposed, and the two read identically until they are on one line. The capability matrix lists every feature, including the ones this policy never asks for; it is what makes an enforced elsewhere in the report mean something.

Time ceilings retain their exact value, including fractional seconds: a 1.5-second limit is shown as 1.5s, and a limit below one second is never rounded down to 0s.

The workspace root is the only path printed. Roots, the working directory and the policies are named by the digests the record keeps, so a report can be pasted into an issue without pasting your tree with it. Two reports over the same directory produce the same digests, which is what makes them comparable between machines.

Where a backend is found but its matrix will not hold this workspace's policy, the matrix is printed, then what was asked for, then no command could be run here with the reason. An unsupported row in the matrix is usually the whole explanation. Where no backend is found at all, the report says no sandbox backend was found with the reason, rather than printing a matrix of claims belonging to nobody. Neither is an error: both are the answer, and the run ends successfully.

crucible sandbox inspect --json writes the same report as one JSON document on one line, with format_version 1, kind sandbox-inspection, and a status of ready, refused, unavailable or failed. It carries no path at all, not even the workspace root. Its fields and exit statuses are listed under Command line.

Writable roots and publication

On enforcing Linux a writable root is not handed to the command directly. The command writes into a private projection of that root, and the host publishes the changed paths back only after the command has ended. Until then nothing the command wrote is visible outside the sandbox, and a reader of the workspace sees the root exactly as it was when the command started.

That projection is an overlay, and Bubblewrap 0.11 takes no mount options for one, so it is the single mount inside the sandbox that would honour a setuid bit or a device node. Neither is a way out. The namespace maps one unprivileged identity and nothing above it, so a file the command marks setuid can only ever name the identity it already runs as, and a mknod for a host disc is refused because the command holds no capability to make a device with. Every other mount, including the read-only runtime and the whole of /dev apart from the device nodes themselves, is mounted nosuid and nodev. The tests assert the behaviour rather than the flag, because the flag is the one thing here that cannot be set.

Because the projection is an overlay over the root itself, a publication into that root while the command runs changes what the command sees beneath its own writes. The command's own publication is then refused, as described below, so such a change can confuse a running command but cannot reach the root through it.

Publication is decided by how the command ended:

  • An ordinary exit, zero or nonzero, publishes the changed paths. A failing build still leaves the files it wrote, as it would have without confinement.
  • Termination by a signal, a deadline, Esc, an output ceiling or a refusal discards the projection. Nothing partial reaches the workspace. A command that has already ended is not stopped by a deadline or Esc while its writes wait their turn to publish, until the ceiling below passes.
  • A command that wrote into a root which changed after it started, whether another command published into it or something outside the sandbox wrote to it, publishes nothing. The delta is discarded rather than merged, and the result says so: the call's own result for a command that was waited for, the note about its ending for a command left running, and the restart or disposal for a confined MCP server. A root the command left unchanged is not checked, so a command that wrote nothing into a changed root is not refused for it.

Publication itself is transactional. The changed paths are staged in this user's private sandbox state directory under /var/tmp, which no other user can read or enter, journaled in a checksummed write-ahead log, applied, and verified. A projected file larger than 8 GiB is refused rather than digested, so a sparse file cannot make publication read through the whole of it. A failure between those steps rolls the root back to its pinned baseline. Where a rollback cannot itself be proved, the staged content is retained as quarantine evidence and the cleanup outcome reports it, rather than deleting what cannot be accounted for. That evidence records the root as its own transaction last saw it, and another command may have published into the root since, so compare it with the root's current content before restoring anything from it. The next preparation recovers any transaction an earlier process abandoned.

The state directory's shipped name, /var/tmp/crucible-code-sandbox-<uid>-v1 where <uid> is the user's numeric id, is predictable, so another local user can create it first, whether a directory, a symlink or a plain file. crucible refuses to use what it finds there and reports that the sandbox backend is unavailable, naming which of those it is without the path or the other user's numeric id; removing it, which may need an administrator, is what lets the sandbox be used. The directory can instead already be this user's own but carry the wrong group or permissions, for example after running crucible under a different group; crucible refuses that too, but says to restore its group and mode instead of removing it, since the directory holds this user's own unrecovered transaction journals and quarantine evidence.

Commands that can write run side by side, including a command left running in the background and a confined MCP server, and none of them holds up another while it runs. Publication is what has to happen alone. A host-owned lock, one for each user, is held while a command's publication is checked against its baseline, applied and verified, and while a starting command takes its baselines, so no two publications overlap and no baseline is taken halfway through one. A command that ends cleanly while another is publishing waits for that publication to finish before its own begins; one killed, or stopped by a limit, has nothing to publish and does not wait. Waiting commands are not served in order, and each keeps its place among the commands allowed to run at once until its own publication finishes. Every wait for the lock has a ceiling: a command being prepared is refused if the lock does not come free within a minute, a command whose own deadline has passed is stopped after a minute, and five seconds is the bound everywhere somebody is waiting: a cancelled turn, a command left running whose report is overdue, a confined server being restarted or disposed of at a turn's end, and the run itself ending.

A command that ran while another published into one of its roots publishes nothing, whatever the root looks like afterwards: this user's state directory remembers how many publications have touched each root, and a command compares that count with the one it recorded when it took its baseline. A crucible from before this release takes the same lock, so the two still publish one at a time, but it does not keep that count. Against such a peer the comparison has nothing to see, and only the check against the root's own content remains. The holder may be another crucible of this user, including one from before this release, which keeps the lock for as long as its commands run; a wait with no end would hold up the turn, the cancel and the exit instead.

A detached command follows the same rules when it ends later. Its start result is accepted only after it is durably stored, and its terminal publication is journaled under the same call identity, so a host restart between the two neither loses the command nor publishes it twice.