Sandbox
Claude Code and Codex sandboxes, Docker
What it covers
Limits on where any code can reach: files, network, system calls, including code inside an interpreter.
What it leaves open
Whether this particular command should run at all.
Open-source prototype · Linux 5.19+
WardenClaw stops the risky commands of an AI agent on your Linux server right in the kernel and runs them only after you sign them on your phone. The key never lives on the server, so even a hijacked agent can't approve itself.
Take back control.
Approve on Android, iPhone or in a terminal. Any agent harness; the OpenClaw plugin is optional.
The code opens with the first release. Until then the docs describe the prototype as it is.
Where it fits
Claude Code and Codex sandboxes, Docker
What it covers
Limits on where any code can reach: files, network, system calls, including code inside an interpreter.
What it leaves open
Whether this particular command should run at all.
Claude Code, Codex, OpenClaw
What it covers
A question before each tool call the agent reports, answered with a button, a chat message or an allow-list.
What it leaves open
The answer is not signed and is checked on the same host as the agent. A child process of an approved script never shows up.
Under any agent harness on Linux
What it covers
Every program start in the agent's tree, checked in the kernel. The risky ones wait for a signature made off the server.
What it leaves open
What a running process does inside itself: interpreters, file writes, network. That stays with the sandbox.
Use them together: the sandbox limits how far anything can reach, WardenClaw asks you about the risky steps inside those limits.
Scope
execve and execveat stops in the kernel and is checked against the rules, including starts from scripts and child processes. Starts through memfd_create, fexecve, a direct call of ld.so or a 32-bit call are gated or refused too.redteam/hardened-check.sh checks this as the agent user.busybox cat, busybox rm), a file reader outside the list (hexdump) or a system program copied under another name runs as an ordinary start: journaled, no card. --policy-mode root asks about every new command instead, at the cost of many more cards.python -c, node -e), shell redirections, LD_PRELOAD, file writes and network inside a running process are not. Work handed to a daemon or another host (docker, ssh, systemd-run) is out of sight once approved, and work started outside wardend's tree (another login, cron) is not seen at all. Keep a sandbox around the agent.wardend verify-journal catches a record removed from the middle, not the last records removed from the end.The problem
An agent on a server keeps working while you sleep. Its permission prompt waits at a screen you are not looking at, so people widen the allow-list or switch on an auto mode, and whatever passes runs with the agent's rights.
An approval is usually just an RPC call. Any client with the right scope, or a /approve typed in chat, resolves it. There is no signature, so the gate can't tell your "yes" from one produced inside the loop.
Server-side auto-reviewers run on the same host, often on the same model, as the agent they review. If the host is compromised, it approves itself.
A plugin hook guards only the calls the agent framework chooses to report. A child process started by an approved script never shows up there.
How it works
Two gates share one envelope format. wardend checks every program start in the agent's process tree at the syscall level. The optional OpenClaw plugin gates OpenClaw's own tool calls. Both send cards to the phone and accept nothing but a valid signature.
CONTINUE or EPERMwardend runs the agent under a seccomp user-notification filter. Every execve in the tree stops in the kernel while the supervisor reads argv, cwd, the real binary and the parent chain from the stopped process and checks them against the rules. What trips no rule continues at once and is journaled. The plugin gates OpenClaw's tool calls (exec, Bash, Write, Edit) through before_tool_call.
A command that trips a rule becomes a canonical envelope and a sha256 digest, byte-identical in Go, JavaScript and TypeScript (shared test vectors). wardend serves it to the phone from its own HTTP endpoint behind a tunnel you choose. Cards and decisions are signed, so the tunnel can't forge them.
The app recomputes the digest itself, shows what the command does, and signs allow or deny with the device's Ed25519 key. If the digest doesn't match the envelope, the only thing it will sign is deny.
wardend checks the trusted device, a ±60 s time window, a single-use nonce and the digest, then answers the kernel with CONTINUE or EPERM. What the approved command starts is checked on its own; only the same tool family passes for 10 minutes (ssh under scp). Every exec and decision lands in a hash-chained, signed journal.
Who answers on the phone
Every card waits for you. One line in plain words, a risk color, two big buttons: Allow once or Deny. The exact command, argv and process chain are one tap away.
A judge proposes a verdict with a reason and a 0 to 100 risk score, and you still press the button. The journal records whether it matched you. Watch it for a while before you trust it.
The judge allows below your risk threshold and denies what it finds dangerous; everything else stays with you. A deterministic block-list (rm -rf on roots, dd, mkfs, curl | sh, gateway config, secrets) goes to you before any model sees it. Text aimed at the reviewer inside a command means deny. Three denials in a row fall back to Observe.
What needs a signature
server · policy_mode
Whether the answer is enforced
server · mode
deny_always rules block Who answers on the phone
app · Mode tab
The server's observe and the phone's Observe are different things: the first enforces nothing, the second only means the judge suggests.
The judge runs on the phone's side: the app asks a model at a URL you set; the agent host never sees its prompt, its keys or its journal. Stop in the Mode tab switches back to Manual at once, and autopilot turns itself off after an hour.
Security, honestly
EPERM.ENOSYS until the service restarts.EAGAIN, never an automatic allow.The kernel copies execve arguments after the supervisor answers, so another thread can swap argv in between. We ran a hostile binary that flips argv[1] between SAFE-ARG and EVIL-ARG in a loop, 40 runs per mode, with the device approving only SAFE-ARG.
| Mode | Approved | Ran with swapped argv | Killed by post-check |
|---|---|---|---|
off | 5 | 4 | n/a |
poll | 7 | 3 | 7 |
stop | 8 | 0 | 6 |
stop (the default for commands that need a signature) sends SIGSTOP to the caller before CONTINUE. The new image halts before its first instruction, /proc exe and cmdline are compared with the signed envelope, and a mismatch is killed. A swapped program never executes a single instruction. Cost: about 1 to 1.5 ms per approved root.
Only if it runs as the same user. A malicious model that controls the harness shares that user's rights, so in a same-user install it could signal, ptrace or overwrite wardend without any exec. The strong guarantee needs a hardened install: wardend as a root system service, root-owned binary and config, the harness as an unprivileged child under its own user. That is the install we recommend. Every vector it closes is covered by a test.
curl -fsSL https://wardenclaw.dev/install.sh | sudo sh One command, hardened by default, starts in observe mode. The release signature is checked before anything is written.
Recommended: hardened install Why the model can't turn it off Verify releases
Numbers
per exec under the supervisor (1000 × /bin/true: 0.51 s → 1.06 s)
cards a day, root mode → tripwire, replaying a 25-hour journal of a busy agent (98,053 execs; peak 34 cards in 2 minutes)
added to a whole claude-cli turn
median kernel receive-to-reply inside the supervisor (p95 510 µs)
Compared
| WardenClaw | agentsh | grith | Claude Code Remote Control | OpenClaw approvals | |
|---|---|---|---|---|---|
| Where the gate sits | Kernel (seccomp on execve) + tool hook | Kernel (seccomp user notify, Landlock, FUSE, eBPF) | ptrace + seccomp | Agent's permission prompt | Gateway tool approvals |
| Where you approve | Phone app or wardenctl in a terminal | Local TTY, TOTP, WebAuthn, API | Telegram, Slack, Discord, WhatsApp buttons | Phone or web | Mobile app, Control UI, /approve in chat |
| Decision is a signature the gate verifies | Yes: Ed25519, key off the host | Local WebAuthn or HMAC on the same machine | Nonce + HMAC webhook | No | No |
| Automatic judge | On the phone side, outside the host | Policy rules | LLM scoring on the host | No | Server-side reviewer on the gateway host |
| Audit trail | Hash chain + Ed25519 signatures, verifiable by a third party | HMAC-chained audit | Not assessed | Not assessed | Approval history (30 days) |
| Status | Prototype | v0.20.5, Apache-2.0 | v0.3.4, MPL-2.0 + paid tiers | Generally available | Built into OpenClaw |
Based on public docs and repositories as of September 2026. agentsh is the closest in mechanics and worth a look. The difference is where the decision is made and who holds the key.
Components
Supervisor on seccomp user notification. No cgo, no libseccomp, Linux 5.19+; in the recommended install a root system service with the agent as its own user. Tripwire rules with harness packs (claude-cli, OpenClaw) or the older root mode, TOCTOU stop-verify, its own HTTP endpoint and pairing for the phone, a signed journal and wardend replay.
The code opens with the first release
Optional. Gates OpenClaw's own tool calls through before_tool_call in observe or enforce mode, verifies signed decisions and keeps its own signed journal. It can also relay wardend exec cards through the gateway.
The code opens with the first release
Pairs with wardend by QR code as a device with its own Ed25519 key in the phone's secure storage. Feed, Mode, Journal and Connection tabs; a risk judge on a model at a URL you set; an append-only journal with a hash chain. Android and iPhone.
The code opens with the first release
Status
FAQ
Keep it. A sandbox limits where any code can reach, including code inside interpreters that tripwire doesn't see. It doesn't decide whether this particular command should run, and its own prompt is answered on the same host. WardenClaw asks about the risky steps inside those limits, with a signature made off the host. The two stack; isolation around the gate (Landlock) is on the roadmap.
The agent stops; it doesn't run free. Without a signature nothing new executes. On the host, run wardend pair revoke for that device (or remove it from trusted_devices, and from trustedDeviceIds if you use the plugin) and pair a new phone. The signing key lives in the phone's secure storage, never on the agent's host.
The recommended install runs wardend as a root system service and the agent as its own unprivileged user, so the model can't turn the guard off. A single-user trial needs no root, but there the model could stop wardend. The gate itself is no_new_privs plus a seccomp filter with a notification listener: no cgo, no libseccomp, Linux 5.19 or newer.
Not yet. The gate is built on Linux seccomp user notification. A macOS backend is planned, with no date.
By default only commands that trip a rule wait for your signature, such as remote execution, package installs, publishing, reading secrets, destructive commands or changes to the agent's settings. Everything else runs and is journaled. How many cards that makes depends on the agent: replaying a busy agent's 25-hour journal gave 819 a day (3,162 in the older root mode). Run your own journal through wardend replay before you enforce, and let Autopilot take the routine.
No. wardend supervises any Linux process and ships harness packs for claude-cli and OpenClaw. wardend serves the phone app itself over its own HTTP endpoint. The OpenClaw plugin is optional: it adds cards for OpenClaw's own tool calls.