Open-source prototype · Linux 5.19+

Two-factor approval for AI agents.

WardenClaw stops the risky commands of an AI agent on your Linux server right in the kernel and runs them only after you sign them on your phone. The key never lives on the server, so even a hijacked agent can't approve itself.

  • Below the agent. The kernel sees every program the agent starts, including the ones launched by scripts and child processes.
  • A signature, not a button. Ed25519 over the exact command, made on your phone. The agent can't produce one.
  • Quiet until it matters. Routine work runs and is journaled; your phone gets what trips a rule. Keep your sandbox: this works on top of it.

Take back control.

Approve on Android, iPhone or in a terminal. Any agent harness; the OpenClaw plugin is optional.

The code opens with the first release. Until then the docs describe the prototype as it is.

Where it fits

Sandbox, built-in approvals, WardenClaw: three different jobs

Sandbox

Claude Code and Codex sandboxes, Docker

What it covers

Limits on where any code can reach: files, network, system calls, including code inside an interpreter.

What it leaves open

Whether this particular command should run at all.

Built-in approvals

Claude Code, Codex, OpenClaw

What it covers

A question before each tool call the agent reports, answered with a button, a chat message or an allow-list.

What it leaves open

The answer is not signed and is checked on the same host as the agent. A child process of an approved script never shows up.

WardenClaw

Under any agent harness on Linux

What it covers

Every program start in the agent's tree, checked in the kernel. The risky ones wait for a signature made off the server.

What it leaves open

What a running process does inside itself: interpreters, file writes, network. That stays with the sandbox.

Use them together: the sandbox limits how far anything can reach, WardenClaw asks you about the risky steps inside those limits.

Scope

What it closes, and what it doesn't yet

Closes

  • Every program start in the agent's tree. Each execve and execveat stops in the kernel and is checked against the rules, including starts from scripts and child processes. Starts through memfd_create, fexecve, a direct call of ld.so or a 32-bit call are gated or refused too.
  • Forged and replayed approvals. A decision is an Ed25519 signature over the exact command with a single-use nonce and a ±60 s window, made with a key that never lives on the agent's host. The agent can't approve its own card through wardend's HTTP endpoint.
  • A command swapped after approval. For a signed command wardend stops the new program before its first instruction and compares it with the signed envelope; a mismatch is killed.
  • The model turning the guard off, in the hardened install. wardend runs as a root system service and the agent as its own user, so the kernel refuses the agent's signals, ptrace and memory writes against wardend, and its binary, config, key and unit belong to root. redteam/hardened-check.sh checks this as the agent user.

Not closed yet

  • Tripwire knows programs by name and by list. A multi-call binary (busybox cat, busybox rm), a file reader outside the list (hexdump) or a system program copied under another name runs as an ordinary start: journaled, no card. --policy-mode root asks about every new command instead, at the cost of many more cards.
  • Only exec is gated. Interpreters (python -c, node -e), shell redirections, LD_PRELOAD, file writes and network inside a running process are not. Work handed to a daemon or another host (docker, ssh, systemd-run) is out of sight once approved, and work started outside wardend's tree (another login, cron) is not seen at all. Keep a sandbox around the agent.
  • A single-user trial can be turned off. Without the hardened install wardend and the agent share a user, so the model can stop or rewrite wardend without a single exec.
  • No post-check without a card. Starts that need no card are not re-checked after the kernel continues them. The TOCTOU stop-and-compare covers only commands that were signed.
  • Same-inode swaps. Changing a file's content in place between check and exec is not caught; the binary hash is not part of the envelope yet.
  • Only listed variables are shown. The card shows the environment variables from a list of those that change what a program does: the loader, shells, interpreters, git, proxies. The rest of the environment, including anything that looks like a secret, is signed only as a hash, so a variable outside the list can change an approved command without your seeing it.
  • A cut journal tail passes the check. wardend verify-journal catches a record removed from the middle, not the last records removed from the end.

Threat model The vector table and its tests

The problem

The approval lives on the host it guards, and the agent can reach it

Nobody at the screen

An agent on a server keeps working while you sleep. Its permission prompt waits at a screen you are not looking at, so people widen the allow-list or switch on an auto mode, and whatever passes runs with the agent's rights.

A button the agent's side trusts

An approval is usually just an RPC call. Any client with the right scope, or a /approve typed in chat, resolves it. There is no signature, so the gate can't tell your "yes" from one produced inside the loop.

A reviewer inside the blast radius

Server-side auto-reviewers run on the same host, often on the same model, as the agent they review. If the host is compromised, it approves itself.

Hooks only see what they are shown

A plugin hook guards only the calls the agent framework chooses to report. A child process started by an approved script never shows up there.

How it works

One signed decision per risky command, enforced below the agent

Two gates share one envelope format. wardend checks every program start in the agent's process tree at the syscall level. The optional OpenClaw plugin gates OpenClaw's own tool calls. Both send cards to the phone and accept nothing but a valid signature.

  1. Agentstarts a program, itself or through a script
  2. Kernel and wardendstop the start and check the rules; if none trips, it runs and is journaled
  3. Phoneshows the card of a start that tripped a rule
  4. Yousign allow or deny with the phone's key
  5. wardendchecks the signature and answers the kernel: CONTINUE or EPERM
  1. Freeze

    wardend runs the agent under a seccomp user-notification filter. Every execve in the tree stops in the kernel while the supervisor reads argv, cwd, the real binary and the parent chain from the stopped process and checks them against the rules. What trips no rule continues at once and is journaled. The plugin gates OpenClaw's tool calls (exec, Bash, Write, Edit) through before_tool_call.

  2. Describe

    A command that trips a rule becomes a canonical envelope and a sha256 digest, byte-identical in Go, JavaScript and TypeScript (shared test vectors). wardend serves it to the phone from its own HTTP endpoint behind a tunnel you choose. Cards and decisions are signed, so the tunnel can't forge them.

  3. Sign

    The app recomputes the digest itself, shows what the command does, and signs allow or deny with the device's Ed25519 key. If the digest doesn't match the envelope, the only thing it will sign is deny.

  4. Verify and run

    wardend checks the trusted device, a ±60 s time window, a single-use nonce and the digest, then answers the kernel with CONTINUE or EPERM. What the approved command starts is checked on its own; only the same tool family passes for 10 minutes (ssh under scp). Every exec and decision lands in a hash-chained, signed journal.

Detailed diagrams

Who answers on the phone

Start by pressing buttons. Hand over the routine when the numbers say so.

01 You decide

Manual

Every card waits for you. One line in plain words, a risk color, two big buttons: Allow once or Deny. The exact command, argv and process chain are one tap away.

02 Judge suggests

Observe

A judge proposes a verdict with a reason and a 0 to 100 risk score, and you still press the button. The journal records whether it matched you. Watch it for a while before you trust it.

03 Judge decides, within rules

Autopilot

The judge allows below your risk threshold and denies what it finds dangerous; everything else stays with you. A deterministic block-list (rm -rf on roots, dd, mkfs, curl | sh, gateway config, secrets) goes to you before any model sees it. Text aimed at the reviewer inside a command means deny. Three denials in a row fall back to Observe.

Three switches, set separately

What needs a signature

server · policy_mode

  • tripwire default only a start that trips a rule
  • root every start outside an approved tree and every delegating spawn

Whether the answer is enforced

server · mode

  • observe default nothing waits; the journal records what would need a card
  • deny-list only the deny_always rules block
  • ticket what needs a signature waits for one

Who answers on the phone

app · Mode tab

  • Manual default you
  • Observe the judge suggests, you press
  • Autopilot the judge decides within your rules

The server's observe and the phone's Observe are different things: the first enforces nothing, the second only means the judge suggests.

The judge runs on the phone's side: the app asks a model at a URL you set; the agent host never sees its prompt, its keys or its journal. Stop in the Mode tab switches back to Manual at once, and autopilot turns itself off after an hour.

Security, honestly

What it guarantees and what we measured

Fail-closed by default

  • No signed allow within the TTL (120 s): the exec gets EPERM.
  • The supervisor dies: every exec in the tree gets ENOSYS until the service restarts.
  • The queue is full (64 pending): EAGAIN, never an automatic allow.
  • Plugin hook times out, is aborted or throws: the tool call is blocked.
  • Forged, replayed, stale or foreign-key tickets are rejected (covered by tests).
  • A process in the tree can't install its own notification filter to answer "yes" around wardend.

TOCTOU: the race the kernel leaves open

The kernel copies execve arguments after the supervisor answers, so another thread can swap argv in between. We ran a hostile binary that flips argv[1] between SAFE-ARG and EVIL-ARG in a loop, 40 runs per mode, with the device approving only SAFE-ARG.

Adversarial argv swap, 40 runs per mode
ModeApprovedRan with swapped argvKilled by post-check
off 5 4 n/a
poll 7 3 7
stop 8 0 6

stop (the default for commands that need a signature) sends SIGSTOP to the caller before CONTINUE. The new image halts before its first instruction, /proc exe and cmdline are compared with the signed envelope, and a mismatch is killed. A swapped program never executes a single instruction. Cost: about 1 to 1.5 ms per approved root.

Can the model turn wardend off?

Only if it runs as the same user. A malicious model that controls the harness shares that user's rights, so in a same-user install it could signal, ptrace or overwrite wardend without any exec. The strong guarantee needs a hardened install: wardend as a root system service, root-owned binary and config, the harness as an unprivileged child under its own user. That is the install we recommend. Every vector it closes is covered by a test.

curl -fsSL https://wardenclaw.dev/install.sh | sudo sh

One command, hardened by default, starts in observe mode. The release signature is checked before anything is written.

Recommended: hardened install Why the model can't turn it off Verify releases

Numbers

Measured on a Raspberry Pi 5

+0.35–0.55 ms

per exec under the supervisor (1000 × /bin/true: 0.51 s → 1.06 s)

3,162 → 819

cards a day, root mode → tripwire, replaying a 25-hour journal of a busy agent (98,053 execs; peak 34 cards in 2 minutes)

~20 ms

added to a whole claude-cli turn

310 µs

median kernel receive-to-reply inside the supervisor (p95 510 µs)

Compared

Close neighbours, and where WardenClaw differs

WardenClaw compared with related tools
WardenClawagentshgrithClaude Code Remote ControlOpenClaw approvals
Where the gate sits Kernel (seccomp on execve) + tool hookKernel (seccomp user notify, Landlock, FUSE, eBPF)ptrace + seccompAgent's permission promptGateway tool approvals
Where you approve Phone app or wardenctl in a terminalLocal TTY, TOTP, WebAuthn, APITelegram, Slack, Discord, WhatsApp buttonsPhone or webMobile app, Control UI, /approve in chat
Decision is a signature the gate verifies Yes: Ed25519, key off the hostLocal WebAuthn or HMAC on the same machineNonce + HMAC webhookNoNo
Automatic judge On the phone side, outside the hostPolicy rulesLLM scoring on the hostNoServer-side reviewer on the gateway host
Audit trail Hash chain + Ed25519 signatures, verifiable by a third partyHMAC-chained auditNot assessedNot assessedApproval history (30 days)
Status Prototypev0.20.5, Apache-2.0v0.3.4, MPL-2.0 + paid tiersGenerally availableBuilt into OpenClaw

Based on public docs and repositories as of September 2026. agentsh is the closest in mechanics and worth a look. The difference is where the decision is made and who holds the key.

Components

Three pieces, one envelope

wardend Go

Supervisor on seccomp user notification. No cgo, no libseccomp, Linux 5.19+; in the recommended install a root system service with the agent as its own user. Tripwire rules with harness packs (claude-cli, OpenClaw) or the older root mode, TOCTOU stop-verify, its own HTTP endpoint and pairing for the phone, a signed journal and wardend replay.

The code opens with the first release

wardenclaw-gate OpenClaw plugin

Optional. Gates OpenClaw's own tool calls through before_tool_call in observe or enforce mode, verifies signed decisions and keeps its own signed journal. It can also relay wardend exec cards through the gateway.

The code opens with the first release

WardenClaw app Expo · TypeScript

Pairs with wardend by QR code as a device with its own Ed25519 key in the phone's secure storage. Feed, Mode, Journal and Connection tabs; a risk judge on a model at a URL you set; an append-only journal with a hash chain. Android and iPhone.

The code opens with the first release

Status

Prototype, running on a Raspberry Pi 5 and a Pixel 9 Pro

Works today

  • Kernel gate for the whole tree: tripwire rules decide which starts need a signature
  • Signed tickets end to end: phone → wardend → execve, with signed responses
  • TOCTOU stop-verify, fail-closed paths, hash-chained journals
  • Phone app with Manual, Observe and Autopilot; wardenctl as an approver in a terminal

FAQ

Questions people ask first

Why not just use a sandbox?

Keep it. A sandbox limits where any code can reach, including code inside interpreters that tripwire doesn't see. It doesn't decide whether this particular command should run, and its own prompt is answered on the same host. WardenClaw asks about the risky steps inside those limits, with a signature made off the host. The two stack; isolation around the gate (Landlock) is on the roadmap.

What if I lose my phone?

The agent stops; it doesn't run free. Without a signature nothing new executes. On the host, run wardend pair revoke for that device (or remove it from trusted_devices, and from trustedDeviceIds if you use the plugin) and pair a new phone. The signing key lives in the phone's secure storage, never on the agent's host.

Do I need root?

The recommended install runs wardend as a root system service and the agent as its own unprivileged user, so the model can't turn the guard off. A single-user trial needs no root, but there the model could stop wardend. The gate itself is no_new_privs plus a seccomp filter with a notification listener: no cgo, no libseccomp, Linux 5.19 or newer.

Does it work on macOS?

Not yet. The gate is built on Linux seccomp user notification. A macOS backend is planned, with no date.

Won't I drown in prompts?

By default only commands that trip a rule wait for your signature, such as remote execution, package installs, publishing, reading secrets, destructive commands or changes to the agent's settings. Everything else runs and is journaled. How many cards that makes depends on the agent: replaying a busy agent's 25-hour journal gave 819 a day (3,162 in the older root mode). Run your own journal through wardend replay before you enforce, and let Autopilot take the routine.

Is it only for OpenClaw?

No. wardend supervises any Linux process and ships harness packs for claude-cli and OpenClaw. wardend serves the phone app itself over its own HTTP endpoint. The OpenClaw plugin is optional: it adds cards for OpenClaw's own tool calls.