Skip to main content

Cut a leg of the lethal trifecta

Can untrusted content steer your agent into moving private data somewhere you don’t control? A coding agent reads untrusted content, reaches private data, and has a way out, all by default. You don’t choose whether that’s true. You choose which leg to cut, and only a structural control cuts one.

Turn capabilities and controls on and off, and watch whether a complete path survives. This page has no server and sends nothing anywhere.

Predict first

This is the Careful team: a firm line in CLAUDE.md telling the agent to ignore instructions in untrusted content, auto mode on, and a deny rule for Bash(curl *).

Is this agent still exploitable?

Presets

A firm line in CLAUDE.md, auto mode, and a curl deny. It looks careful, and none of it is a boundary.

Verdict

Make your prediction above to see the verdict.

The diagram

Tick a node to give the agent that capability. A cut edge has a dashed border and names the control that cut it. Edges on the highlighted path are red.

Untrusted content (leg intact)

  • Live

  • Live

  • Live

  • Live

  • Off

  • Live

Private data (leg intact)

  • Live

  • Live

  • Live

  • Off

  • Live

Agent context

Untrusted content and private data meet here, and actions leave.

Ways out (leg intact)

  • Live

    curl is denied. wget, python, node, and the rest are not.

  • Live

  • Live

  • Live

  • Live

  • Off

  • Off

Every live path, in words

Make your prediction above to see the live paths.

Controls

Hover or focus a control to highlight the edges it removes.

Structural

Properties of the OS, the network, or a permission rule. These are the controls that cut.

Architectural

How the work is split, so the payload never reaches the part that acts.

Human gate

Holds while a person reads every prompt. Listed as a residual risk, not a cut.

Prompt-only and best-effort

They rely on the model’s judgment, so they cut nothing.

Loading the allowlist…

Share or save

The summary is Markdown: the verdict, the path, the controls in place by kind, and the residual risks. The link holds the toggles, never your settings files.

Loading the share buttons…

Load your settings

Upload or paste your user, project, and local settings files, and the controls fill in from what’s there, each with the line that justified it. Anything a settings file can’t settle stays as your toggle.

Loading the settings reader…

Agents in CI

A CI agent is a credential-bearing process whose prompt is partly written by anyone who can open an issue. Each choice below adds or cuts a leg.

Loading the CI scenarios…

Why a reader/doer split works

Only the fields of a strict schema pass from the reader to the doer. Edit the text and play it through.

Loading the reader/doer demo…

Vectors you might not have considered

Show me turns the matching nodes on in the diagram.

Loading the vectors…

Notes and sources

The trifecta

The lethal trifecta is Simon Willison’s: an agent is exploitable when it reads untrusted content, reaches private data, and has a way to send data out, all at once. Every control that works here is a property of the OS, the network, or the architecture. Every one that fails is a property of the model’s judgment.

Auto mode

Anthropic called auto mode “a convenience feature backed by a best-effort classifier, not a security guarantee,” in response to Johann Rehberger’s published attack chain, which succeeded 60–80% of the time. It cuts down on prompt fatigue. It doesn’t stop a determined payload.

Secrets

The only safe credential is one the agent can’t read. If something leaks, rotate first and investigate second.

From the documentation

These come from the Claude Code sandboxing documentation, checked on 2026-10-04:

  • The sandbox constrains Bash commands and the processes they start.
  • Read, Edit, Write, and WebFetch follow permission rules instead. allowedDomains doesn’t limit WebFetch, and denyRead doesn’t stop the Read tool.
  • Commands in excludedCommands run outside the sandbox.
  • An allow rule like Bash(curl *) also approves an unsandboxed retry, unless allowUnsandboxedCommands is false.
  • strictAllowlist only counts in user, managed, or --settings settings. A repository’s files can’t set it.

What this doesn’t model

  • Every path has the same length, so the highlighted one is the first in the order the columns list them, which puts the most common vectors first.
  • It doesn’t scan repositories or test payloads against a live agent. Exercise 6 does that, safely.
  • A settings file can’t say whether you run in a container, split reading from doing, plan first, or set core.fsmonitor, so those stay your toggles.