Skip to main content

Experiments

Small interactive tools I've built to answer specific questions, mostly about what AI models cost and how coding agents spend their context.

Agentic Coding Patterns

Browse a library of named patterns for working with coding agents, each with the part most catalogues leave out: when not to use it. Or open your own folder of notes.

Cache Break-Even

Work out whether switching models or effort levels mid-session pays back the cost of re-caching everything already in context, from your own session or from numbers you type.

Compact or Clear

Project what keeping going, compacting, and clearing cost from where your coding session is now, and see after how many turns each choice pays for itself.

Delegation Economics

See how much faster a task finishes when you split it across subagents or an agent team, and how many more tokens and dollars it costs than one session.

Lethal Trifecta

Turn a coding agent’s capabilities and controls on and off, and see whether untrusted content can still steer it into sending private data somewhere you don’t control, and which of your controls actually prevent that.

Loop Governor

Simulate thousands of runs of an agent loop to see how often it finishes honestly, stops on a false “done”, or runs away with your budget, and which governors change that.

Measurement Noise

Bring timings from two ways of working and find out whether the data can tell them apart, with the uncertainty drawn out, a sample-size planner, and a warning when you’re measuring activity instead of outcomes.

Mechanism Picker

Decide whether something belongs in a prompt, an instruction, a skill, a subagent, a hook, a workflow, a goal, a loop, a routine, or CI, and lint a CLAUDE.md for rules that need a stronger rung.

Model Pricing Calculator

Compare what the same token usage costs across AI models, or drop in a Claude Code or Codex session to price its real usage on every model.

Review Capacity

Compare how much code your parallel agents open each day with how much one person can review well, and see the backlog or the escaped defects that pile up over two working weeks.

Session Log Auditor

Drop in your Claude Code session transcripts and count what actually goes wrong: recurring tool failures clustered by session, how many come from your environment rather than the model, what the sessions cost, and whether the problems you fixed stay fixed.

Usable Evidence Budget

See how much of a model’s context window is left for the files you are working on after instructions, history, tools, reserved output, and compaction margin each take their share, then drop in files to check whether they fit.

Which Model Actually Runs

Work out which model a Claude Code subagent really runs on, on any version, and audit your own agent files to see what changes when you upgrade.

Workflow Fan-Out

Animate a multi-stage workflow under pipeline() and parallel(), see what .filter(Boolean) hides in the results, estimate what the run costs by model, and lint a pasted orchestration script.