Skip to main content

Breaking the cache: model and effort calculator

Switching models, or changing the effort level, in the middle of a session makes the provider reprocess everything already in your context. That reprocess is billed as a cache write on the destination. Set where you’re coming from, where you’re going, how much context you’ve built, and how much output work is left, and the chart updates live.

Start from your session

Fill in your context, your current model, and an estimate of the work left from a Claude Code session, or from a context readout you paste. Everything is read in your browser, and nothing you drop or paste is sent anywhere.

Use my session

Files are read in your browser. Nothing is uploaded.

  • Claude Code saves sessions as ~/.claude/projects/<project>/<session>.jsonl. That folder is hidden. In the macOS file picker, press ⌘⇧. to show it.
  • Choose several files, or a folder, and the newest turn wins.

Paste a line such as 312k/1000k tokens, or the output of claude -p "/context". The first pair is offered as N.

Your setup

Where you’re coming from
Where you’re going
Cache TTL

no effort change

R is measured at your current settings.

Cost to make the change

$2.00

re-cache 500K at Sonnet 5 input rate

Value of remaining work

$2.25

$15.00 per MTok of output

Net

+$0.25

ahead if you change now

Switching to Sonnet 5 already pays for itself, by $0.25. That stays true as long as your context doesn’t grow past 563K tokens.

Where you stand

Loading the chart…

Every option from here

Loading the table…

Sensitivity map

Net for every combination of context (N) and remaining work (R), with the line where the change breaks even. Choose a cell to set both.

Loading the map…

Prices and assumptions

Loading the price editor…

How this works

The two levers

Changing the model and changing the effort level both reprocess everything already in context. Either one costs N × destination input price × write multiplier. Right now that is 500,000 × $2 × 2 ÷ 1,000,000 = $2.00. The multiplier is 2× on a 1-hour TTL and 1.25× on a 5-minute one, because the coding tool caches automatically. You’re on the 1-hour TTL.

Where they differ

A model switch changes the price of each output token. An effort change changes how many tokens get generated. The value of what’s left is R × (from output − ratio × to output), which is 150,000 × ($25 − 1.00 × $10) ÷ 1,000,000 = $2.25. That’s why an effort-only change on an expensive model can cost more to carry out than switching to a cheaper one: the re-cache is priced wherever you land.

Effort ratios

Each factor is output volume relative to high:

  • low: 0.25
  • medium: 0.50
  • high: 1.00
  • xhigh: 1.50*
  • max: 2.20*

* Placeholder. No published runs back this number, so override the ratio when you know better.

In the published long-horizon coding runs behind the sourced factors, medium effort costs roughly half as much per task and gives up about 2 points of pass rate. Low costs about a quarter and gives up about 8 points. Effort curves depend on the workload: research and knowledge work is much flatter than coding.

Simplifications

  • R counts output tokens only.
  • Fresh input you add later costs less on a cheaper model, and future turns read your context at the destination’s cache-read rate. Both make real break-evens more favorable than the ones shown here.
  • Some models also invalidate the tools and system caches on an effort change. This page treats everything as a full re-read.