Skip to main content

What’s actually left for evidence

A context window is a capacity limit. It is not a promise that every token gets used equally well. Before any of it holds the thing you’re working on, five other claims are already against it. This tool is that subtraction, made visible.

usable evidence budget = context capacity
                       − trusted instructions
                       − task and retained history
                       − exposed tool definitions
                       − reserved generation
                       − operational margin

This page has no server. What you paste or drop is read in your browser and never sent anywhere.

Usable evidence budget

810K

81.0% of the window, after 190K is claimed. Largest single draw: operational margin.

Start here

About the best case on a 1M model. The largest draw is the margin, which is a choice rather than a cost.

Where the window goes

Each draw floats down from what is left after the one before it. Hover or focus a bar to read it, and select a draw to jump to its control.

0250K500K750K1M1M−18K−40K0−32K−100K810KContextcapacityTrustedinstructionsTask andretainedhistoryExposed tooldefinitionsReservedgenerationOperationalmarginUsableevidencebudget
  • Capacity
  • Draws against it
  • Usable
  • Usable when nothing is left

Pin the scenario you have as A, change something, and the chart ghosts A behind the new bars while the table shows the difference.

Every figure in the chart with its share of the window.
TermTokensShare of the window
Context capacity1M100.0%
Trusted instructions18K1.8%
Task and retained history40K4.0%
Exposed tool definitions00.0%
Reserved generation32K3.2%
Operational margin100K10.0%
Usable evidence budget810K81.0%

Adjust the terms

Sliders step by 1K, and the boxes accept shorthand such as 25k, 1.2m, or 25,000.

The whole window, before anything claims part of it.

18K

System prompt, project instruction files, memory, and the skill listing.

40K

Prompts, replies, file reads, and tool results.

0

Tool schemas loaded into the prefix.

32K

Room held back for the reply itself.

100K

Slack before automatic compaction fires.

Or set it from the autocompact threshold

Autocompact at a number of tokens or a share of the window. The margin is whatever is left after it.

Will my evidence fit?

Drop the material you want the model to work with and see how much of the usable budget it takes. Token counts are estimates from each file’s length.

Loading the file check…

Getting your own numbers

Run /context in Claude Code to see where your window is going. This table maps each term to what that readout calls it.

What each term is called in the coding tool’s context readout, and how to move it.
TermWhat /context calls itHow to move it
Context capacityThe total in the header, such as the 1M in 120k/1M tokens.Choose a model with a different window.
Trusted instructionsSystem prompt, Memory files, Skills, and Custom agents. That covers the system prompt, environment info, user- and project-level instruction files, auto memory, and skill descriptions.Trim instruction files and memory. Skills flagged to stay out of model invocation also stay out of the listing.
Task and retained historyMessages: prompts, replies, file reads, and tool results.This is the term that grows on its own. Compacting replaces it with a summary, and subagents keep large reads out of it.
Exposed tool definitionsSystem tools and MCP tools. A row marked deferred means tool search is active, and the page shows it without counting it.This is the most reducible term. Turn tool search on, or load fewer MCP servers.
Reserved generationNot a readout row. It is the room the reply needs.It is set by what you ask for: a longer answer needs more room.
Operational marginNot a row in every version. Where the readout has an Autocompact buffer row, the paste box uses it. Otherwise it is the slack before automatic compaction.Set it through the autocompact threshold, with the two boxes under the margin control.

Loading the readout box…

What to keep in mind

  • This is an accounting model, not a measured performance formula. Tokenization varies, and only the host’s own report for the assembled request is authoritative.
  • What’s left is not uniformly useful. Research on position sensitivity and on degradation as input grows points the same way. Neither establishes a universal threshold. A large remaining budget is permission to add evidence, not a guarantee that adding it helps.
  • The terms are not equally movable. Tool definitions are the most reducible, and turning tool search on is a configuration change. History grows on its own, so compacting and subagents are the levers. Instructions move when you trim files and memory. Reserved generation follows from what you ask for. The margin is a choice, set through the autocompact threshold, and shrinking it buys room at the price of less slack before compaction fires.