Is the difference real?
You tried a new way of working and it felt faster. Bring the timings from before and after, and this page tells you whether the data can actually tell the two apart, how wide the uncertainty is, and how many more tasks it would take to know. It also tells you when you’re measuring the wrong thing.
Predict first
Five tasks under workflow A and five different tasks under workflow B. The averages:
- A, mean minutes per task
- 48.0
- B, mean minutes per task
- 40.8
Your data
Start from a preset, or bring your own: one row per task, with its condition and how long it took. Everything is read in your browser and never sent anywhere.
Five different tasks each way, and B’s mean is exactly 15% lower. It looks like a win. With tasks this varied and this few, it’s noise.
Loading the data inputs…
Loading the data summary…
Can the data tell them apart?
Make your prediction above first. The analysis appears once you’ve answered, so the answer can’t lead you.
How many tasks would it take?
Loading the planner…
Feeling faster isn’t a measurement
Loading the perception chart…
What to measure, and what not to
Measure
- Time to an accepted result, not time to a first draft.
- Rework rate: how often a person has to touch it again afterward.
- Review minutes.
- Cost per accepted result, not cost per run.
- Escaped defects: the bugs that got past review.
Don’t measure
- Lines of code.
- Suggestion acceptance rate. It goes up when you stop reading.
- Raw token counts.
- The number of agents you have running.
Pick one endpoint before you start, and measure a baseline before you change anything. “15% faster on five tasks” is noise. “I feel faster” isn’t a measurement. “We didn’t measure it” is a perfectly legitimate answer.
If it matters, make it executable. Then look at the data to see whether it actually changed.
Everything here runs in your browser. Nothing you upload, paste, or type leaves this page, and a copied link holds only the settings and the preset, never your data.