Field guide
What actually drives your Claude Code bill
Three things I expected to matter turned out not to, and three I did not expect turned out to dominate. All of it measured from 17,424 real usage events on one working machine.
1. Model choice costs 4.5x per turn — and not because of length
The obvious guess is that Opus costs more because it writes more. That is not what the data shows. Opus turns used fewer tokens on average than Sonnet turns, and still cost 4.5 times as much each:
| Model | Turns | Cost | Per turn | Tokens / turn |
|---|---|---|---|---|
| claude-sonnet-5 | 16,015 | $2,593.14 | $0.162 | 399,308 |
| claude-opus-5 | 1,383 | $1,013.16 | $0.733 | 349,294 |
The gap is pure rate, not verbosity. Which means the lever is which model handles which turn, and nothing about how you write prompts will move it.
2. 8% of turns produced 28% of the cost
Those 1,383 Opus turns were 7.9% of all turns and 28.1% of all cost. If you are trying to bring spend down, that ratio is the whole game: a small number of model-selection decisions carries almost a third of the bill, and the other 92% of your activity is comparatively cheap noise.
This is also why “use Claude Code less” is bad advice. Using it less uniformly cuts mostly cheap turns. Moving a category of work from Opus to Sonnet cuts the expensive ones.
3. Spend is Pareto across projects, and not where you would guess
Across 32 projects, the single largest accounted for 35.7% of all spend, and the top three together for 77.1%. The remaining 29 projects shared under a quarter.
The useful part is that the top three were not the three I would have named before measuring. Perceived effort and actual token weight are only loosely related — a repo where you did a lot of thinking but little agent iteration is cheap; a repo where the agent ground through a large codebase repeatedly is expensive, even if it felt like a quiet week.
What does not drive the bill
- Prompt length. Across the same dataset, raw
input_tokenswere 0.00% of all tokens. What you type is a rounding error. - Output length. 0.15%. Also a rounding error.
- Number of sessions. Cost tracks context size multiplied by turn count, not how many times you opened a terminal.
Cache reads were 97.47% of all tokens — the same context being re-read on nearly every turn. That mechanism is covered in more detail in the usage-log field guide.
Checking your own numbers
Every figure above came out of local session logs, and you can produce the same breakdown for yourself in one command — no key, no account, nothing uploaded:
npx github:THEMANJH/agentspend-upload --report
It prints cost by project, by model and by day, and flags the cache ratio when it spots it.
Doing this for a team
The command above covers one machine. AgentSpend is the hosted version: everyone's spend in one dashboard split by person and project, with an email alert before you cross a budget. Flat $19/mo for up to 10 people, never metered on the spend you track.
Measured 23 August 2026 across 17,424 usage-bearing events from one heavily-used machine. Costs are estimated at API list price; on a Max subscription the real bill is flat, so these numbers show relative weight rather than an invoice. One machine is one data point — the method transfers, the exact ratios will not.