Grow your agent usage 7x without growing the bill.
Most agent spend is not bought by a model being expensive. It is bought by context re-sent every turn, schemas nobody calls, and retries that were never going to succeed. We measure yours and hand back a policy.
See if it applies to youUber Engineering reported this over February to August 2026. Their result, in their environment, reported by them. Not a cheaper model and not a better rate — the model was held constant. Your workloads will differ.
Read their write-up ↗What each task family actually costs to complete
Not cost per token or cost per call — cost to finish the job, including the retries and the failures, measured across the model set at real settled prices.
Where the money goes that buys nothing
Context that is re-sent every turn, tool schemas nobody calls, prose in descriptions, retries that were never going to succeed. Ranked by what it costs you, not by how easy it is to fix.
A routing and spend policy you can deploy
Which task family goes to which model at which budget ceiling, as configuration you can ship — not a slide recommending that you consider optimising.
- You give us 20 to 50 of your real agent tasks — the ones that actually run, not a benchmark suite.
- We run each across the model set through our gateway at hard budget caps, at real settled prices.
- Success is program-verifiable only: tests pass, exact match, correct end state. Never a model grading a model.
- You get the numbers and the policy. We do not need your source, your data, or access to your systems.
Our own MCP server was putting 12,900 tokens of tool schema into every session before a single message was exchanged. More than half of that was prose sitting in descriptions, and about a fifth was removable in a day and a half. We published the number and the method before offering to measure anyone else.
Read the self-audit →This is worth your time if agents are already running and the bill is already real.
Answer all three to continue. We use them to skip the discovery call.
Questions
What do you need from us to start?
A list of 20 to 50 real agent tasks and a way to tell whether each one succeeded. No source access, no data access, no infrastructure changes.
How is success measured?
Program-verifiable outcomes only — tests pass, exact match, correct end state. We do not use a model to grade another model's output, because that measures agreement rather than correctness.
Do we have to migrate anything?
No. The assessment runs through our gateway, and the output is a routing and budget policy. Deploying it is your decision and does not require moving your workloads.
Is this just telling us to use a cheaper model?
No. The result worth having is the one where the model is held constant and the cost per completed task still falls, because that is a change you keep when you upgrade models.