BlockRun
Assessment

Grow your agent usage 7x without growing the bill.

Most agent spend is not bought by a model being expensive. It is bought by context re-sent every turn, schemas nobody calls, and retries that were never going to succeed. We measure yours and hand back a policy.

See if it applies to you 
It has been done, publicly
9.4x
growth in weekly agent requests
flat
total AI spend over the same window
52%
lower cost per session, with the model held constant

Uber Engineering reported this over February to August 2026. Their result, in their environment, reported by them. Not a cheaper model and not a better rate — the model was held constant. Your workloads will differ.

Read their write-up ↗
What you get
01

What each task family actually costs to complete

Not cost per token or cost per call — cost to finish the job, including the retries and the failures, measured across the model set at real settled prices.

02

Where the money goes that buys nothing

Context that is re-sent every turn, tool schemas nobody calls, prose in descriptions, retries that were never going to succeed. Ranked by what it costs you, not by how easy it is to fix.

03

A routing and spend policy you can deploy

Which task family goes to which model at which budget ceiling, as configuration you can ship — not a slide recommending that you consider optimising.

How it runs
  1. 01You give us 20 to 50 of your real agent tasks — the ones that actually run, not a benchmark suite.
  2. 02We run each across the model set through our gateway at hard budget caps, at real settled prices.
  3. 03Success is program-verifiable only: tests pass, exact match, correct end state. Never a model grading a model.
  4. 04You get the numbers and the policy. We do not need your source, your data, or access to your systems.
We published our own bill first

Our own MCP server was putting 12,900 tokens of tool schema into every session before a single message was exchanged. More than half of that was prose sitting in descriptions, and about a fifth was removable in a day and a half. We published the number and the method before offering to measure anyone else.

Read the self-audit →
Start here

This is worth your time if agents are already running and the bill is already real.

Three questions, then book
Roughly what do you spend on models each month?
What do you run in production today?
How many repeated task families could you name right now?

Answer all three to continue. We use them to skip the discovery call.

Questions

What do you need from us to start?

A list of 20 to 50 real agent tasks and a way to tell whether each one succeeded. No source access, no data access, no infrastructure changes.

How is success measured?

Program-verifiable outcomes only — tests pass, exact match, correct end state. We do not use a model to grade another model's output, because that measures agreement rather than correctness.

Do we have to migrate anything?

No. The assessment runs through our gateway, and the output is a routing and budget policy. Deploying it is your decision and does not require moving your workloads.

Is this just telling us to use a cheaper model?

No. The result worth having is the one where the model is held constant and the cost per completed task still falls, because that is a change you keep when you upgrade models.