Value per token: making AI cost a first-class product decision

AI spend keeps climbing, and waiting for price cuts is not a plan. Usage compounds, and newer models spend more tokens per task. The teams in control treat value per token as a product metric: attributed to real work, proven by evals, and managed with the same discipline as latency or uptime.

One token, weighed against the cost of serving it.

Why the bill keeps growing

Ask anyone running an AI budget where it is heading and the honest answer is up. Frontier prices hold firm while volume compounds underneath them, and the newest models, the ones everyone actually wants, run more verbose and reasoning-heavy, spending more tokens per task. The units get cheaper, the unit count grows faster, and the bill still climbs.

You can't price-cut your way out of that. So the question changes. It moves from what a token costs to whether each token earns its keep, and that is what tokenomics means.

Flat prices are small comfort when volume climbs like this.

What tokenomics is, and why the token is the unit

Tokenomics is knowing what every token costs you and what it earns you, then ensuring that return on token investment is sufficient, and acting when it isn't.

Tokens matter more than dollars because the token count is where you can act: compress prompts, cap output, cache what repeats, route work to the right model. You can't optimise an invoice directly. Keep both in view, though: tokens are the lever you pull, but the dollar each one costs is the reason you're pulling it.

A token is roughly four characters, about three-quarters of a word. Input and output are billed separately, and output runs four to six times dearer. With reasoning models you also pay for "thinking" tokens: internal reasoning, billed at output rates, that never appears in the response.

Two things make the token a treacherous unit. First, tokens are model-specific and shift under you. A tokeniser change can produce materially more tokens for the same text at an unchanged headline price. Second, this is not cloud FinOps rebadged. A vCPU-hour is stable and a token is not, so you are dealing with sparse telemetry, a moving billing unit, and agents that can burn through a budget faster than any cloud overspend ever could.

The sharpest difference is predictability. Cloud consumption is broadly forecastable: instances run for known hours, and last month is a decent guide to next month. Token consumption is not. The same request can cost pennies or pounds depending on how much context it drags in, how long the model reasons, and how many loops an agent takes. One badly scoped agent can spend more in an afternoon than a team did all month, so forecasts built on averages will mislead you: the money sits in the tail.

A token is roughly four characters. Output is billed 4–6× dearer than input, and reasoning models bill their thinking at output rates.

Why frontier prices behave this way

For a fixed capability, prices do fall; hardware, quantisation and competition see to that. The frontier resists the pull because it follows a pharmaceutical model: a flagship is enormous fixed R&D recovered through near-marginal-cost token "manufacturing", so each launches at a premium to fund the next run, and only the previous frontier gets discounted toward commodity. Demand compounds the effect: consumption outruns any price decline. Waiting for the frontier to get cheap is not a strategy.

What you can use today is the spread. Inside Anthropic alone, Haiku 4.5 at $1/$5 to Fable 5.1 at $10/$50 is a 10× span, and it is wider still against GPT-5.6 Luna at $0.20 input. A gap that large inside one vendor is why routing between models matters more than any single default choice. We come back to it in the levers below.

Labels show input / output price per million tokens. The gulf is the argument for routing.

Value, the hard part

Measuring cost is arithmetic. Measuring value is not, and the evidence so far is not encouraging.

Two large studies land in the same place. MIT's Project NANDA finds that almost all enterprise GenAI pilots show no measurable P&L impact. The headline number is contested, built on a small sample and a narrow definition of success, but it is hard to dismiss on direction. McKinsey, by a different method, finds that fewer than half of adopters report any EBIT impact at all, and only around six per cent report material gains (both are shown below). The methods differ and the conclusion is the same: most AI spend is not demonstrably earning.

Value rarely hides in clean cost-substitution. It surfaces as experimentation velocity, and as people shifting to higher-level work, second-order effects that never appear on the token bill. If you only look for headcount savings, you will conclude the spend earns nothing, and you will be measuring the wrong thing.

Experimentation velocity is the clearest example, because it is value that hides in plain sight. Trying an idea used to mean a spike of engineering time; now it can mean a few dollars of tokens and an afternoon. That shift is bigger than it sounds. When testing a hypothesis is nearly free, you stop rationing experiments and start running them side by side, which means you find the approach that genuinely works rather than committing to the first one that seemed plausible. The token cost of ten discarded prototypes is trivial against the cost of shipping the wrong design, but the bill records the ten prototypes, while the disaster you avoided appears nowhere. Measured on spend alone, cheap experimentation looks like waste; measured on decisions reached faster and better, cheap experimentation is often where the real return sits.

One counterintuitive pattern is worth naming. In our experience, more spend does not reliably buy more accuracy. Teams often find quality peaking at an intermediate cost point, with longer reasoning budgets adding tokens faster than they add correct answers. Treat "spend more, get more" as a hypothesis to test rather than a rule.

Both methods point the same way: most AI spend is not demonstrably earning.

Measuring value: the eval loop

Start with granularity. You can attribute spend per app, per team, per user, per repo, per pull request. The core shift is from inputs like tokens and seats to outcomes: cost per resolved ticket, per shipped PR, per dollar of pipeline.

Two metrics dominate. At the model layer, intelligence per dollar, which independent benchmarks such as Artificial Analysis track and which moves quarterly. At the application layer, cost per outcome, the number the boardroom actually thinks in, translated into unit margin.

Most tokenomics writing stops before this step. Use a judge model to audit a sample of production traces and score three things: path efficiency (did the agent take a clean route, or burn tokens in wasted loops?), outcome success, and the human counterfactual (what would the same task have cost a person unaided?). AI cost against human cost for the same outcome is the ROI number. Take off the human time spent prompting, reviewing and correcting, and you have Net AI Leverage: how much work the AI did, net of the effort it took to steer it.

Be honest about how far that number goes, though, because value is genuinely hard to measure and the counterfactual is where it gets slippery. Asking what a human would have charged assumes the work would have happened at all. Much of what AI produces is work nobody would have commissioned: the analysis that wasn't worth a fortnight of someone's time, the prototype that would never have cleared a prioritisation call. Score that against a human day rate and you manufacture savings on work that was never in the budget. The reverse case is just as real: a task that genuinely displaces two days of skilled effort might barely register as tokens, and its true worth is whether the team could then do something it otherwise could not.

So put a second measure next to it: Effective AI Leverage — the share of that work that actually advanced something, whether an issue, a bug, a task, a team goal or the organisation's mission. The gap between the two is Wheelspin: output that burned tokens and effort but advanced nothing. A cheap output that advances nothing is a poor trade at any token price, and an expensive one that unblocks a strategic programme is a bargain. Which is why Net AI Leverage should never be reported on its own; without Effective AI Leverage beside it, Wheelspin looks like productivity. Cost per outcome is measurable and worth tracking; value against the mission needs human judgement, and no judge model will hand it to you.

Wheelspin is output that burned effort but moved nothing forward. Proportions are illustrative.

Three caveats apply. The judge has a meta-cost, so batch the eval runs. Judge reliability needs human spot-checks, and verbosity bias is the one to watch. The human baseline is an estimate, so the ROI is only as good as the loaded-cost guess underneath it.

Gamify it

One lever most teams miss entirely is making the numbers public inside the organisation. A metric only changes behaviour if the people doing the work can see it, yet tokenomics usually lives in a finance dashboard the delivery team never opens. Put the scores where the work happens, team by team: a prompt-quality rating, a token-efficiency figure, a cost per outcome. Then give those numbers meaning, because a score of 0.42 means nothing without context. Show what good looks like, what poor looks like, and where each team sits against the others. A little visible, friendly competition tends to improve efficiency faster than any mandate from above, because people would rather beat a benchmark than be told to hit one.

Make the scores visible team by team, and the numbers start to move on their own.

Measurement itself has a human cost. Per-user token attribution curdles into productivity-policing, so attribute at team, workflow or feature level, and frame the counterfactual as what you freed people from rather than who you could replace. Then track engagement, retention and after-hours toil as outcomes in their own right rather than a safety check. They are the human canary: a team reporting its work got more engaging is often the earliest evidence a deployment is working, well before cost per outcome moves. Watch only the ledger and you see the effect late.

A judge model audits production traces; its scores feed routing, prompts and budgets.

Controlling cost: levers and guardrails

The technical levers are largely solved. Caching, batching and their smaller cousins (semantic caching, retrieval, output caps) are increasingly table stakes, pulled already or built silently into the gateways and platforms you buy; it is worth re-checking the rates, though, since vendors keep moving them (Fable 5.1 cache reads now bill at 2.5% of input, not the usual 10%). Routing is the one still worth active attention, because the 10× spread from earlier is the whole case for it: send simple queries to cheap models, escalate only the hard ones, and the cascade savings shown below follow with little loss of quality.

The lever that has not been properly considered is human: educating the people spending the tokens. Most remaining waste comes from users who don't know that a sharper prompt gets a better answer in one attempt instead of five, or which tasks deserve the expensive model. A modest investment in enablement raises both sides of the ratio at once: less spend per task, and more people getting real value from the tools at all. The leaderboard from earlier only works if people know how to move their score.

Build versus buy has a threshold. Self-hosting only starts to pay at sustained, GPU-saturating volume: for most teams that is hundreds of millions of tokens a day, not millions. Below it, you are buying yourself an SRE problem rather than a saving.

Then the guardrails. None of this is novel; subscription tools and the better AI gateways already ship stacked rate limits. Even so, teams tend to stop at a single monthly budget and miss what the stack is for. Each window does a different job. A per-minute limit protects infrastructure, smoothing bursts so one workload cannot starve everyone else of capacity. A per-session limit, say a five-hour working block, catches the runaway agent or the user stuck in a loop before a single session does real damage. The longer windows, per day, per week, per month and per project, are the budget caps: they are not about capacity at all, they exist so that spend cannot quietly drift past what you agreed to.

Set each of them per team, per feature and per customer, so every limit doubles as an attribution boundary and an alert can point to where spend spiked. The short windows keep you safe operationally, the long windows keep you solvent, and you want both. A system that only checks a monthly ceiling can burn a quarter's budget in a bad afternoon and not say a word until month end. And keep one blunt rule in reserve: if a workflow's token use grows faster than its measurable output, pause it and re-scope.

The judge's path-efficiency score from the eval loop feeds these routing and budget decisions directly, which is why measurement and control are really one system.

Caching, batching and routing compound to 60–80% off, so audit that your stack is actually delivering it.

The maturity ladder

Four steps, in order. First, visibility, because you cannot manage what you cannot attribute. Second, confirm that the table stakes of caching, batching and routing are actually delivering their 60–80%, and invest in enablement where the remaining waste lives. Third, govern your agents with budgets, owners and kill-switches. Fourth, re-anchor measurement on outcomes, with the eval loop doing the auditing and Effective AI Leverage, not Net AI Leverage alone, as the number you report.

What would change our mind? If frontier prices start following the commodity curve down and agentic growth plateaus, the balance tips toward spend freely, optimise later. If power stays the binding constraint, aggressive governance and provisioned commitments win. Watch three indicators: intelligence per dollar on independent benchmarks, your own cost per outcome, and the warning sign of the two diverging: cost per outcome improving while engagement or retention slides.

Value per token is now a first-class delivery decision. The teams that treat it that way will deliver more, and more sustainably, than the ones still waiting for the bill to fix itself.

Four steps: from attributing every token to anchoring the whole system on outcomes.

Next
Next

Cloudscaler achieves Select partner status in Claude Partner Network Services Track