In April 2026, Uber's CTO confirmed the company had burned through its entire 2026 AI budget in four months after rolling Claude Code out to roughly 5,000 engineers. A month later, ServiceNow's CIO disclosed nearly the same thing: the full-year Anthropic budget, gone in the first few months. She called it "a really hard problem."
ServiceNow is worth sitting with. This is a company that sells AI governance tooling, and whose own chief customer officer had publicly warned that tokenmaxxing was "a short-lived hype cycle" because "there's a bill to pay for those tokens." They saw it coming, said so out loud, and still blew through the budget.
When the companies most equipped to manage this can't, it stops being a discipline problem.
Cheaper tokens, bigger bills
Token prices have fallen more than 50 times in two years. And spend keeps climbing anyway.
That has a name. In 1865, William Stanley Jevons noticed that as steam engines got more efficient, Britain burned more coal, not less, because efficiency made coal worth using for things nobody would have bothered with before. Satya Nadella said it directly when DeepSeek shipped: "Jevons paradox strikes again." Every price drop moves a tier of previously uneconomic ideas into the obviously-worth-doing column. Prices already fell. The bill went up.
I've been building Optimizely Mark, our agent platform for marketers, for close to two years now. For most of that time, the industry was asking what AI could do. That question was arguably the easy one. The one that replaced it is harder: where does AI earn its keep, and at what cost?
The advice is aimed at the wrong person
Most tokenomics guidance reads like a diet plan for end users. Pick a cheaper model. Write shorter prompts. Turn off extended thinking. Don't paste the whole document.
Now picture the person you're giving that to. A marketer running a content refresh across 400 pages. They don't know what a frontier model costs relative to a mid-tier one. They can't see how many tool definitions got loaded into their agent's context. They have no idea their conversation history has been carried forward, in full, for 30 turns.
None of that is their failure. It's a design failure, and it belongs to the harness: the system layer sitting between the user and the model.
Four levers control agent spend. One or two are genuine user choices. The rest should happen automatically, every run, without anyone thinking about it.
The 4 levers of agent token spend
Lever one: Which model is smart enough?
Artificial Analysis publishes a chart worth bookmarking: it plots models on two axes, cost per intelligence on the x-axis, intelligence index on the y-axis. The interesting real estate is the top-left. Low cost per task, high intelligence score. New models are increasingly living there, delivering serious intelligence at a fraction of frontier prices.
That shifts the question from which model is smartest to which model is smart enough. The cost of getting this wrong is not marginal: the gap in input token costs between high scoring models can be upwards of 50x.
Multi-model support is table stakes. But a model picker alone isn’t the answer, it's a tax. Every dropdown asking "which model do you want?" is a question the user isn't equipped to answer. The system should classify the task and route it, with per-agent overrides for teams that want explicit control.
Lever two: The volume is in the input, not the answer
Context is where the bill sneakily grows, and it's almost invisible from the user's seat.
We created an industry-wide benchmark for AI models (Mark-Bench) based on 285 marketing tasks across 15 functions and 6000+ criteria. This rubric scores AI model quality and cost. We measured 55 times more input tokens than output tokens per task. While output costs more per token, input typically drives the bill.
Rich input is the price of quality output, and telling people to send less context to save money is telling them to accept worse results in many cases. So, the question is how to curate the input intelligently.
Here are five things our AI harness can do: