Your marketers shouldn't have to be token economists

Nikita BokilNikita Bokil
Sep 27, 2026

TL;DR

  • Token prices have fallen more than 50x in two years and enterprise bills continue to climb. That's Jevons paradox, and it means waiting for prices to fall is not a cost strategy.
  • Four levers control an agent's token spend:
    • Right-size the thinking. Not every task needs the most capable model. The right system routes automatically.
    • Curate the context. Control what enters the model before it starts thinking. Less noise, fewer wasted tokens, better output.
    • Guide the agent. Watch for anomalies in real time, rescue inefficient paths, and stop unrecoverable runs before they burn budget.
    • Skip the LLM. For tasks that don't need language understanding, use code or purpose-built decision models. No LLM call, no token costs. 
  • Controlling AI spend is the system's job. Your marketers already have one.

 

In April 2026, Uber's CTO confirmed the company had burned through its entire 2026 AI budget in four months after rolling Claude Code out to roughly 5,000 engineers. A month later, ServiceNow's CIO disclosed nearly the same thing: the full-year Anthropic budget, gone in the first few months. She called it "a really hard problem."

ServiceNow is worth sitting with. This is a company that sells AI governance tooling, and whose own chief customer officer had publicly warned that tokenmaxxing was "a short-lived hype cycle" because "there's a bill to pay for those tokens." They saw it coming, said so out loud, and still blew through the budget. 

When the companies most equipped to manage this can't, it stops being a discipline problem. 

 

Cheaper tokens, bigger bills 

 

Token prices have fallen more than 50 times in two years. And spend keeps climbing anyway.

That has a name. In 1865, William Stanley Jevons noticed that as steam engines got more efficient, Britain burned more coal, not less, because efficiency made coal worth using for things nobody would have bothered with before. Satya Nadella said it directly when DeepSeek shipped: "Jevons paradox strikes again." Every price drop moves a tier of previously uneconomic ideas into the obviously-worth-doing column. Prices already fell. The bill went up.

I've been building Optimizely Mark, our agent platform for marketers, for close to two years now. For most of that time, the industry was asking what AI could do. That question was arguably the easy one. The one that replaced it is harder: where does AI earn its keep, and at what cost? 

 

The advice is aimed at the wrong person 

 

Most tokenomics guidance reads like a diet plan for end users. Pick a cheaper model. Write shorter prompts. Turn off extended thinking. Don't paste the whole document.

Now picture the person you're giving that to. A marketer running a content refresh across 400 pages. They don't know what a frontier model costs relative to a mid-tier one. They can't see how many tool definitions got loaded into their agent's context. They have no idea their conversation history has been carried forward, in full, for 30 turns.

None of that is their failure. It's a design failure, and it belongs to the harness: the system layer sitting between the user and the model.

Four levers control agent spend. One or two are genuine user choices. The rest should happen automatically, every run, without anyone thinking about it. 

 

The 4 levers of agent token spend 

 

Lever one: Which model is smart enough? 

 

Artificial Analysis publishes a chart worth bookmarking: it plots models on two axes, cost per intelligence on the x-axis, intelligence index on the y-axis. The interesting real estate is the top-left. Low cost per task, high intelligence score. New models are increasingly living there, delivering serious intelligence at a fraction of frontier prices. 

That shifts the question from which model is smartest to which model is smart enough. The cost of getting this wrong is not marginal: the gap in input token costs between high scoring models can be upwards of 50x.

Multi-model support is table stakes. But a model picker alone isn’t the answer, it's a tax. Every dropdown asking "which model do you want?" is a question the user isn't equipped to answer. The system should classify the task and route it, with per-agent overrides for teams that want explicit control. 

 

Lever two: The volume is in the input, not the answer 

 

Context is where the bill sneakily grows, and it's almost invisible from the user's seat.

We created an industry-wide benchmark for AI models (Mark-Bench) based on 285 marketing tasks across 15 functions and 6000+ criteria. This rubric scores AI model quality and cost. We measured 55 times more input tokens than output tokens per task. While output costs more per token, input typically drives the bill. 

Rich input is the price of quality output, and telling people to send less context to save money is telling them to accept worse results in many cases. So, the question is how to curate the input intelligently. 

Here are five things our AI harness can do: 

1

Selective context enrichment: The right context beats all the context. Better answer, lower cost.

2

Selective tool loading: Marketing agents touch a lot of systems. That doesn't mean handing the model 1,000 tool definitions when the task needs 12.

3

Prompt and response caching: The biggest single win. In Mark it cuts token cost by more than 90% on repeated patterns.

4

Context compaction: Long-running agents shouldn't carry their entire history on every turn.

5

Sub-agent delegation: Self-contained work gets handed to a sub-agent with its own isolated context, avoiding context bloat. 

 

None of this replaces good prompting or thoughtful agent design. But prompt craft can't cache a repeated system prompt, prune a 30-turn history, or keep raw research out of the main chat thread. Those are system decisions, and they apply to every run regardless of who wrote the prompt. 

 

Lever three: Reliability is a cost lever 

 

Reliability normally gets filed under quality. It belongs under cost too. When an agent goes off the rails at 2am it burns tokens and compute until something stops it, and if nothing is watching, that something is the invoice.

Guiding the agent means acting at three moments: 

Before the run: input guardrails block work the agent was never built to do.

During the run: behavioral monitoring watches for anomalies in real time and triggers a rescue or course correction.

After the run: output evals validate quality before results reach users. 

 

Most teams have the third. Far fewer have the first and second. I've written more about why stopping a failing agent and saving the run are not the same thing, and what a system that can do both actually looks like, here. 

 

Lever four: Does this need an LLM at all? 

 

This one is the most counterintuitive lever when talking about agents. Using an LLM for classification, routing, or validation is like driving a Ferrari to get groceries. It technically works, but it also costs way more than it should. Code handles deterministic logic and routing. Classical ML handles classification and scoring. LLMs handle language understanding, creative generation, and open-ended reasoning, and that's where they should stay. 

 

TypeSafe AI released Jev, which gives up free-form text entirely. You can't draft a campaign brief with it, but you can classify 50,000 support tickets, score lead quality, or check brand compliance before a human sees it. I suspect more advances are coming in this area, and they will reshape how agentic systems get designed.

A lot of what we route through an LLM isn't a language problem, it's a decision problem wearing language as a costume. Code handles the deterministic version. Purpose-built decision models are coming for the probabilistic one. Both are cheaper than asking a frontier LLM to think it over. 

 

The bar worth holding your platform to 

 
The four levers are the framework. In Optimizely Agent Platform, they run by default on every agent: multi-model support, auto model routing, caching, selective tool loading, context compaction, sub-agent delegation, behavioral monitoring, input guardrails, output evals, and a code execution environment for work that shouldn't touch a model at all. None of it requires configuring tokenomics per agent — including the one a marketer built on a Tuesday afternoon without telling anyone.
 
Our aforementioned benchmark report, Mark-Bench says that's worth something. Same model (Opus 5-XHigh), three setups: our harness ran 88% cheaper per task than the raw model — $4.91 against $40.52 — and scored higher on quality, 67% to 65%. Cheaper and better, from the layer around the model.
 
Efficiency that depends on user discipline isn't efficiency. It's a policy you're hoping people follow.
 
When you evaluate an AI platform, the useful question isn't just which models it runs. It's how much of this it handles without asking you. Does it route on its own? Does it cache by default? Does it know when an agent has gone off course, and do something about it? Does it know when not to call an LLM at all?
 
Controlling AI spend without sacrificing quality is a real goal, and those two things are not in tension. But it's the system's job. Your marketers have one already.