Introducing Mark-Bench: The marketing AI benchmark we built because one didn't exist

Leah MessengerLeah Messenger
Sep 29, 2026

Every serious field that uses AI has a way to measure it. Coding has HumanEval and SWE-bench. Medicine has MedQA. Law has Harvey's legal benchmark; — a document-based eval that tests whether AI can handle the actual tasks lawyers do, not just abstract questions about law. There's even a site that catalogues the full landscape of AI benchmarks across domains.

Marketing had nothing.

Without a benchmark, "this AI is good at marketing" is a claim with no real meaning. Every vendor can say it. There's no shared standard to push back with, and no way to know if a model is genuinely strong at content strategy or just better at sounding like it is.

The longer you sit with that, the more it makes your brain tick and your eyes do that whole pondering-squint thing.

How do you have a real conversation about quality without one? How do you say your model is good, or your harness is good, without any objective way to measure it? It started to feel a little nonsensical to me. So we built one.

Shafqat Islam|President at Optimizely

That's Mark-Bench, Optimizely's marketing AI benchmark. And we're here to tell you what it is, how it works, and why it matters.

Why marketing benchmarks are harder to build than they look

The reason a marketing benchmark didn't exist isn't that nobody thought of it. It's that building one properly is genuinely hard.

The first problem: There's no agreed-upon map of what marketing covers. Coding has well-defined task types. Law has established document categories. Marketing sprawls across brand strategy, campaign planning, SEO, CRO, PR, creative production, analytics, email, field events, paid media, personalization... the list goes on. Every vendor covers a different slice. Some focus on ads, some on customer experience, some on the marketer's day-to-day workflow. Nothing covers everything, end to end.

The second problem: Quality in marketing is subjective in ways that quality in code or law isn't. A piece of code either compiles and runs correctly or it doesn't. A campaign brief is harder to grade. You need rubrics — clear, specific criteria — and you need multiple evaluators to stop any one model's taste from becoming the standard.

The third problem: Benchmarks get gamed. Companies build models to score well on their own tests. A benchmark designed by the same team shipping the model is more of a product demo than a benchmark. 

We built Mark-Bench with these three problems in mind. 

What is Mark-Bench?

Mark-Bench is a standardized evaluation for AI performance across marketing tasks. It covers 17 marketing domains — everything from brand strategy and campaign planning to conversion optimization, creative production, PR, and field marketing —

It's always-evolving, currently assessing against 285+ tasks, 15 functions, and 6,523 criteria informed by real marketers, designed to reflect what marketers actually do, not what software vendors wish they did.

For each domain, Mark-Bench gives an AI agent a realistic scenario: a real brief, the source files it would actually need (brand guidelines, style guides, source data), and a clear expected output format. That output is then scored against a rubric of 12 to 30 criteria, evaluated by multiple LLM judges (both Claude and Gemini) so no single model's preferences can skew the results.

The output: a clear, comparable score across models and tasks. You can see which models perform well at which tasks. You can see where the harness — AKA the system and context the AI operates within — makes a difference. You can finally put a number on what 'good' looks like for marketing.

How we built it, and what we learned

Before you can benchmark marketing, you need a complete map of what marketing covers. That map didn't exist off the shelf. So the team built it themselves, drawing from four distinct sources:

1

The open-source Marketing Skills repository: A community-built catalogue covering areas like analytics, competitive analysis, and AI SEO, representing how practitioners in the field define the scope of the discipline

2

Zapier's Automation Bench: Approximately 100 marketing automation tasks, which gave the team a concrete view of what tool-based marketing workflows actually look like in practice

3

Scott Brinker's State of Martech survey: An annual practitioner study by the editor of chiefmartec.com that tracks how marketers are actually using AI, based on real-world usage rather than vendor positioning

4

Optimizely's own internal agent library: Combined with direct sessions with practitioners across personalization, CRO, and other specialisms

AI handled the initial extraction, humans reviewed everything. That step, as the team puts it, was non-negotiable.

Before we get into how this was built, let us quickly introduce Proggo Pratik, Product Manager and the builder behind Mark-Bench. (Round of applause, please.)

For the task structure, Proggo drew on Harvey's legal benchmark as a structural reference. Harvey's approach — building evals around realistic documents and task briefs rather than abstract questions — translates well to marketing.

Each scenario has a brief, a set of input files, an expected output format, and a rubric with 12 to 30 criteria. Scoring is done by LLM judges. We use both Claude and Gemini to balance out any bias a single model might introduce.

Proggo Pratik|Product Manager

The team started with 15 domains, 13 of which made the first cut. After running the benchmark and reviewing results, they removed data governance, merged one domain into email marketing, and added creative production, PR and communications, and field marketing. The benchmark now sits at 17.

One early lesson came quickly: v1 rubrics were too easy. Across roughly 7,000 individual criteria, models passed more than 95% of them. The real differentiation only showed up in the all-pass rate (whether a model cleared every criterion in a scenario), which ranged from about 40% to 60% across models. Version 1.5 removed the criteria every model was clearing and replaced them with harder ones. "We kept doing that," says Proggo, "until the gap between a good and a bad model was clearly visible."

What it measures (and why that distinction matters)

Mark-Bench works two ways. Hold the harness constant and switch the models, and you get inter-model performance; which AI is actually best at your marketing tasks. Hold the model constant and switch the harness, and you understand how much your context, prompting, and workflow setup is doing the work. There's a cost impact on this, but more on that in the next section. 👀

That second dimension is more significant than it sounds. Most teams evaluate AI models in isolation, as if the model is the only variable that matters. In practice, the harness (the instructions, the context, the tools the model has access to) can swing results just as dramatically. Knowing which variable is driving performance changes where you invest.

Shafqat sees task-level granularity as the key unlock:

My hypothesis — still unproven — is that certain models will turn out to be significantly better at specific marketing tasks than others. A single pass rate tells you something. But knowing a model is strong at content strategy and weaker at conversion optimization, that's the level of granularity that actually changes how you work.

Shafqat Islam|President at Optimizely

The end-to-end vision is automatic model routing. Rather than asking marketers to choose the right AI for the job, the product chooses for them, based on evidence.

When a new model launches, we can run it through Mark-Bench within hours, understand exactly where it's strong and where it isn't, and automatically route different marketing tasks to the right model based on what we just learned. That's the goal. I don't think we're far away.

Shafqat Islam|President at Optimizely

What Mark-Bench tells us about cost

Choosing the right model isn't just a quality question, it's a cost one. And the numbers here are significant.

Nikita Bokil, part of the product team at Optimizely, ran a direct comparison using the same model in three different configurations. The result: the Optimizely harness ran 88% cheaper per task than hitting the raw model (Opus 5X-High) directly — $4.91 per task against $40.52 — and scored higher on quality in the process (67% versus 65%). Cheaper and better, from the layer around the model, not the model itself.

This is not a marginal efficiency gain. It's an order-of-magnitude difference driven entirely by how intelligently the harness uses the model; what context it passes in, what it routes elsewhere, and what it skips the LLM for entirely.

This is where Mark-Bench's task-level data becomes practically useful. Not every marketing task needs the same model — some need a fraction of the horsepower. The intelligence is in knowing which is which, and routing accordingly.

Using Optimizely's own AI product stack as an illustration of how that spectrum works in practice:

 

A simple chatbot query: A marketer asking what a metric means, or requesting a quick content summary. A lightweight, fast model handles this. Low cost, near-instant response, no heavy reasoning required.

An agent task: The Experiment Ideation Agent generating test hypotheses from your existing programme results. A mid-tier model with more context and reasoning, scoped to a single well-defined output.

A workflow agent: The Campaign Brief Agent running a multi-step sequence: research, structure, draft, review. The harness routes different steps to different models. Planning steps use heavier models; formatting outputs use lighter ones. The marketer doesn't need to configures none any of this.

Virtual Teammates: A Marketing Analyst running overnight anomaly detection, flagging performance drops before anyone has logged on. The harness makes every sub-task routing decision autonomously, matching model to task type in real time. The marketer didn't ask; the work just happened.

 

The point isn't that cheaper is always better, because no, it isn't. It's that the right model for the task is always better. Virtual Teammates operating at scale, running hundreds of sub-tasks across a workflow, need that routing intelligence built in. Without it, you're either overpaying for simple tasks or underpowering complex ones. Mark-Bench gives us the evidence base to get that routing right.

As Nikita's analysis so finely puts it: marketers shouldn't have to be token economists. The harness handles it. Mark-Bench is what tells us which model belongs where, using data rather than guesswork.

What it means for people doing the work

For SVP of Product, Kevin Li, the implications go further than picking the right model. A machine-readable rubric doesn't just help humans choose better, it gives AI a target to improve toward.

In most current AI workflows, improving outputs requires a human in the loop: a person reads the result, decides what's wrong, rewrites the prompt, and tries again. That loop is expensive and it doesn't scale. A rubric changes the equation. When quality criteria are written in a form a machine can check, the system can run its own improvement cycles — trying variations, keeping what scores higher, discarding what doesn't.

Once a rubric is machine-readable, the improvement loop doesn't need a human in it," Kevin says. "Give an agent a score and it has a goal to seek. Scores a 55, varies its approach, hits a 57, keeps the change. Over enough cycles the agent self-approves on work that clears the bar.

Kevin Li|SVP, Product at Optimizely

What that means for marketers in practice: less time rewording prompts, more time on work that actually requires human judgment.

Humans move up a level. Not inside the loop adjusting outputs. Above it — defining the standard, and showing up for the moments that actually need judgment. That's the creative work people came here to do. And it's the first thing to disappear when your week goes to rewording prompts.

Kevin Li|SVP, Product at Optimizely

On staying honest about it

Benchmark gaming is a known problem in AI evaluation. Models get fine-tuned to score well on specific tests. The benchmark stops reflecting real-world performance and starts reflecting how much effort a team put into optimizing for it. The result is numbers that look good and tell you nothing.

The team working on Mark-Bench team is explicit about not doing this. The goal is to run every model through the same test — including models Optimizely didn't build — and publish results that are genuinely useful to anyone trying to understand what works for marketing AI.

We want to be an honest, unbiased voice on what's actually good for marketing. That only works if the benchmark is as transparent and objective as possible.

Shafqat Islam|President at Optimizely

Human review is part of how they keep it honest. "You can use AI to generate a benchmark, but if you don't have subject matter experts checking the domains, the tasks, the rubrics, you just end up with more slop," says Shafqat. "The human judgment in the loop is what makes the benchmark worth anything."

What's next

Mark-Bench is a living benchmark. Rubrics will keep getting harder as models improve. New models will be tested as they launch, including models released after this post. Results will be shared publicly, with the methodology open for scrutiny.

The next version is already in progress. The rubric will be harder, the results will be different, and the team will share them when they are. That's kind of the point.

Want to hear more about what we're up to? Head to optimizely.com/ai.