That's Mark-Bench, Optimizely's marketing AI benchmark. And we're here to tell you what it is, how it works, and why it matters.
Why marketing benchmarks are harder to build than they look
The reason a marketing benchmark didn't exist isn't that nobody thought of it. It's that building one properly is genuinely hard.
The first problem: There's no agreed-upon map of what marketing covers. Coding has well-defined task types. Law has established document categories. Marketing sprawls across brand strategy, campaign planning, SEO, CRO, PR, creative production, analytics, email, field events, paid media, personalization... the list goes on. Every vendor covers a different slice. Some focus on ads, some on customer experience, some on the marketer's day-to-day workflow. Nothing covers everything, end to end.
The second problem: Quality in marketing is subjective in ways that quality in code or law isn't. A piece of code either compiles and runs correctly or it doesn't. A campaign brief is harder to grade. You need rubrics — clear, specific criteria — and you need multiple evaluators to stop any one model's taste from becoming the standard.
The third problem: Benchmarks get gamed. Companies build models to score well on their own tests. A benchmark designed by the same team shipping the model is more of a product demo than a benchmark.
We built Mark-Bench with these three problems in mind.
What is Mark-Bench?
Mark-Bench is a standardized evaluation for AI performance across marketing tasks. It covers 17 marketing domains — everything from brand strategy and campaign planning to conversion optimization, creative production, PR, and field marketing —
It's always-evolving, currently assessing against 285+ tasks, 15 functions, and 6,523 criteria informed by real marketers, designed to reflect what marketers actually do, not what software vendors wish they did.
For each domain, Mark-Bench gives an AI agent a realistic scenario: a real brief, the source files it would actually need (brand guidelines, style guides, source data), and a clear expected output format. That output is then scored against a rubric of 12 to 30 criteria, evaluated by multiple LLM judges (both Claude and Gemini) so no single model's preferences can skew the results.
The output: a clear, comparable score across models and tasks. You can see which models perform well at which tasks. You can see where the harness — AKA the system and context the AI operates within — makes a difference. You can finally put a number on what 'good' looks like for marketing.
How we built it, and what we learned
Before you can benchmark marketing, you need a complete map of what marketing covers. That map didn't exist off the shelf. So the team built it themselves, drawing from four distinct sources: