The experimentation lifecycle, re-imagined
The fundamental experimentation lifecycle has not changed. Teams still ideate, plan, personalize, build, and analyze their tests. But how that work gets done within the lifecycle is drastically different with AI, resulting in much greater automation, velocity, and quality of output. This paper is organized along that lifecycle. Before walking through the stages, we start with the biggest shift in our data: the arrival of agentic experimentation itself.

1. Ideate
A good experiment starts with a good idea.
Over 90% of experiments target five common metrics: CTA clicks, pageviews, checkout, registration, and add-to-cart.
Our research indicates that some of the most impactful primary metrics are the least tested.
2. Plan
Setup quality is the single biggest predictor of win rate.
Sound experiments win 23 percentage points more often than poor ones (58% vs. 35%). Investing in setup quality is one of the highest-leverage improvements teams can make.
When choosing a hypothesis, most teams default to a "Clarify Value" hypothesis, the safe first draft of a test plan. It's the choice in over a quarter of all experiments, and it only wins 41% of the time. Findability and Urgency hypotheses barely register by comparison, just 14% of tests combined, yet post the two highest win rates in the dataset: 60% and 56%.
Testing still clusters at the bottom of the funnel: 43% of all tests target it, where gains are the most incremental and hardest to detect. Post-conversion testing barely exists, at 4%. Scroll and engagement metrics sit at just 1.4% of tests but deliver one of the highest ATIs of any funnel metric, 2.1%. Revenue-per-visitor and purchase rate, by contrast, are used over 20% of the time but deliver only a third to a fifth of the value.
3. Personalize
Audience-targeted experiments generate 17.4% higher ATI on the audiences they target, though that can be diluted if the audience is small relative to total traffic.
This impact might be mitigated by the reach of the audience if the targeted audiences sum to less than the total traffic.
Also, after adopting Optimizely AI, companies run 65.4% more personalized experiments per year.
4. Build
The number of variations you build and the substance of the changes you make are among the strongest predictors of how much a test returns. AI changes the economics of building.
3 out of 4 experiments are still simple A/B tests yet multi-variant tests deliver up to 1.7x more impact.
The gap is widest in multi-variant tests: at 4 variations, Optimizely AI experiments show 4.1% average test impact (ATI) against 1.6% for non-Optimizely AI experiments.
The mix of what gets built also differs: Optimizely AI customers run a higher share of 3- and 4+-variation tests.
Multivariate tests are the extreme version, testing every possible combinationat once. They're rare, just 0.15% of experiments. But the ones that finish hint at real potency: of 44 meeting quality criteria, 8 produced significant winners, with a median uplift of 57%.
Complexity pays. Average test impact (ATI) rises with the number of distinct change types a variation combines.
Optimizely's Web Experimentation supports seven distinct change types that can be freely combined: attributes, custom code, redirects, insert image, insert HTML, widgets, and CSS. More complex experiments show greater returns, and Optimizely AI customers build a higher share of them.
In the year since Optimizely AI launched, customers have begun shifting towards higher complexity tests.
Bandits raise the ceiling and the floor if you give them options. At two variations, a bandit behaves like an A/B test: it beats an even traffic split just 49.7% of the time. With only one alternative to the baseline, there is almost nowhere for the algorithm to shift traffic. Add variations and the bandit starts working: at three or more, it beats an even split 63.9% of the time, 29% more often than at two variations.
5. Analyze
An experiment only creates value when its result is understood and acted on. The analysis stage covers the data foundation behind experimentation: the analytics used to explore results and form the next hypothesis, and the customer data that connects tests to business outcomes.
ODP customers show higher impact (ATI of 1.8% vs 1.4%) while velocity stays flat, consistent with richer customer data supporting more involved, more targeted tests.
For ODP customers who also use Optimizely AI, that trade-off disappears: they show an ATI of 2.1% alongside 29.1% higher velocity than the no-ODP baseline.