Don't bet on crawls that never pay out

Anna PleshakovaAnna Pleshakova
19. Aug. 2026

You know AI agents crawl your site invisibly. But most marketers miss what matters: whether the AI is training on your content or actively retrieving it for a buyer right now. Learn which bots are buying, and which are just borrowing.

You probably already know that AI agents are visiting your website. You probably also know that your traditional web analytics dashboards aren't showing you any of it. That gap is real, and it matters.

But here's the thing most teams miss once they start looking at the data: not all AI traffic is the same. A bot that visits your site to harvest content for model training is doing something fundamentally different from a bot that visits because a buyer just asked it a question. One is background noise. The other is an active conversation about your product, happening right now, without you.

Most organizations, when they first get visibility into their AI traffic, are surprised to find how much of it falls into the first category. Their content is feeding the machine. But it isn't getting a seat at the table.

That distinction - between content that gets indexed and retrieved on a buyer's behalf versus content that gets harvested and then sits in a training queue - is the most useful frame I've found for turning AI traffic data into actual decisions. The rest of this piece explains how to read that signal and what to do with it.

A bot that visits your site to harvest content for model training is doing something fundamentally different from a bot that visits because a buyer just asked it a question.

Anna Pleshakova|Senior Manager of Product Management, Optimizely

Traditional web traffic analytics have a blind spot

Here is the technical reason this problem exists, and I promise it is less boring than it sounds.

Traditional web analytics tools are client-side. They work by running a tiny script inside a browser. A human visits your page, the script fires, a data point lands in your dashboard. Clean, simple, reliable.

AI agents do not use browsers. They hit your site at the server level — the same layer where your CDN lives. No browser, no script, no data point. They come in through the back door, take what they need, and leave without signing the guestbook. Invisible to every client-side tool you own.

Cloudflare however, the CDN that the vast majority of Optimizely CMS customers already run on, is an example of a tool that catches everything at that server level. Those logs have always existed. We built a way to make them even more useful. Agent Visibility Analytics surfaces that infrastructure-level data directly inside Optimizely Analytics — real request-level signal: which agents visited, which pages, at what frequency, and with what intent. Not simulated behavior. Not a synthetic panel. Actual traffic.

This is where I want to push back on what already exists in the market.

Most tools operating in this space rely on synthetic testing. They prompt ChatGPT and Perplexity manually and observe outputs. That approach is slow, expensive, and directionally misleading — it tells you what an AI says about you, not what AI systems actually do when they visit your site. It's like getting a security camera that shows you someone came to your door, but not whether they were delivering a package or casing the place. Other tools that rely on factual log data show the request volume but miss an important question.

The question that actually matters is: why did the request happen?

We classify every AI request into one of three intent types:

1

Training — bots like GPTBot, ClaudeBot, and CCBot harvesting content to improve a model. This is batch and passive.

2

Search/Indexing — bots like OAI-SearchBot, Claude-SearchBot, and Googlebot navigating the public internet to build, index, and organize web content so it can be retrieved it later as needed. These bots crawl in advance to improve the speed and relevance of AI search results.

3

User Action (real-time retrieval) — bots like ChatGPT-User or Claude-User visiting your site directly to fetch content and answer a human's question in real time — not indexing for later, but retrieving now on a buyer's behalf. This, in my opinion, is the intent most marketers would want to optimize for.

Search and user action traffic have a direct, fast feedback loop — content retrieved today influences AI-generated answers today. Training traffic operates on a much longer horizon, tied to model retraining cycles that can span months. Critically, content that is structured well enough to be reliably retrieved by search agents will also make good training data — optimizing for retrieval is the active investment; training inclusion follows as a passive byproduct.

If your site is flooded with training requests but light on indexing and retrieval, your content is feeding the machine without getting a seat at the table. Knowing that ratio, and watching how it shifts as you make changes, is what turns a log file into a strategy.

What else would be useful in these kinds of tools?

Here is something I find genuinely exciting, and I think gets undersold: the ability to add custom dimensions to your request data.

Every business slices data differently. A hospitality company might want to segment AI requests by location. A B2B SaaS team might care deeply about which funnel stage their content falls into — are AI platforms retrieving your awareness pages? Your comparison pages? Your demo request pages?

Of course you can create and analyze this taxonomy manually but with a product like Optimizely's Agent Platform, you can define that classification taxonomy much faster. You describe what matters to your business — funnel stage, content topic, campaign goal, whatever — and an agent classifies your pages accordingly and feeds that context back into your dashboard. The result is not just knowing that AI is on your site, or even why, but understanding what it is interested in, mapped to the content taxonomy that reflects your actual business goals. For example, I can now see that my pages about the statistical significance topic get significantly more retrieval requests than others — so I know exactly where to invest attention next. That is a meaningful upgrade from "3,000 requests to /ab-testing-overview page URL" by itself.

Visibility is the foundation, not the destination. 

Imagine leveraging the power of AI to ask: “show me the top ten pages that are getting a lot of training traffic but almost no indexing or retrieval.” That is instantly a prioritized content action list – one that you can schedule to go out into your inbox weekly. Those are the pages your team should look at first — refresh the content, update the metadata, add FAQ schema, improve the structure. On the other hand, a page getting over 90% training requests but almost no indexing or retrieval is a signal the content is duplicative or not adding enough value.

The iteration cycle then looks something like this: AI detects agent behavior trends, surfaces prioritized recommendations to marketers, executes improvements using specialized agents after getting approval, measures & reports on the pre- and post-change impact, repeats.  This is something you can facilitate by stitching multiple products across each other today but Optimizely Agent Platform makes this easy to do all in one place.

Traffic signs illustration: getting traffic is the start, not the destination.

What happens in a year (and why the three-year question broke my brain)

Someone asked me recently how marketers will talk about AI visibility in three years the way they talk about organic search rankings today. Three years in this space feels like asking someone in 2005 how they would talk about social media strategy in 2015. I have guesses, not answers.

What I am more confident about is the next twelve months. I think proper experimentation on agentic traffic becomes real within a year. The same discipline that Optimizely has always been built around — controlled tests and defined metrics — applied to agent retrieval. The unit of measurement changes. Instead of a human click or conversion, you are measuring whether an agent successfully retrieved your content, what kind of retrieval it was, and which AI system performed it. Different methodology, same underlying principle: stop guessing and start knowing.

Beyond that, the longer arc is more fundamental. Static content, in my view, does not survive in its current form. The website's job is changing — from a destination for human visitors to a content repository, a retrieval system, and an interaction layer for AI agents. The teams building visibility and retrieval infrastructure now are not just keeping up. They are laying the foundation that everything else will run on.

The data you manage today — the intent signals, the content you structure for retrieval, the visibility you build right now — becomes the raw material for how buyers discover your products and make decisions in the world that is coming. Start before you feel ready!