Understanding Your Data

AI bots

Citation, indexing and training are three different jobs behind one word, and telling them apart decides what you should do about them.

"AI traffic" is usually spoken about as one thing. It is not. Three quite different jobs hide behind the phrase, they arrive from different crawlers with different names, and they give you back three different things. Getting them confused is how sites end up blocking the traffic they wanted and paying for the traffic they did not.

#A subcategory of verified bots

AI is not a fifth verdict alongside human, verified, suspicious and fake. It sits one level below them.

Honeylog first decides whether a request is a verified bot at all, by checking the name it gave against where it actually came from. Only then does it ask what that bot is for, and AI is one answer among many. The same verified group holds site and network monitors, feed fetchers, link validators, security checkers, social media agents collecting a preview, read-it-later services and more.

What makes AI and SEO different is not that they are the only ones, it is that they are the two most sites need to act on, so the dashboard gives each of them a report of its own. Everything else verified sits together in the general Verified Bots report.

So nothing reaches this report without passing the identity check first. A request that writes an AI crawler's name on itself and comes from somewhere that name does not own is counted as fake, never here.

Inside the AI family, Honeylog keeps three categories apart. The dashboard labels them AI Citation, AI Indexing and AI Training.

#AI Citation

An assistant is fetching your page right now because somebody asked a question a moment ago. ChatGPT-User, Claude-User, Perplexity-User, DuckAssistBot, Gemini Deep Research and Google NotebookLM all work this way: the page is retrieved on demand and used to compose an answer for one person, a technique usually called retrieval augmented generation.

This is the closest thing in the whole bot world to a visit. Someone is being read your page, out loud, in the present tense.

Two things follow from that. The first is that whatever you return at that moment is what the person gets. A 403, a 500, a consent wall or a page that only assembles itself once a browser runs your JavaScript means you are simply absent from an answer that was about to include you. The second is that this is the category most likely to end in an attributed link and a click back to you, which is what Bot Conversions measures.

Its traffic also looks different from everything else in the product. It follows what people are asking about, not a schedule, so it arrives in bursts with quiet stretches in between. A flat week here is not a fault to chase.

#AI Indexing

These are the crawlers building the indexes that answer engines search. OAI-SearchBot, PerplexityBot, Claude-SearchBot, Applebot, Amazonbot and YouBot do for AI search roughly what Googlebot does for traditional search: they come back on their own schedule, systematically, to keep an index current. The rhythm is less predictable than a search engine's, often sporadic or bursty, but the intent is the same.

What this category buys you is eligibility. An assistant can only cite what its index already knows about, so being crawled here is the condition for being quoted later. It is the least visible of the three and the easiest to lose by accident, because losing it produces no error and no complaint, just a slow disappearance from answers you never see.

#AI Training

These crawlers download your content to build datasets that models are trained on. GPTBot, ClaudeBot, Google-CloudVertexBot, Ai2Bot, Diffbot and cohere-training-data-crawler are the familiar names.

What you get back is nothing. No index entry, no citation, no link, no visit, which is why Honeylog marks this as the one AI category that does not lead to human traffic. What you do get is the bill: bandwidth, CPU and, on a large site, real volume. Operators rarely publish how they choose sites or how often they return, and these crawlers can be heavier handed than search engines.

That makes it a question about cost and rights rather than visibility. It is also the reason to keep the record: if you ever want to raise licensing, or simply know what was taken and when, your own logs are the only evidence you have.

#Why the three must not be treated as one

Category Triggered by What you get back What blocking it costs
AI Citation Someone's question, right now A chance of a citation and a click Absence from answers people are already asking for
AI Indexing The crawler's own schedule Eligibility to be cited later A quiet disappearance from answer engines
AI Training A dataset being built Nothing Nothing you can see in your traffic

The practical consequence is that "block AI" is not a decision, it is three decisions taken carelessly at once.

Blocking GPTBot does not keep you out of ChatGPT's answers, because the crawler that reads your page for a user is ChatGPT-User and the one that indexes you is OAI-SearchBot. They are separate names with separate jobs. The opposite mistake costs more: a blanket rule in robots.txt or at the edge takes out indexing and citation along with training, and you vanish from answer engines while the traffic you actually minded keeps arriving under a name your rule never mentioned.

Training is the one you can refuse without a visibility cost. The other two are the ones that can send people back. Deciding per job is only possible once you can see which of the three is actually on your site, and in what proportion, which is what this report is for.

#Why these numbers hold up

Everything here is verified traffic. AI crawler names are among the most impersonated on the web, and an impersonator never reaches this report.

That restraint is the point. It is easy to publish a large AI traffic number, and much of what circulates elsewhere is inflated by requests that did nothing more than write a famous name on themselves. A smaller number you can act on beats a flattering one you cannot.

#Where to go next

The AI Bots report under Bots Traffic holds this traffic on its own. Visibility Radar follows a single crawler over time and lets you filter the list down to AI only. Bot Conversions picks up the other half of the story, the people who arrive after an assistant read your page.