AI Agent Traffic: Crawler, Referral or Shopping Agent?

By Puneeth · · 8 min read
AI Agent Traffic: Crawler, Referral or Shopping Agent?

Short answer: Three different kinds of AI traffic reach your site, and most stacks treat them as one thing. A crawler reads your pages to train or index a model and will never buy. A referral is a human who arrived from an AI answer and behaves like any other visitor. An agent is software browsing on behalf of a named person, which can add to cart and check out. The distinction matters because your analytics and your bidding treat all three as noise or as signal, with no middle setting. GA4 silently removes traffic on the IAB known-bot list, and Google states you cannot see how much it removed. Anything not on that list, including agents acting for real buyers, stays in your data and trains your campaigns.

Three species of AI traffic

Crawler AI referral Shopping agent
Who starts it The model provider, on a schedule A person, who then clicks through A person, who delegates the task
Examples GPTBot, OAI-SearchBot, Google-Extended, Google-CloudVertexBot A visit from a ChatGPT, Perplexity or Gemini answer ChatGPT-User, Google-Agent and similar user-triggered fetchers
Runs JavaScript Usually not Yes, it is a browser Often yes
Can convert No Yes Yes, and increasingly does
Respects robots.txt Yes, when well behaved Not applicable Often not: Google says user-triggered fetchers "generally ignore robots.txt rules"
What you want Let it read, keep it out of analytics Measure it as a channel Let it through, and attribute the sale correctly

The middle column is well covered elsewhere. The first and third are where measurement goes wrong.

What the platforms actually publish

This is documented, not speculative.

OpenAI lists four separate agents with different jobs. GPTBot crawls content "that may be used in training our generative AI foundation models". OAI-SearchBot powers search results. OAI-AdsBot checks pages submitted as ads. ChatGPT-User handles "certain user actions in ChatGPT and Custom GPTs", and OpenAI notes those actions "are initiated by a user", which is why robots.txt rules may not apply to them.

Google splits its fleet the same way. Google-Extended governs whether your content trains Gemini, and does not affect Search inclusion. Google-CloudVertexBot crawls sites at their owner's request for building Vertex AI agents. Separately, Google documents user-triggered fetchers, including Google-Agent for agents on Google infrastructure, and states plainly: "Because the fetch was requested by a user, these fetchers generally ignore robots.txt rules."

Read those two paragraphs again from a measurement point of view. The providers already distinguish between a machine reading your site and a machine acting for a person. Most marketing stacks do not.

What each one does to your data

Crawlers inflate what you never see. A crawler that executes no JavaScript never reaches GA4, so it shows up in server logs and CDN bills rather than reports. A crawler that does execute JavaScript, and is not on the IAB list, lands in your analytics as a one-page session with no conversion. At scale this drags down conversion rate, inflates bounce, and quietly skews any model trained on those sessions, in the same way invalid traffic does.

AI referrals arrive mislabelled. A person who clicks through from an AI answer often carries no useful referrer, so the session lands in Direct. If your Direct channel has grown without explanation over the past year, part of that is AI referral traffic you are not crediting.

Agents break identity, not just counting. This is the expensive one. An agent session may complete a purchase for a real customer while looking nothing like that customer: a datacentre IP, a fresh browser profile, no cookie history, no prior sessions. Block it and you have blocked a sale. Let it through unexamined and you attribute a real purchase to an anonymous one-session visitor, which is a poor training example for every bidding model you feed. That is an identity problem before it is a reporting one.

Why your current defences don't resolve it

GA4 excludes "traffic from known bots and spiders" automatically, using the International Spiders and Bots List maintained by the IAB. That is sensible, and it has two consequences people miss. First, Google states you "cannot disable known bot traffic exclusion or see how much known bot traffic was excluded", so the volume is invisible rather than reported. Second, the list names known robots. A user-triggered agent running a real browser for a real buyer is not a known robot, so it stays in.

So the default setting is backwards for both cases. Crawlers that should be counted and understood are removed without a number. Agents that need careful handling are treated as ordinary humans, while the bots acting for themselves are a separate problem again.

How to tell them apart

Four signals, in order of reliability.

  1. User agent. The easiest and the least complete. GPTBot, OAI-SearchBot, ChatGPT-User and Google-Agent all identify themselves in the string, and the providers publish them. Maintain the list; it changes.

  2. Reverse DNS and published IP ranges. The defence against spoofing. A user agent can claim anything; provenance is harder to fake.

  3. Behaviour. Crawlers read broadly and never return. Agents move with purpose: a search, a product page, a cart, often faster than a person and without the hesitation patterns humans show.

  4. Outcome. The decisive one. Did a session complete a purchase or a qualified action? An agent that buys is a customer interaction, whatever its user agent says.

Running signal four alongside the first three is what separates measurement from blocking. Security tools stop at signals one and two, because their question is whether to allow the request. The measurement question is different: what should this session mean to your attribution?

A decision framework: block, allow, or label

For each class of AI traffic, decide once and apply consistently.

Traffic Serve it? Count it in analytics? Send it to ad platforms?
Training crawler Your call, it's a content licensing question No, segment it out Never
Search and answer crawler Yes, this is how you get cited No Never
AI referral Yes Yes, as its own channel Yes, it's a human
Shopping agent, no conversion Yes Yes, labelled as agent No
Shopping agent, converted Yes Yes, labelled and stitched to the customer Yes, with the identity resolved

The last row is the one worth engineering for. An agent-completed purchase is real revenue, and it should reach your ad platforms as a conversion tied to a real person, not as an anonymous session that teaches your bidding to chase datacentre traffic.

What to do this quarter

  1. Measure before you decide. Segment your traffic by user agent and look at what each group does. Most teams discover the volumes are not what they assumed, in both directions.

  2. Give AI referrals their own channel group in GA4, so they stop hiding inside Direct and you can see whether they convert.

  3. Label agent sessions rather than blocking them. A label is reversible; a block costs you a sale you never learn about.

  4. Resolve identity on agent-completed purchases. Email at checkout, order ID, or an existing customer record ties the sale to a person. Ingest ID is built for this, and it is the difference between a useful conversion signal and a misleading one.

  5. Score what's left. Ad Shield gives every visitor a Traffic Quality Score and acts on rules you set, which is how you keep crawler and low-intent sessions out of the conversion data you send to Google, Meta and the rest.

  6. Write down your policy. This changes every few months. A documented position is easier to revise than a set of inherited filters nobody remembers creating.

What this looks like in a year

Agent-completed purchases are a small share of ecommerce today and a growing one. The teams that will handle it well are the ones that can already answer a simple question: of the sessions on our site last month, how many were human, how many were agents acting for humans, and how many were neither? Most stacks cannot answer that today, and the default settings are designed so that you never notice.

If you want the answer for your own site, start with the free Site Intelligence audit. It reports what is firing, what is missing, and what share of your traffic does not look human.

Frequently asked questions

What is AI agent traffic?

AI agent traffic is software browsing your site on behalf of a person, rather than a crawler collecting data for a model. Examples include OpenAI's ChatGPT-User and Google's user-triggered fetchers such as Google-Agent. Unlike crawlers, agent sessions are initiated by a user, can run JavaScript, and can complete actions including purchases.

Is an AI agent the same as a bot?

No, and the distinction has money attached. A bot acts for itself or its operator. An agent acts for a named person who asked it to do something. Blocking bots protects your budget; blocking agents blocks customers. The practical test is outcome: a session that completes a purchase for a real buyer is a customer interaction regardless of what the user agent string says.

Does GA4 filter AI agent traffic?

GA4 automatically excludes traffic from known bots and spiders using the International Spiders and Bots List maintained by the IAB. Google states you cannot disable that exclusion or see how much was excluded. Agents that are not on the list, including user-triggered ones, are not filtered and appear in your reports as ordinary sessions.

Should I block AI crawlers in robots.txt?

It depends what the crawler does. Blocking a training crawler such as GPTBot is a content licensing decision. Blocking a search crawler such as OAI-SearchBot removes you from the answers people see, which is the opposite of how AI engines read your site. And it is not a reliable control for agents: Google states that user-triggered fetchers generally ignore robots.txt rules, because a person asked for the fetch.

How do I see AI referral traffic separately?

Create a channel group in GA4 that matches the referrers for ChatGPT, Perplexity, Gemini, Copilot and the rest, rather than leaving them to fall into Direct. Review it monthly, because the referrer behaviour of these products changes without notice.

Will blocking agent traffic hurt my conversions?

It can, and the loss is invisible because a blocked agent never becomes a session you can analyse. That asymmetry is the argument for labelling rather than blocking: a labelled session can be excluded from reporting later, while a blocked one is simply gone.

How do I attribute a purchase made by an agent?

Resolve the identity at the point of purchase, then send the conversion with that identity attached. The order carries an email or customer record, so the sale can be tied to a person rather than to an anonymous datacentre session. Without that step, your ad platforms learn from a conversion that looks like a bot.

Sources