Asva AIPower the future
By use case · Measure AI traffic

Stop losing AI revenue to the Direct bucket

AI referral traffic is systematically under-counted. Some assistants strip the referrer, some open links in contexts analytics cannot attribute, and the result lands in Direct alongside everything else of unknown origin. Teams then under-invest in a channel that is quietly working, because the dashboard says it sends nothing.

The other half of the problem runs the opposite way. AI crawlers and agents generate a meaningful share of raw requests, and when those are counted as sessions they inflate engagement and corrupt conversion rates. Measuring AI traffic properly means doing both jobs: taking the machines out of the human numbers, and putting the humans AI sent back into the channel they came from.

Why attribution needs its own approach here

Direct is hiding a real channel

When AI-assisted visits collapse into Direct, the channel has no measurable return and loses every budget argument by default. Recovering them does not create traffic that was not there; it moves visits that already convert into a bucket that can be reported, compared against paid and organic, and defended.

Crawler traffic is not audience traffic

A meaningful share of raw server hits are AI crawlers and agents. Counting them as sessions inflates page views, distorts bounce and time-on-page, and depresses conversion rate on exactly the pages crawlers like most. Separating them at the request level is the only way to keep human metrics honest.

Agent demand is a leading indicator

Which pages the retrieval agents fetch, and how often, tells you what the engines are currently reading about you, usually before it shows up in answers. A pricing page OAI-SearchBot starts fetching daily is about to be cited on pricing questions. A section no agent has touched in a month is not in any answer.

Retention is the silent failure

Edge providers expose AI crawler requests but keep them briefly, commonly around a day at lower tiers. A team that discovers this channel in month three has no history to compare against. Persisting the log from the first day is the cheapest decision in the programme and the one most often skipped.

How AI traffic measurement works

Separating agent traffic from human traffic, then recovering the human visits that AI actually sent.

  1. 1

    Separate agents from people at the request level

    Classify every request by user agent into human, retrieval agent, training crawler and other bot. Retrieval agents (OAI-SearchBot, ChatGPT-User, PerplexityBot, Perplexity-User, ClaudeBot) and training crawlers (GPTBot, CCBot, Google-Extended, Applebot-Extended) should be reported separately: one is an engine reading you to answer a question now, the other is a corpus being built. Do this at the edge or the server, where the raw request is visible, not in the analytics tool, which only sees what fired a tag.

  2. 2

    Identify AI-assisted human sessions

    Recover visits that arrived via an AI surface but lost their referrer, using the signals that survive the hand-off: an intact referrer where one exists, tracking parameters some assistants append, landing-page patterns only AI answers link to, and timing that follows a retrieval fetch of the same URL. Reclassify those sessions out of Direct into an AI-assisted channel.

  3. 3

    Persist the crawler log

    Ship the classified request log to durable storage from day one. Edge providers retain request analytics for a short window, and once it is gone the history is unavailable. The log is what lets you say which pages the agents fetched last quarter, whether fetch rate rose after a fix, and whether a citation appeared after a fetch.

  4. 4

    Join fetches to citations and sessions

    Line up three timelines per URL: when retrieval agents fetched it, when it appeared as a cited source in tracked answers, and when AI-assisted sessions landed on it. Fetches without citations mean the engine read the page and chose something else. Citations without sessions mean the answer was complete enough that nobody clicked, which is still visibility and should be reported as such.

  5. 5

    Tie recovered sessions to outcomes and report on a cadence

    Connect the AI-assisted channel to the same conversion events every other channel uses, attributed on the same model so the number is comparable, with assisted conversions reported separately from last-click. Publish it monthly next to organic, paid and referral, with the confidence tiers visible, the crawler-versus-human split stated, and fetch rate by page as the leading indicator. A report that states its limits is one leadership can act on.

What this answers

Is AI search sending us anything?

A number you can take to a budget conversation, with the confidence of each tier stated, rather than an anecdote about a customer who mentioned ChatGPT on a sales call.

Which pages do the models fetch?

Crawler demand by URL shows what the engines are actively reading about you and how often. It surfaces the pages the models care about, which are not always the pages the marketing team does.

Are our metrics bot-inflated?

Separating agents from humans often materially changes reported engagement and conversion rates, particularly on documentation, pricing and comparison pages that retrieval agents favour. The corrected figures are the ones every other channel report should have been using.

Did the citation work pay?

Connect a change in mention or citation rate on the tracked prompt set to a change in recovered sessions and revenue on the cited pages. That is the closest thing to a closed loop this channel offers, and it justifies the next quarter of content and outreach work.

Are we blocking something we meant to allow?

A retrieval agent that never appears in the log for pages you expect to be cited is a diagnosis. The cause is usually a CDN rule, a robots.txt line or a firewall signature written for scrapers, and the log finds it faster than an audit.

What differs per engine

How each engine hands off a visit, and what that means for what you can measure.

ChatGPT

Retrieval fetches arrive as OAI-SearchBot and ChatGPT-User, and human clicks from a browsing answer may or may not carry a referrer depending on the client. Expect a mix of attributed referrals and Direct-bucket visits on the same pages.

Perplexity

Perplexity fetches via PerplexityBot and Perplexity-User and shows numbered citations, and clicks from those citations are more likely to carry a referrer than most. It is often the cleanest engine to attribute and the one where fetch-to-session timing lines up most easily, which makes it the place to validate classification rules first.

Google AI Overviews / AI Mode

AI Overviews sit inside Google Search, so clicks arrive as ordinary Google referrals and are indistinguishable from classic organic in the referrer alone. There is no separate retrieval agent to watch; Googlebot serves both. Measuring this surface means pairing Search Console with citation tracking, since the traffic folds into organic.

Gemini

Gemini's grounded answers link to sources, with referrer behaviour that varies by surface and app. Clicks may arrive attributed, unattributed or through an intermediate redirect. Treat it as a mixed-confidence engine and lean on citation tracking and landing-page patterns more than on the referrer field.

The numbers to report

The figures this channel should report, and what each is safe to claim.

AI-assisted sessions
Human visits attributed to an AI surface, split by confidence tier and per engine where the referrer allows. This is the headline volume figure and the one most likely to be challenged, so the tiers need to be visible on the report itself.
Recovered share of Direct
The proportion of what was previously Direct that has been reclassified as AI-assisted. It shows how much of the unknown bucket this channel accounted for, and changes how the rest of the Direct number is read. Track it as the classification rules improve.
Retrieval fetch rate by URL
How often retrieval agents request each page, per agent. Rising fetch rate after a fix confirms the engine is reading the page; zero fetches on a page you expect cited is an access problem. This is the leading indicator and belongs on the same report as the lagging session figures.
AI-assisted conversions
Signups, demos or orders from AI-assisted sessions, attributed on the same model as other channels, with assisted conversions separate from last-click. This is the return figure. Read it alongside citation rate so answers that convinced without a click are not counted as zero.

Common mistakes and the fix

Mistake · Treating Direct as unknowable

Fix · Classify what can be classified and state the confidence. A portion of Direct is AI-assisted and recoverable with the signals that survive the hand-off. A channel with no figure has no budget.

Mistake · Counting crawler hits as sessions

Fix · Separate agents at the request level before anything reaches the analytics layer. Retrieval agents favour the same pages humans convert on, so their hits depress conversion rate exactly where it matters most.

Mistake · Discovering the retention window too late

Fix · Export the classified log to durable storage from the first day, even if nobody reads it yet. The comparison you want in month six depends on data from month one, and edge providers will not have kept it.

Mistake · Summing retrieval and training fetches

Fix · Report them separately. Retrieval fetches predict citations and precede human sessions; training fetches predict nothing about this quarter's answers. A combined "AI bot traffic" figure hides the signal in the noise.

A worked example: a D2C skincare brand

A hypothetical D2C skincare brand sees a large Direct bucket, flat organic, and a growing number of customers who mention an AI assistant in post-purchase surveys. Its analytics show nothing from any AI surface. Request-level classification at the edge shows retrieval agents fetching the ingredient and comparison pages daily, and training crawlers touching the whole catalogue on a slower cycle. Page views on the ingredient pages fall once the agents are removed, and conversion rate on those pages rises, because the visitors who remain are people.

Session recovery finds two tiers. A certain tier arrives with intact referrers or the tracking parameter one assistant appends. An inferred tier lands on ingredient pages no other channel links to, within hours of a retrieval fetch of the same URL, with no referrer. Together they account for a visible share of what was Direct, and the inferred tier is reported as inferred. Joined to orders on the same attribution model as paid social, the channel now carries a revenue figure with its uncertainty stated.

The following quarter, fetch rate becomes the leading indicator. When a rewritten comparison page starts being fetched daily by PerplexityBot, it appears as a cited source within weeks and AI-assisted sessions follow. When a category page shows no retrieval fetches at all, the log reveals a CDN rule blocking anything with "bot" in the user agent on that path. The channel report goes to leadership monthly, next to paid and organic, with the crawler split and confidence tiers on the same page.

Frequently asked questions

Why does AI traffic show up as Direct?+

Several assistants either strip the referrer header or open links from a context that carries none, such as a native app or a sandboxed view. Analytics has nothing to attribute the visit to, so it falls into Direct.

Can I just filter bots in GA4?+

GA4's known-bot filtering catches classic crawlers but not the full set of AI agents and retrieval fetchers, which change frequently and often never execute the tag. Server-side or edge-level classification is more reliable because it sees the raw request. GA4 remains the reporting layer; the classification has to happen upstream.

Why is the crawler log time-limited?+

Edge and CDN providers typically expose recent request analytics on a short retention window, often around a day at lower tiers. Any longitudinal view has to be built by exporting continuously into your own storage. If you have not started, start now; the history before today cannot be recovered.

Does this replace GA4?+

No. It corrects it. GA4 stays the system of record for sessions and conversions; this supplies the AI-channel classification and crawler separation GA4 cannot infer on its own, and joins them back so the reports you already run show the channel properly.

What is the difference between a retrieval agent and a training crawler?+

A retrieval agent fetches a page because an engine is answering a question now and may cite it: OAI-SearchBot, ChatGPT-User, PerplexityBot, Perplexity-User, ClaudeBot. A training crawler collects pages to build a corpus: GPTBot, CCBot, Google-Extended, Applebot-Extended. Retrieval fetches predict citations; training fetches do not, and the two should never be summed.

How confident can we be in recovered sessions?+

It depends on the signal. An intact referrer or a known tracking parameter is certain. A landing pattern or a fetch-then-visit sequence is inferred. Report the tiers separately; a channel figure with its uncertainty stated survives scrutiny, and one without does not.

What if the answer is complete and nobody clicks?+

Then the channel is delivering visibility without visits, which is real and should be reported alongside the citation data. Pair recovered sessions with mention and citation rate so the report reflects both the visits and the answers that made a visit unnecessary.

Should we block training crawlers to reduce noise?+

That is a content-policy decision, not a measurement one, and it has no effect on the human traffic figures once classification is in place. If you do block them, keep the retrieval agents allowed, or the visibility side of the channel disappears along with the training use.

Explore the features behind this solution

Go deeper

Find out what AI search is really sending you

Separate agents from people, recover the misfiled sessions, and put a number on the channel.

← All solutions