AI in ecommerce operations

Is AI referral traffic big enough yet to justify dedicated tracking?

Visible AI referral traffic still sits around 0.2% of sessions on the broadest public benchmark, and the undercount is wider than the channel. Here is a proportionate tracking ladder for D2C operators.

A thin lit filament of measured referral sessions inside a wider translucent volume of unattributed traffic, with the surrounding haze larger than the filament.

Key takeaways

Expect visible AI referral traffic in the region of 0.2% of sessions on the broadest cross-industry benchmark (Contentsquare, April 2026), with one ecommerce dataset implying roughly 0.08% for ChatGPT-referred sessions on published counts. Those figures are floors, not measurements: published dark-traffic estimates span about 30, 50% to 70.6% of AI sessions landing in Direct, and they do not agree. Enterprise headlines such as ChatGPT at 20% of Walmart referral clicks are shares of a referral slice that is itself under 5% of visits, not shares of sessions. A GA4 custom channel group is worth shipping now because it costs minutes and uses the single extra slot; Direct diagnostics and server-side capture should wait for a session-share or revenue trigger you write down before opening the property. Conversion quality is unresolved for D2C: Adobe and Shopify report a premium, while one 200-site ecommerce cohort finds AI converting worse than Google organic.

You have been asked whether the brand should stand up dedicated tracking for ChatGPT, Perplexity, Gemini, Claude and AI Overviews. The decision is not whether those products exist. It is whether the session share they send today is large enough, and stable enough, to justify a measurement build, and how much build is proportionate to a stream that may still be a rounding error in checkout.

The number that actually changes the decision is share of sessions, not share of referral clicks, and not a year-on-year growth rate on a tiny base. On the broadest published benchmark, AI-referred traffic represented 0.2% of total visits after 632% year-on-year growth (Contentsquare, April 2026). Contentsquare describes that sample as billions of sessions across thousands of websites; Elmo puts that sample at 99 billion sessions (S2). It is the conservative public anchor. It is cross-industry, not an ecommerce-only distribution, and it counts only visits that still arrive looking like referral.

One ecommerce dataset implies an even smaller visible share. Ethercycle published 44,737 ChatGPT-referred sessions against 57.7 million regular sessions. The dossier does not specify the date range for those counts. That is roughly 0.08% of those sessions, arithmetic on their published counts, ChatGPT-only rather than all AI engines, and not a replica of the Contentsquare study. If your GA4 property shows something in that neighbourhood, you are not behind a hidden industry median. You are looking at the same order of magnitude as the public floor.

What does getting the share wrong actually cost?

The commercial cost is not "missing AI". It is spending scarce operator and engineering time on a stream that still does not move contribution, or ignoring a stream that might, because the denominator was wrong.

Enterprise headlines make the first error easy. Similarweb data from August 2025, reported by Modern Retail in September 2025, put ChatGPT at one in five of Walmart's referral clicks, over 20% for Etsy, nearly 15% for Target and 10% for eBay. The same reporting states that referral clicks account for less than 5% of total site visits for those retailers. Twenty per cent of a slice that is itself under 5% of visits is under 1% of sessions. Treat those figures as session share and you will over-build tracking, brief generative-engine work as if it were a media channel, and argue about budget against a number your checkout will never feel.

The opposite error is cheaper this quarter and expensive if the channel later starts to matter. A custom channel group occupies the single extra slot GA4 gives you. Server-side referrer capture has an unpriced engineering cost in the public record, no source in this research prices it, but it is not free, and hostname lists decay. Time spent here is time not spent reading blended media efficiency against contribution, or checking whether a Meta spend increase would magnify a system that already works.

There is a third cost: quality confusion. If AI sessions convert differently from Google organic, a sub-1% share can still change merchandising, feed or content priorities. If they do not, the same share is noise. The evidence on that point does not currently agree, which is itself a reason not to spend as if a quality premium were settled.

Why the usual answers do not settle the decision

Most of what you will be handed does one of four things, and none of them answers whether dedicated tracking is worth it now.

The first is a growth-rate headline without the base. 632% growth of a 0.2% share is still 0.2% of visits. Shopify's Q1 2026 platform analysis, cited by Ethercycle, reports AI referral sessions growing 8x year on year. Growth on a rounding-error base can be real and still too small to reallocate paid search, paid social or engineering.

The second is a ten-minute regex tutorial presented as strategy. Those guides are useful for the cheapest rung of tracking. They are not a cost-benefit case. Several of them also assert that GA4 custom channel groups reclassify historical sessions; at least one states that existing sessions will not reclassify. That conflict is unresolved in the public material. Do not plan a backfill on a blog claim. Test it on a live property if historical comparison matters to you.

The third is a single dark-traffic multiplier applied to Direct. You will see estimates that 30 to 50% of AI-driven sessions land in Direct (S9), and you will see 70.6% recirculated across posts as if it were independently corroborated. The 70.6% figure traces to Loamly's analysis of 446,405 visits, not to a family of studies. Engine-level tests disagree with any blended multiplier. A small Retailgentic test identified Gemini iOS app traffic as AI referral in only 5 of 56 visits, with 91% landing in Direct; Perplexity stripped least, with roughly 30% landing as dark traffic; ChatGPT was mixed. A 56-visit test is illustrative, not a rate. Its value is that undercount varies by engine, which breaks a single multiplier.

The fourth is treating Google Search Console's Generative AI performance report as a traffic proxy. At launch the report omits clicks, CTR, triggering queries, position within the answer, citation placement, and any conversion or revenue attribution. Google has said click data is coming "over time", with no confirmed date in the material we have. It is a reach signal. It cannot tell you how many sessions AI Overviews sent, and it cannot justify a tracking build on its own.

Is the published share a measurement, or a floor?

Read every AI referral percentage as a floor. GA4 only sees what arrives with a usable referrer. Three mechanisms strip that signal: in-app WebView and WKWebView browsers that do not pass the Referer header; HTTP redirect chains; and engines that suppress referrers. Pasting a URL into a browser produces the same Direct/(none) bucket.

That is why Contentsquare's 0.2% and the Ethercycle 0.08% arithmetic are not "the size of AI traffic". They are the size of AI traffic that still looks like referral in analytics. The undercount estimates then refuse to collapse into one number. One account puts 30 to 50% of AI-driven sessions in Direct. Another puts 70.6% on the Loamly visit set. Attrifast, on a vendor-run cohort of 200 Stripe-connected SMB sites through May 2026, reports a median 34% of GA4 Direct/(none) traffic as AI-referred once server-side fingerprinting is applied, and 41% for B2B SaaS. That method is not independently verified. Treat it as one cohort's finding.

That pairing is the diagnosis. A 0.2% visible share with a 30 to 70% undercount still leaves AI as a low-single-digit percentage of sessions at most, and quite possibly still well under 1% after correction. You cannot honestly say it is nothing. You cannot honestly say it is already a media channel. You can say the measurement problem is larger than the measured object, which is a reason to buy cheap instrumentation first and expensive instrumentation only when a trigger you wrote in advance actually fires.

This is the same discipline you already use elsewhere: AI can widen retrieval without replacing judgement, and ROAS can steer delivery while profit decides whether growth is worth buying. Tracking spend should be judged by the ecommerce work it is meant to improve, a cleaner source report, a decision about content or feed quality, a decision not to reallocate paid budget, not by whether a vendor chart is up and to the right.

What tracking is proportionate at each rung?

Penang Media's operating position, as a method rather than as a claim about your numbers, is a three-tier ladder. The figures above are cited. The ladder is how we would gate effort against them. Write your triggers down before you look at the property, or you will move the goalposts.

Start with a GA4 custom channel group, and do that now.

GA4 allows a maximum of two custom channel groups per property, one of which is the default, leaving a single extra slot. Setup is on the order of five minutes per new property once a documented regex exists. The AI channel must sit above Referral in the group order, or GA4 assigns matching sessions to Referral first. Review the regex quarterly; engine hostnames move. That is cheap enough to be unconditional. It does not recover dark traffic. It does give you a named channel for the sessions that still arrive with a referrer, which is the floor this whole argument depends on.

Use the slot deliberately. You only get one extra group. If that slot is already doing something that changes weekly decisions, do not casually overwrite it for a 0.2% stream. If it is idle, fill it.

Next comes a Direct diagnostic rather than a multiplier, and it should be gated.

Seresa describes a practical comparison: engagement rate, pages per session and duration on Direct against organic. If Direct matches or exceeds organic engagement, a meaningful share of Direct may be AI-referred rather than typed-in brand traffic. That is a diagnostic, not a reclassification. Do not apply 70.6% or 34% as a constant. Attrifast's 34% is a median from a fingerprinted SMB cohort; Loamly's 70.6% is one visit set. Your mix of ChatGPT, Gemini, Perplexity and in-app browsers will not match either sample.

A reasonable trigger, which you should rewrite in your own units: only run this diagnostic when Tier 1 AI referral plus a conservative uplift, for example, assuming the low end of the 30 to 50% undercount range, would still be large enough to change a content, feed or merchandising decision. If even the generous case cannot change a decision, skip it.

Seresa also claims a roughly 30-minute custom GA4 channel setup recovers 50 to 70% of misattribution, with server-side needed for the rest. Those recovery percentages are unsourced in the material we have. Treat the minutes as plausible; do not treat the recovery range as a measured yield.

Server-side capture sits last, gated harder still.

This is where you persist referrers before the browser strips them, or fingerprint sessions the way the Attrifast cohort describes. No public source here prices the engineering. No public source tests whether standing up a tracked AI channel later changed a budget decision. Those two gaps are the point. Do not buy an unpriced build to instrument a channel whose visible share is 0.2% and whose conversion quality is disputed.

Until a new classification actually changes an action, server-side capture is an unpriced build. A channel that does not change a decision is just another report.

I do not have a published figure for hours per quarter to keep server-side referrer capture alive, and I will not invent one. In practice it does not get a standing job without a named owner and a maintenance estimate. If it is more than a light tag job, it waits until the channel is moving money.

A reasonable trigger: Tier 1 visible share crosses a session or revenue line you would actually reallocate paid search, paid social or content time to serve, and you have already decided what you would do with a trustworthy number. If you cannot name the decision, you do not need the pipeline.

  1. Tier 1: GA4 custom channel group, minutes of work, do it now unless the single extra slot is already earning its keep.
  2. Tier 2: compare Direct engagement with organic; do not apply a blended dark-traffic multiplier; gate on a written session-share trigger.
  3. Tier 3: server-side referrer capture, gated on a decision you can name and a visible-share line you wrote down in advance.

Does AI traffic convert better for ecommerce?

This is the spine of the tracking decision, and the public evidence does not resolve it for D2C brands.

Adobe Digital Insights, reported for Q1 2026, said that in March 2026 AI referral traffic converted 42% better than non-AI traffic across more than a trillion visits from 130-plus North American retailers, reversing from 38% worse a year earlier. A secondary compilation reports Adobe finding AI-referred US retail visitors converting 54% better than non-AI in May 2026. The premium is a moving figure, vendor-stated, and drawn from large North American retailers rather than D2C brands of typical operator size. Contentsquare's early behavioural data, also vendor-stated and directional, says AI-referred visitors bounce less and convert more than average, while a 2025 Contentsquare cut put AI-referred conversion at 1.3%, below email at 1.9%.

Shopify's Q1 2026 platform analysis, cited by Ethercycle, reports AI referral sessions converting roughly 50% better than organic search. Ethercycle's own first-party finding is that AI beat Google by 46% on orders-per-session (2.56% versus 1.75%), on those 44,737 ChatGPT-referred sessions versus 57.7 million regular sessions. Ethercycle also reports that two-thirds of those AI-referred visitors landed directly on a product page. That is that source's finding, not a house measurement.

Set that against Attrifast's 200 Stripe-connected SMB sites through May 2026: on ecommerce, Google organic converted at 2.1% versus AI at 1.6%, AI underperformed. On B2B SaaS the pattern reversed (2.7% versus 1.4%). That ecommerce/B2B split is the most useful published contradiction in this set. The study is vendor-run; the "server-side fingerprinting" method is not independently verified. It is still the only cohort here that isolates ecommerce conversion against Google organic and finds AI worse.

Useful for

  • Adobe reports AI-referred retail visitors converting better than non-AI traffic in 2026, with the published premium moving between cuts.
  • Shopify platform analysis and Ethercycle first-party data report AI converting better than organic search on orders-per-session.

Watch for

  • A 200-site Stripe-connected cohort finds AI converting worse than Google organic on ecommerce specifically.
  • Samples differ in size, sector and method, so the evidence does not settle conversion quality for D2C brands.

Plausible reconciliations, sample mix, B2B versus D2C, relative premiums versus absolute rates, ChatGPT-only versus all-engine, are inference, not evidence. Elmo's compilation makes the methodological point that Adobe reports a relative premium while Contentsquare reports an absolute rate, so the two are not directly comparable. Nobody in the dossier resolves the contradiction. Until someone does, do not build Tier 3 on the assumption that AI sessions are premium demand. Do not skip Tier 1 on the assumption they are junk. Instrument the floor, then judge quality on your own orders, contribution and repeat rate.

That last step is how we operate, not a ranking we borrowed. Last-90-day AI-referred conversion has not been compared to Google organic by engine on our own orders. The public sets contradict each other, so neither is the rule. Until there is an order-level read with a surviving referrer, AI-referred traffic stays a research path.

Engine mix is not a stable denominator either, which is why maintenance cost, not build cost, is the real objection to anything above Tier 1. Attrifast's cohort put ChatGPT at 71% of AI sessions, Gemini 12%, Perplexity 8%, Claude 6% and AI Overviews 3%. Lantern data, cited by Retailgentic, put ChatGPT at 87.4% of AI referral traffic. Goodie's Wave 1 (May, August 2025, 2,802,519 AI referral sessions across 41 brand sites) put ChatGPT at 89.1%; Wave 2 (March, April 2026, brand-averaged and explicitly labelled B2B) put ChatGPT at 62.6%, Claude 18.5%, Gemini 10.6% and Perplexity 7.3%. Do not average those figures. They are different samples. Wave 2 is B2B and must not be presented as an ecommerce mix. Together they only support one operational claim: the source list you write in Q2 will be wrong in Q4 if you do not revisit it.

That is also why a tracking spec expires. Guides recommend quarterly regex review because hostnames move. We do not have a sourced miss rate for a stale list, only the recommendation. If in-chat purchase surfaces keep appearing and retreating, the instrument you need remains the on-site session, not a closed-loop checkout pixel you do not control. Session-level tracking is the right object for now. It is not a reason to over-build it.

If AI-referred visitors do land on product pages more often, the on-site problem they create is still a merchandising and lifecycle problem, not a tracking-stack problem. Fix the welcome or post-purchase flow on its own evidence. Do not wait for an AI channel report to tell you whether the first order experience works.

What should you commit to before you look at the data?

Open the property. If the extra custom channel group slot is free, create an AI referral group, place it above Referral, and save the regex next to your UTM conventions. Put a quarterly review in the calendar. That work is minutes.

On the same page, write the session-share or revenue line at which Direct would be worth comparing with organic, and the line at which server-side capture would be worth scoping. Use the units that would actually move you. One working example: do not change a paid, content or merchandising decision on session-share of AI traffic. The line that would move the decision is contribution, or enough volume to move blended MER. There has not been enough volume to set a dedicated threshold, so extra tracking waits until that line is in sight. If you cannot name a decision those lines would change, stop. The cheap channel is enough until one of those lines is crossed.

Then go back to contribution, stock and creative supply, the constraints that already decide whether the next pound of acquisition is worth buying.

Useful answers

Questions operators ask

What percentage of sessions should we expect from AI referral traffic?
Treat published figures as a range and a floor, not a point estimate for your brand. The broadest cross-industry benchmark sits at 0.2% of total visits (Contentsquare, April 2026). One ecommerce dataset implies roughly 0.08% when you divide 44,737 ChatGPT-referred sessions by 57.7 million regular sessions (Ethercycle's published counts, July 2025, June 2026). Those are visible referral shares only. Dark-traffic estimates then span about 30, 50% to 70.6% of AI sessions landing in Direct, so a corrected share could be higher, and the sources do not agree on the uplift. Do not extrapolate a figure for a specific D2C brand from these samples.
Why do enterprise retailers report ChatGPT at 20% when my share is under 1%?
The denominator is different. Similarweb data from August 2025, reported in September 2025, put ChatGPT at one in five of Walmart's referral clicks (and high teens or tens of per cent of referral clicks at Etsy, Target and eBay). The same reporting states that referral clicks account for less than 5% of total site visits for those retailers. Twenty per cent of a slice under 5% of visits is under 1% of sessions. Those headlines are shares of referral, not shares of sessions, and the cut is already dated.
How much AI traffic is hidden in my Direct bucket?
The public range is unresolved. One account estimates 30, 50% of AI-driven sessions land in Direct; another puts 70.6% on Loamly's analysis of 446,405 visits, a figure that is widely recirculated from that single study. A vendor cohort of 200 SMB sites reports a median 34% of GA4 Direct/(none) as AI-referred after server-side fingerprinting, which is not independently verified. Undercount also varies by engine: a small Gemini iOS test put 91% of visits in Direct, while Perplexity stripped least. Do not apply a blended multiplier. Compare Direct engagement, pages per session and duration with organic as a diagnostic instead.
Is it worth building dedicated AI traffic tracking now, or should we wait?
A GA4 custom channel group is worth doing now if the single extra slot is free: setup is on the order of minutes, the channel must sit above Referral, and it only captures sessions that still arrive with a referrer. Direct diagnostics and server-side capture should wait for a session-share or revenue trigger you write down before looking at the data. That ladder is Penang Media's operating position on effort, not a claim that your property will match any published share. No public source here prices server-side engineering or shows that a tracked AI channel later changed a budget decision.
Does AI traffic actually convert better for ecommerce?
The evidence does not settle it for D2C brands. Adobe has reported AI referral converting 42% better than non-AI traffic in March 2026 (and a secondary source reports 54% better in May 2026) among large North American retailers. Shopify's Q1 2026 analysis and Ethercycle's first-party cut report AI converting better than organic search on orders-per-session. A 200-site Stripe-connected cohort finds the opposite on ecommerce: Google organic at 2.1% versus AI at 1.6%, with the pattern reversing on B2B SaaS. Populations and methods differ. Instrument the floor, then read your own orders.
Can Google Search Console tell me how much traffic AI Overviews sends?
Not at present. The Generative AI performance report shows reach without clicks, CTR, triggering queries, position within the answer, citation placement, or conversion and revenue attribution. Google has said click data is coming over time, with no confirmed date in the material we have. Treat it as a reach signal, not a traffic or performance proxy, and verify current rollout status in your own Search Console property.

About the author

Eddie Cheng

Eddie Cheng founded Penang Media and co-owns VIBAe. He writes from the agency and brand sides of ecommerce growth, connecting paid acquisition with stock, margins, cash flow and contribution profit.

More from Eddie Cheng

The operating context

Growth from the agency and brand sides.

Eddie Cheng writes about profit-first ecommerce growth from both sides of the work: Penang Media, the performance agency he founded, and VIBAe, the footwear brand he co-owns. His articles connect paid acquisition with stock, margins, cash flow and the decisions that determine profitable growth.

About the publication