AI SEO for Ecommerce: How Online Stores Get Found in AI Search
Ishant
Published : September 25, 2026 at 2:34 pm
Updated : September 25, 2026 at 4:28 pm
Ishant
Ishant Sharma is the Founder and CEO of Hustle Marketers, a Google Partner digital marketing agency. With 12+ years of experience in Google Ads, Meta Ads, SEO, and e-commerce PPC, he has helped 2,500+ brands generate $780M+ in trackable sales. Upwork Top Rated Plus with 100% Job Success Score. Ishant Sharma is the digital marketing specialist, not the Indian cricketer of the same name.
Somebody has told you that AI search is going to kill your organic traffic, and that the fix is a new service with a new acronym. Somebody else has told you to add llms.txt and write in an authoritative tone.
Here is what the primary documentation actually says. Google publishes that there are no additional requirements to appear in AI Overviews or AI Mode and no special structured data you need to add. The paper that invented the term generative engine optimization found that keyword stuffing scored worse than doing nothing, and that writing more authoritatively produced no significant improvement. The generic block-the-AI-bots robots.txt snippet that circulates also deletes your store from ChatGPT search answers, because training crawlers and search crawlers are different bots and most advice conflates them.
This post is built from vendor documentation and fetched research papers only. Google Search Central, Chrome’s Lighthouse documentation on agentic browsing, OpenAI’s bot and commerce documentation, Anthropic’s crawler article, Perplexity’s bot guide, the Bing Webmaster blog, the llms.txt specification, and three research studies read in full rather than summarised from someone else’s blog post. Where a number could not be traced to a source, it is not here, and there is a list at the end of what was left out and why.
Written by Ishant Sharma, founder of Hustle Marketers, working in Google Ads, Microsoft Ads and ecommerce SEO since 2013. Checked September 2026.
Key takeaways
- Google states there are no additional requirements to appear in AI Overviews or AI Mode, and no special schema.org structured data you need to add. The eligibility condition is that the page is indexed and eligible to be shown with a snippet.
- Blocking GPTBot means OpenAI says your content should not be used to train its foundation models. Blocking OAI-SearchBot removes you from ChatGPT search answers, in OpenAI’s own words. They are different lines in the same file and the generic snippets that circulate usually include both.
- Google-Extended does nothing to AI Overviews or AI Mode. Google states it does not impact inclusion in Search and is not a ranking signal. Its scope is Gemini Apps and Vertex grounding.
- The real AI kill switch is nosnippet and max-snippet. Google states nosnippet prevents content being used as a direct input for AI Overviews and AI Mode, and that max-snippet limits how much may be used. A restrictive max-snippet set years ago is an active throttle today.
- Pew’s metered study of 900 US adults and 68,879 searches found users clicked a source cited inside an AI summary on 1 percent of visits to search pages that had one.
- One and two word searches produced an AI summary only 8 percent of the time. Question-shaped searches produced one 60 percent of the time. Your catalogue pages are far less exposed than your buying guides.
- The famous 40 percent GEO figure is a relative gain on a text-share metric the authors invented, on their own harness wrapping gpt-3.5-turbo, submitted November 2023. On the one commercial engine tested the gain was 9 percent on one of two metrics.
- ChatGPT product results come from a gated feed pipeline, and OpenAI’s published ranking criteria are availability, price, quality and whether you are the maker or primary seller. None of those is content.
- Bing Webmaster Tools reports AI citations at page level. Google merges AI Overviews and AI Mode into the Web search type with no documented way to separate them.
Table of Contents
- What is actually different about AI search for an online store?
- What does Google say you have to do to appear in AI Overviews and AI Mode?
- Which robots.txt lines remove you from AI answers, and which only stop training?
- Are you already blocking yourself from AI answers without knowing?
- Does llms.txt actually do anything?
- What does the research actually show about optimising for AI answers?
- How much traffic does an AI citation actually send?
- Which pages on a store are actually candidates for AI citation?
- How do you get a product into ChatGPT’s shopping results?
- Why does Bing matter more here than its market share suggests?
- How do you measure any of this?
- What should an ecommerce store actually do about AI search?
- How should this change your SEO and PPC budget?
- Why work with Hustle Marketers on AI search for ecommerce?
- AI SEO for ecommerce FAQs
What is actually different about AI search for an online store?
Less than the marketing suggests on the Google side, and more than you would expect on the assistant side. The two need separating before anything else makes sense.
What are the surfaces, and do they work the same way?
Four things get called AI search and they behave differently.
AI Overviews sit at the top of an ordinary Google results page. Google describes them as an AI-generated snapshot with key information and links to dig deeper. The results page underneath is still there.
AI Mode is a separate Google experience. Google calls it its most powerful AI search experience, where a user can ask anything and go deeper through follow-up questions.
ChatGPT search is OpenAI’s own product. It rewrites your question into one or more targeted queries and sends them to search partners. OpenAI names Microsoft and Shopify among those partners.
Assistants that browse, which covers Claude, Perplexity and ChatGPT when a user asks about a specific page, fetch pages live in response to a user request rather than from an index they built earlier.
The reason this matters is that the controls are different for each one, and a single robots.txt edit can affect one and not the others.
Is AI Overviews the same as AI Mode?
No, and you should stop reading anything that treats them as one thing.
They are different surfaces. AI Mode is a separate conversational experience where you can ask follow-up questions. AI Overviews sit above a conventional results page that is still there underneath. Google describes query fan-out, dividing a question into subtopics and searching each one simultaneously, across its AI features rather than as the thing that separates the two.
They do share their eligibility rules and their measurement treatment in Search Console, which is covered below. But an article that says “AI Overviews, also known as AI Mode” has not read Google’s documentation.
Where do ChatGPT’s product results come from?
Not from your page copy. This is the single most important thing in this post for a store owner.
OpenAI runs a separate commerce pipeline for product results, and it is a feed pipeline. It is documented as currently available to approved partners and works from daily snapshots. Shopify is named as a ChatGPT search partner.
OpenAI also publishes the criteria it uses to rank merchants: availability, price, quality, and whether they are the maker or primary seller. It states separately that product results are selected independently and are not ads.
Read that list again. Availability. Price. Whether you made the thing. None of those is content optimisation, and no amount of rewriting your product description changes any of them. That section has its own treatment further down.
Which parts of this can you influence?
Three buckets, and being honest about which is which is most of the value here.
Things you control outright: whether the right crawlers can reach you, whether your snippet directives are throttling you, whether your content is in text rather than locked in images or tabs, whether your feed is accurate.
Things you influence: whether your pages are the ones worth citing for the questions your customers actually ask. That is ordinary content work with a different target.
Things that are structural: whether you are the manufacturer, what your price is, whether the query your customer typed even triggers an AI surface. You can change the first two as a business decision. You cannot change them with marketing.
What does Google say you have to do to appear in AI Overviews and AI Mode?
Nothing new. That is not a simplification, it is close to a quotation.
Is there anything extra to do?
Google’s AI features documentation states that there are “no additional requirements to appear in AI Overviews or AI Mode, nor other special optimizations necessary.”
The best practices it does list are these seven, and not one of them is AI-specific: allow crawling in robots.txt; make content findable through internal links; provide a good page experience; make sure important content is available in textual form; support text with high-quality images and video; make sure structured data matches the visible text; and keep Merchant Center and Business Profile information up to date.
That last one is the only one with an ecommerce edge to it, and it is a feed and profile instruction rather than a content one.
Do I need special schema for AI search?
No. Google’s wording on the same page is that there is “no special schema.org structured data that you need to add.”
This is worth being firm about, because a lot of ecommerce AI search advice is built on the opposite claim. Structured data is still worth doing, for the reasons it was always worth doing: Product, merchant listing, product variants, review snippet, return policy, shipping policy, breadcrumb and Organization markup all drive documented search appearances in Google’s structured data gallery. Google’s own instruction on the AI features page is that your structured data should match the visible text. What it does not say, anywhere, is that adding schema gets you into AI Overviews.
If someone is selling you schema work as an AI Overviews lever, ask them for the vendor statement. Google’s says the opposite. The one statement pointing the other way is Microsoft’s, which names JSON-LD among the things it says help content get into Bing and Copilot answers. That is a different engine, a formatting recommendation, and not a measured effect.
What is the actual eligibility condition?
One sentence, and it is the whole gate. To be shown as a supporting link in either surface, a page “must be indexed and eligible to be shown in Google Search with a snippet.”
Two conditions. Indexed, and snippet-eligible. Fail either and you are out, and the second one catches stores far more often than anyone expects. That is the next section.
What is query fan-out, and what does it mean for a store?
AI Mode breaks a question into subtopics and searches each one at the same time across multiple data sources, then combines the results. Google’s framing of what that produces is a wider and more diverse set of links than a single query would return.
The practical reading for a store is that you do not have to rank for the literal question a customer typed. If someone asks a broad question about choosing between two product types, the fan-out generates subquestions about materials, sizing, care, cost and use cases, and a page that answers one of those subquestions well is a candidate even though it would never have ranked for the head query.
That is an argument for depth on the specific questions your customers ask, not for another page targeting the head term.
When does AI Mode actually include a link?
Google publishes one line about this that nobody in ecommerce seems to have read. AI Mode is trained to decide when to include hyperlinks if it is likely the user may want to take action or finish a task on a website.
Transactional intent is a documented trigger for including a link. For a store, that is the most useful sentence Google has published about AI Mode, because it says the queries most likely to produce a clickable link are the ones where somebody wants to do something rather than know something.
Google also states that AI Mode will provide a set of web links where there is not high enough confidence in the quality or helpfulness of an AI response. Low confidence is good for you.
Which robots.txt lines remove you from AI answers, and which only stop training?
This is where most stores get hurt, and the damage is silent. Every assistant vendor here runs at least two different bots with completely different jobs, and the advice circulating online treats them as one.
What does each bot actually do?
| Bot | Vendor | What it is for | What blocking it does |
|---|---|---|---|
| GPTBot | OpenAI | Training foundation models | OpenAI says content should not be used for training |
| OAI-SearchBot | OpenAI | Surfacing sites in ChatGPT search | Removes you from ChatGPT search answers |
| ChatGPT-User | OpenAI | Live fetch on a user action | Not an automatic crawler, robots rules may not apply |
| ClaudeBot | Anthropic | Collecting content for models | Signals future material should be excluded from training |
| Claude-SearchBot | Anthropic | Improving search result quality | Stops your content being indexed for search |
| Claude-User | Anthropic | Fetching on a user question | Stops retrieval in response to a user query |
| PerplexityBot | Perplexity | Surfacing and linking sites in results | Stops you appearing in Perplexity results |
| Perplexity-User | Perplexity | Live fetch on a user action | Generally ignores robots.txt |
| Google-Extended | Gemini training and Vertex grounding | Nothing to Search, AI Overviews or AI Mode | |
| Googlebot | Indexing for Google Search | Removes you from Search, and therefore from AI features | |
| CCBot | Common Crawl | Building a public research corpus | Corpus inclusion only, no assistant checked here answers from it |
Every purpose column above is the vendor’s own wording. Four of the consequence cells are not, and you should know which. Perplexity frames allowing the bot as what gets you surfaced rather than stating what blocking costs you. Google states only that Google-Extended does not affect Search inclusion or ranking, and never names AI Overviews or AI Mode at all. Google does not spell out the Googlebot consequence in those words, it follows from the eligibility rule. And the Common Crawl line is bounded to the assistants checked here. Those four are my reading of the documentation rather than quotations from it.
The wording on two of the others is worth quoting because it is unusually direct.
OpenAI on OAI-SearchBot: “Sites that are opted out of OAI-SearchBot will not be shown in ChatGPT search answers.”
Perplexity on PerplexityBot: it is “designed to surface and link websites in search results on Perplexity. It is not used to crawl content for AI foundation models.”
Read the second one twice. PerplexityBot is not a training crawler at all, so blocking it to stop training achieves nothing. What Perplexity does not say is what happens if you block it. It says allowing it is what gets your site surfaced and linked in Perplexity results, so the reasonable reading is that blocking costs you that. But Perplexity has not stated it the way OpenAI has about OAI-SearchBot, and you should know which of those two is a quote and which is a reading.
Why does the standard block-the-AI-bots snippet delete you from ChatGPT?
Because it lists every AI-sounding user agent in one block, and half of them are the search bots.
The snippets that circulate typically disallow GPTBot, OAI-SearchBot, ClaudeBot and PerplexityBot together, and often CCBot and Claude-User as well. The intent is almost always “do not train on my content”. What it actually does is opt out of training, and simultaneously opt out of ChatGPT search answers, very probably Perplexity results, and Claude retrieving your page when somebody asks about it.
If your position is that you do not want to be used for training but you do want to be found, the split is:
Disallow for training: GPTBot, ClaudeBot, CCBot, and Google-Extended if you want Gemini training excluded.
Allow for visibility: OAI-SearchBot, Claude-SearchBot, Claude-User, PerplexityBot, Googlebot.
That is a business decision, not a technical one, and it is worth making deliberately rather than inheriting from a snippet someone posted on a forum.
Does blocking Google-Extended remove me from AI Overviews?
No, and this is the most consequential single misconception in the subject.
Google’s documentation states that “Google-Extended does not impact a site’s inclusion in Google Search nor is it used as a ranking signal in Google Search.” The scope Google names for it is training future Gemini models that power Gemini Apps, and grounding in Gemini Apps and Grounding with Google Search on Vertex AI.
AI Overviews is not in that list. AI Mode is not in that list.
So a store owner who blocks Google-Extended expecting to disappear from AI Overviews has changed nothing about AI Overviews. And a store owner who allows Google-Extended hoping to improve their AI Overviews visibility has also changed nothing, because Google says it is not a ranking signal.
One more detail that catches people auditing their logs. Google-Extended has no user agent string of its own. Google documents that crawling is done with existing Google user agent strings and the robots.txt token is used in a control capacity only. You will never see Google-Extended in a server log, because it does not exist as a fetcher.
What about Common Crawl?
Blocking CCBot is a training and corpus decision, not a visibility decision. No assistant in this research answers live user queries out of Common Crawl. Blocking it removes you from a widely reused research corpus. It does not remove you from anybody’s answers.
How do I check whether a bot is really who it says it is?
Every vendor here publishes a machine-readable IP list, and this is the most underused thing in the whole subject.
OpenAI publishes separate JSON files for GPTBot, OAI-SearchBot, ChatGPT-User and its ads bot. Anthropic publishes one for all three of its bots. Perplexity publishes one each for PerplexityBot and Perplexity-User. Common Crawl publishes one for CCBot. Google documents its user agent strings in full.
That means you can establish from your own server logs exactly which AI systems fetched which pages and how often, and verify that the request really came from that vendor rather than from something spoofing the user agent. More on why that matters in the measurement section.
Are you already blocking yourself from AI answers without knowing?
Quite possibly, and not through robots.txt. Through a snippet directive somebody set years ago for a completely different reason.
What does nosnippet actually do now?
Google’s robots meta tag documentation defines nosnippet as not showing a text snippet or video preview in the search results for the page. Then it adds a sentence that changes everything: it “will also prevent the content from being used as a direct input for AI Overviews and AI Mode.”
Combine that with the eligibility rule from earlier. A page must be indexed and eligible to be shown with a snippet. nosnippet makes it ineligible. The page is therefore out of AI Overviews and AI Mode entirely.
Is max-snippet throttling me?
It might be, and it is the version of this problem that hides best, because a max-snippet value does not look like a block.
Google states that max-snippet applies to “all forms of search results (such as Google web search, Google Images, Discover, Assistant, AI Overviews, AI Mode)” and that it “will also limit how much of the content may be used as a direct input for AI Overviews and AI Mode.”
A value of 0 is equivalent to nosnippet. A value of -1 lets Google choose. Anything in between is a cap.
Stores set restrictive max-snippet values for reasonable reasons that have nothing to do with AI: a publisher agreement, a European press rule, an old recommendation about controlling how much product copy appeared in results, or an SEO plugin default that was never revisited. Whatever the reason, if there is a low max-snippet on your templates today, you are capping how much of your content can feed an AI answer.
Go and look. This takes two minutes and I have found it on real stores more than once.
What is data-nosnippet for?
It is an HTML attribute rather than a meta directive, and it lets you exclude parts of a page rather than the whole thing. It works on span, div and section elements and takes no value.
Google documents several constraints that trip people up: any attribute value is ignored, the HTML must be valid with properly closed tags, unclosed elements affect everything after them, custom elements have to be wrapped in a div, span or section, and you should include the attribute when the DOM element is first created rather than adding or removing it with JavaScript.
Bing added support for data-nosnippet in October 2025 and is explicit about the scope: marked sections do not appear in Bing search snippets or AI-generated answers, while the page stays indexed and eligible to rank. Bing describes that as covering Bing Search and Copilot experiences powered by Bing.
So data-nosnippet is now a cross-vendor way to say “index this page, use it, but do not quote this bit”. For a store, the obvious candidates are internal notes, supplier terms, and anything in a template that you do not want appearing as an answer.
What is the robots.txt trap?
Blocking a URL in robots.txt does not apply the noindex on that page.
Google documents it plainly: if a URL is disallowed in robots.txt, the meta tag rules on that page are never discovered and are therefore ignored. The crawler cannot read the instruction because it is not allowed to fetch the page.
This is an old SEO trap and it is now also an AI trap, because the whole snippet directive family lives in the same meta tag.
What audit should I run?
Six checks, and you can do them in ten minutes.
Fetch your robots.txt and list every user agent block. For each one, decide from the table above whether you meant to stop training or stop visibility.
Check whether OAI-SearchBot, Claude-SearchBot, Claude-User or PerplexityBot are disallowed, and whether that was deliberate.
Check the robots meta tag on your product template, your collection template and your top buying guide for nosnippet or a max-snippet value.
Check the same via the X-Robots-Tag response header, because it can be set at server level and never appear in your HTML or your SEO plugin.
Check whether anything important is disallowed in robots.txt and also carries a noindex, which means the noindex is not being seen.
Check your SEO plugin’s global defaults, not just the per-page overrides, because that is usually where an old max-snippet lives.
Does llms.txt actually do anything?
Probably not yet, and the honest answer is more useful than either of the confident ones.
What is the specification?
llms.txt is a proposed convention for a markdown file at the root of a domain that gives large language models a curated, readable map of a site’s most important content. It is a community proposal with a published specification, not a standard published by any search engine or assistant vendor.
Has any assistant said it reads the file?
Not in anything I could find. This research checked twenty primary pages across Google, OpenAI, Anthropic, Microsoft and Perplexity, including each vendor’s main documentation on crawling, on appearing in AI features, and on making content discoverable for AI search. None of them contains a statement that its assistant or search engine reads llms.txt when answering a user query.
Google goes slightly further than silence. Its AI features page states that “You don’t need to create new machine readable files, AI text files, or markup to appear in these features.” That is not a statement that llms.txt is ignored, but it is Google telling you not to create a file like it in order to appear in AI Overviews or AI Mode.
That is a negative finding from a bounded search. It means those pages were checked and none said it. It does not prove nobody reads it. It does mean the statement is absent from exactly the documentation where it would most naturally live.
The strongest thing that does exist is a Lighthouse audit in Chrome’s documentation, in a section about agentic browsing, which calls llms.txt an emerging convention, marks its absence as not applicable, and describes providing the file as optional at the moment. You can reasonably read that as Google taking the convention seriously enough to audit whether the file is retrievable. You cannot reasonably read it as Google saying anything consumes it.
Should I publish one anyway?
If it costs you an hour, sure. It is cheap, several mainstream platforms and plugins support generating one, and there is no documented downside.
What you should not do is buy it as an AI visibility service, or let it displace the things in this post that do have vendor documentation behind them. If somebody is selling you llms.txt as the lever, ask them which assistant has said it reads the file. On this evidence they will not have an answer.
What does the research actually show about optimising for AI answers?
Less than the industry built on top of it. The foundational paper is worth reading properly, because almost every guide quotes one number from it and none of them quote the caveats.
Where does the 40 percent figure come from?
From a paper called GEO: Generative Engine Optimization, arXiv 2311.09735, submitted in November 2023 by Aggarwal, Murahari, Rajpurohit, Kalyan, Narasimhan and Deshpande, revised to version three in June 2024 and accepted to KDD 2024, which is the version read here. It is the paper that coined the term. Its abstract claims GEO “can boost visibility by up to 40%” in generative engine responses.
That sentence has been repeated into a market. Here is what is underneath it.
What did the paper actually measure?
Not traffic. Not clicks. Not sales.
The headline metric is Position-Adjusted Word Count, which measures how much textual real estate a source occupies inside a generated answer, weighted by where it appears. The other main metric is a subjective impression score assessed by GPT-3.5 across seven dimensions, one of which is the likelihood of a user clicking, scored by the model rather than observed from any user.
The 41 percent improvement the paper reports is a relative gain on Position-Adjusted Word Count, from a baseline of 19.3 to 27.2 for the best method, measured on the authors’ own research harness wrapping gpt-3.5-turbo with Google’s top five results, on a 1,000-query test set that is around 80 percent informational.
What happened on a real commercial engine?
This is the part nobody quotes. The paper also tested on Perplexity, the only deployed commercial engine in the study, on a 200-sample subset. There, the paper reports Statistics Addition improving by up to 9 percent and 37 percent on the two metrics.
Nine percent on one of the two metrics, on a real engine, on 200 samples. That is what the headline reduces to when it leaves the authors’ own test harness.
Which tactics failed?
Two of the most commonly recommended ones, and the paper says so directly.
Keyword stuffing was the worst performer and produced a negative result: 17.7 against the 19.3 baseline, so it made things worse. The paper states that simple methods traditionally used in SEO do not perform well and that keyword stuffing offers little to no improvement.
Writing in a more authoritative tone produced no significant improvement, and the authors suggest generative engines are already somewhat robust to that kind of change.
The methods that did work were adding quotations from credible sources, adding quantitative statistics, and citing sources. Those are the three at the top of the table, and they are all versions of the same thing: put verifiable specifics on the page.
What did the paper explicitly not test?
The authors state it themselves: “Owing to the black-box nature of search engine algorithms, we didn’t evaluate how GEO methods affect search rankings.”
Beyond that, and by simple date arithmetic: Google AI Overviews, Google AI Mode, ChatGPT, Claude, Copilot and Gemini were not tested. No model from 2025 or 2026 was tested. Ecommerce and transactional queries were not reported as a separate segment. Click-through, traffic and revenue were not measured at all.
None of that makes the paper bad. It is a careful piece of early work and it says what it did. It makes the way the paper is quoted bad.
How well supported are AI citations in the first place?
Less than you would assume, on the one study that measured it with human annotators. Note carefully what it measured: whether a cited source actually supports the sentence attached to it, not whether that sentence is true. The authors are explicit that verifiability is not factuality.
A 2023 study by Liu, Zhang and Liang, Evaluating Verifiability in Generative Search Engines, put 1,450 queries through Bing Chat, NeevaAI, Perplexity and YouChat, with 34 trained annotators. On average, 51.5 percent of generated sentences were fully supported by their citations, and only 74.5 percent of citations supported the sentence they were attached to. The authors call those figures concerningly low given what they describe as a facade of trustworthiness.
There is one finding in there that is genuinely uncomfortable and worth sitting with. Citation precision was inversely correlated with perceived usefulness, at minus 0.96 across the four engines tested, so four data points. The systems that cited most accurately were the ones users found least useful, because accurate citation meant copying or closely paraphrasing the source.
Both studies are old in AI terms. Neither tested the surfaces you care about now. That is the honest state of the public evidence, and anyone presenting this field as settled science is selling something.
What do working SEOs think moves AI visibility?
With the published evidence this thin, practitioner belief is worth recording separately, provided nobody confuses it with proof. In September 2026 Cyrus Shepard and Dawn Shepard surveyed 131 SEO professionals across more than one hundred candidate ranking factors. It is a survey about Google organic ranking rather than about AI answers, so it is not direct evidence for anything on this page. Two things in it are still worth your attention.
The first is what sits at the top. Relevance at 57.1 percent, backlinks at 54.8 percent, content quality at 47.6 percent, authority and trust at 36.5 percent, behaviour and click signals at 29.4 percent. Brand signals, which the survey defines as the importance of being a brand or entity that people search for, trust and visit, came sixth at 27 percent, ahead of technical SEO health at 17.5 percent. Nothing in the top six is an AI-specific tactic. If the systems answering these queries are grounded in web search results, as Google describes for AI Overviews and Microsoft describes for Copilot, and as OpenAI describes for ChatGPT search, which it says rewrites your query and sends it to search providers, then the things that decide what those searches return are the things that decide what gets quoted.
The second is a single practitioner observation, reported in that survey by Andrew Shotland: a client picked up a mention and link in a Reuters article, and their visibility in ChatGPT rose roughly fourfold over the following week. That is one anecdote, from one person, with no control and no published method. It is not evidence, and we are not presenting it as evidence. It is worth repeating only because it points the same way as every vendor document quoted on this page. Earned coverage on sources these systems already trust does more for AI visibility than any file you can add to your own server.
The practical reading for a store is that the AI-specific list on this page is short on purpose. Stop blocking the retrieval crawlers, stop throttling your own snippets, get into Bing properly, keep your product data accurate. After that, the work that moves AI visibility is the work that moves organic visibility, and it is mostly not technical.
How much traffic does an AI citation actually send?
On the best available measurement, very little, and the number is worth knowing before you budget anything.
What did Pew find?
Pew Research Center ran a metered browsing study: 900 US adults from a probability-based panel, tracking software installed, monitored through March 2025. Within 2,457,176 page visits they analysed 68,879 unique Google searches, of which 12,593 produced an AI summary.
Two numbers. Users clicked a link in the standard results on 8 percent of visits to pages with an AI summary, against 15 percent of visits to pages without one.
And the one almost nobody quotes: users clicked a source cited inside the AI summary itself on 1 percent of visits to search pages that had a summary. If the plan is to be the cited source rather than the clicked one, the trust work in E-E-A-T for ecommerce is the part that carries over.
One percent. That is the directly measured traffic value of being cited in an AI Overview, in that sample, in that month.
Two caveats that belong with it. The design is observational, so it does not establish that the summary caused the reduction. And pages with summaries differ systematically from pages without them, because, as the next section shows, summaries are triggered by a particular shape of query.
What triggers a summary, and why does that matter for a store?
Query shape, more than anything else, and the split is stark.
| Query shape | Produced an AI summary |
|---|---|
| Question format | 60 percent of the time |
| Ten or more words | 53 percent |
| Full sentence | 36 percent |
| One to two words | 8 percent |
Now think about what a product or category query looks like. “Leather dog collar”. “Merino base layer”. “Standing desk”. One to two words, which in this sample triggered a summary 8 percent of the time.
And what a buying guide query looks like. “How do I choose between leather and biothane for a dog collar”. Question format, ten or more words, which triggered one 60 percent of the time.
Your catalogue is structurally less exposed to AI Overviews than your content is. That is not a reason to ignore the surface. It is a reason to stop panicking about your product pages and to think carefully about which pages are actually in the firing line, in either direction.
One more Pew figure that kills a common assumption: 88 percent of summaries cited three or more sources, and only 1 percent cited a single source. This is not a winner-takes-all surface where one brand owns the answer.
What does Google say about the traffic effect?
Google says people using AI Overviews use Search more and are more satisfied, and that clicks from pages with AI Overviews are higher quality, with users more likely to spend more time on the site.
I am quoting that because it is Google’s position and you should know it. But Google publishes no methodology, no sample, no time period, no market, no query mix and no definition of higher quality alongside those claims in the document where they appear, and it does not publish absolute click volume comparisons between pages with and without AI Overviews. They are vendor assertions. Treat them as Google’s view, not as measurement, in exactly the same way you should treat a vendor’s case study.
Which pages on a store are actually candidates for AI citation?
Not the ones most stores start with. The exposure follows query shape, and query shape follows page type.
| Page type | Exposure to AI surfaces | Why |
|---|---|---|
| Buying guides and comparisons | High | Question-shaped queries trigger summaries most often |
| Sizing, fit and care guides | High | Natural language questions with a definite answer |
| Policy pages: returns, shipping, warranty | High and undervalued | Question-shaped, with an answer that exists only on your site |
| Product pages | Low for AI Overviews, separate path for ChatGPT | Short head queries rarely trigger a summary |
| Collection and category pages | Low | Same reason |
| Blog posts answering one question well | High | Fan-out targets subquestions |
| Home and about pages | Low, but used for brand grounding | Rarely the answer, often the identity check |
What should a buying guide look like now?
The same as it always should have, with three emphases the research supports.
Put verifiable specifics on the page. The GEO paper’s three successful methods were quotations from credible sources, quantitative statistics, and citing sources. Whatever you think of the paper’s limits, those three are also just good content, and they are the opposite of the vague, padded guides most stores publish.
Answer the subquestions, not just the head question. Fan-out looks for the pieces. A guide that answers sizing, care, materials, cost and use case explicitly, each under its own heading, gives the system more surfaces to grab.
Make the answer extractable. Microsoft’s published guidance on content for AI answers names headings that define content slices, question and answer pairs, lists and tables as things that help, and long undifferentiated text blocks, content hidden in tabs or expandable menus, and critical information locked in PDFs or images as things that hurt. That is a vendor telling you how to format for its own system.
Why are policy pages undervalued?
Because they answer the questions people actually ask an assistant before buying, and almost no store treats them as content.
“Can I return this if it does not fit.” “How long does delivery take to Australia.” “Is there a warranty.” Those are question-shaped, natural language, and have a definite answer that lives on your site and nowhere else. They are also exactly the sort of task-completion query Google names when it says AI Mode is trained to decide when to include hyperlinks, if the user may want to take action or finish a task on a website.
Most stores have a returns page that is 200 words of legal boilerplate with no headings. Rewriting it as clear question and answer pairs is a couple of hours of work on a page you already own.
Structured data helps here for its documented purpose rather than for AI: Google publishes merchant return policy and merchant shipping policy markup, and getting those right is documented search-appearance work. It is not an AI lever and I am not going to pretend it is.
What about product and collection pages?
Keep doing the ordinary work, and change your expectations about which surface they win on.
Short head queries rarely trigger an AI Overview. Your product pages are competing in ordinary search, in Shopping, and in the ChatGPT commerce pipeline, which is a feed, not a page. That is the next section.
The technical hygiene still matters, because the eligibility gate is being indexed and snippet-eligible, and stores fail that in ways that have nothing to do with AI. If your product templates are failing there, the platform-specific causes are in Shopify SEO issues and WooCommerce technical SEO, and the speed side of eligibility is in ecommerce Core Web Vitals.
How do you get a product into ChatGPT’s shopping results?
Through a feed, and mostly not through anything you write.
What is the actual mechanism?
OpenAI operates a commerce pipeline that takes merchant product data directly, documented as currently available to approved partners, working from daily snapshots. It sits alongside ChatGPT search, which rewrites a user’s question into targeted queries and sends them to search partners. OpenAI names Microsoft and Shopify among those partners.
So there are two routes into a ChatGPT answer about products. One is being cited as a web source, which works the way the rest of this post describes. The other is being in the product data, which does not.
What does OpenAI say it ranks on?
Four things, in its own words: availability, price, quality, and whether they are the maker or primary seller. Results may carry a “Best price” label. OpenAI states separately that product results are selected independently by ChatGPT, are not ads, and are not influenced by OpenAI partnerships.
That is one of the very few published ranking statements any AI vendor has made about commerce, and it is worth reading for what it is not about. It is not about your copy. It is not about your blog. It is about your offer.
Why is a reseller structurally disadvantaged?
Because being the maker or primary seller is a named factor, and you either are or you are not.
If you are a reseller competing on a commodity product against the manufacturer, OpenAI’s stated criteria put you behind on one factor before anything else is considered, and price and availability are the two you can still move. No amount of content work changes the first one.
That is not a counsel of despair. It is a reason to be honest about where your effort goes. If you sell other people’s products, your differentiated content, your bundles, your service and your own brand are where the return is, and competing head-on for “best price on commodity item” inside an assistant is not a winnable game.
What can you actually do?
Get your feed right, because the feed is the surface. Accurate availability, accurate price, correct identifiers, complete attributes and no disapprovals. That is the same work that decides your Shopping performance, which is why it is worth doing regardless of what ChatGPT does next.
We cover the feed itself in product feed optimization and the rejections that quietly take products out of circulation in Google Merchant Center errors. The overlap with this topic is almost total, and it is the least fashionable and most reliable AI commerce work available right now.
One attribution note from OpenAI’s own merchant guidance, for anyone who does get into the feed programme: add feed attribution parameters to your product URLs, such as a feed value on the medium parameter, and keep internal tracking parameters consistent between snapshots. That is self-attribution on a channel you control, and it is the only documented way to see those clicks.
Why does Bing matter more here than its market share suggests?
Because it is the only place you can currently see whether any of this is working, and because ChatGPT search sends queries to Microsoft.
What does Bing Webmaster Tools show that Google does not?
AI Performance, a report in public preview since February 2026, which shows when your site is cited in AI-generated answers across Microsoft Copilot, AI-generated summaries in Bing, and select partner integrations.
The metrics it exposes are: total citations, the number of times your pages appeared as sources in AI answers; average cited pages per day; grounding queries, which are the key phrases the AI used when retrieving content that was referenced; citation counts for specific URLs; and a visibility trend over time.
Grounding queries is the interesting one. It is a list of what the machine asked, not what a human typed, which is a view of your content that nothing else gives you.
One stated limit, and it matters: the data represents a sample of overall citation activity. Treat the numbers as directional rather than as a census.
Why set this up if most of your traffic is Google?
Three reasons, and none of them is Bing’s traffic share.
Of the consoles checked here it is the only first-party AI citation reporting, and it is free. Google reports AI Overviews and AI Mode inside the Web search type with no way to separate them. Bing names citations as their own metric at page level. If you want any evidence at all that your content is being used in AI answers, this is where it comes from.
OpenAI names Microsoft among the search partners it sends its rewritten queries to, so your Bing presence plausibly touches part of what ChatGPT search sees. OpenAI does not document how much, and it also runs its own crawler in OAI-SearchBot, so do not treat Bing as a back door into ChatGPT.
It gives you a push mechanism. Bing states that freshness signals directly influence how quickly updates are reflected in search results and AI generated answers, prefers XML sitemaps because they carry lastmod, fetches sitemaps immediately on submission and at least once a day after that, and supports IndexNow. For a store where price and stock change constantly, that is the only push mechanism any vendor in this research documents for getting a change into AI answers faster.
If you already run Microsoft Ads, this is a short setup on an account you have. If you do not, we cover the paid side separately in Microsoft Ads management, but the webmaster tools side is free and worth doing either way.
What does Microsoft say about content?
Its published guidance for inclusion in AI search answers asks for content that is fresh, authoritative, structured and semantically clear. Named practices are clear titles and H1s aligned to intent, heading hierarchies that define content slices, question and answer pairs, lists and tables, and JSON-LD schema markup. Named things to avoid are long text blocks, content hidden in tabs or expandable menus, and relying on PDFs or images for critical information.
Note that Microsoft does name schema markup, where Google says no special schema is needed. Those are two vendors with two positions and you should know both rather than picking the one that suits.
How do you measure any of this?
Partly, and being clear about the boundary is the most valuable thing an advisor can give you here.
What does Search Console count?
More than people assume, but not separably.
Google states that sites appearing in AI features are included in overall search traffic in Search Console, reported within the Web search type. For AI Overviews, standard impression rules apply, clicking a link to an external page counts as a click, and an AI Overview occupies a single position with all links in it assigned that same position. For AI Mode, clicks and impressions are counted the same way and position follows the ordinary results page methodology.
Now the limitation. Search Console’s documented dimensions are queries, pages, countries, devices, search appearance and dates. AI Overviews and AI Mode do not appear in the search appearance types. The available search types are Web, Image, Video and News. There is no AI search type.
So AI clicks are in your numbers, and there is no documented way to separate them out. Anyone telling you they can report your AI Overviews traffic from Search Console is either using an inference or making it up.
There is a small diagnostic point worth adding. Google’s own documentation on debugging search traffic drops does not mention AI Overviews or AI Mode as a cause or factor anywhere. It points at the Performance report, data anomalies, crawl stats, page indexing, security issues and manual actions, server availability, robots.txt, content quality, position changes and seasonality. Before you attribute a traffic drop to AI, work through Google’s own list.
Can I see referral traffic from ChatGPT or Perplexity?
Not on any documented contract.
No page from Google, OpenAI, Anthropic, Microsoft or Perplexity that this research fetched publishes a referrer header value, a referral domain list, a UTM convention or a link attribution scheme for traffic sent from its assistant to a website. Your analytics tool may well be grouping referrers that look like assistant domains, and that is often directionally useful, but it is inference from observed strings rather than something any vendor has committed to, and it can change without notice.
The one exception is the OpenAI merchant feed guidance mentioned earlier, where the merchant adds the parameter themselves inside their own feed. That is self-attribution on a channel you control, available only to approved feed partners, and it covers product links from feeds rather than citations of your pages in ordinary answers.
What can I measure completely?
Your server logs, and this is the most underused measurement in the subject.
Four of the five vendors in this research publish a machine-readable IP list for their bots. OpenAI publishes one per bot. Anthropic publishes one covering all three. Perplexity publishes one each for its crawler and its user fetcher. Common Crawl publishes one. Google does not publish one in the documentation checked here, but it does document its user agent strings in full.
That means you can establish, from data you already own, exactly which AI systems fetched which of your pages, how often, and whether the request was authentic rather than something spoofing the user agent. It is complete rather than sampled, it costs nothing but log retention, and it answers a question no console will answer: is anything actually reading this page.
It does not tell you whether you were cited. It tells you whether you were fetched, which is the necessary condition.
What is not measurable at all?
Five things, and you should hear them from your advisor rather than discover them later.
| Question | Can you answer it? | Where from |
|---|---|---|
| How many Google clicks came from AI surfaces | No | No documented dimension separates them |
| How often you are cited in AI Overviews or AI Mode | No | Google publishes no citation count |
| Whether your content was used without a link | No | No vendor reports uncited usage |
| Revenue attributable to ChatGPT or Perplexity | Not reliably | No vendor publishes a referral scheme |
| Your visibility inside Claude, Gemini or Perplexity | No | None checked here publishes a webmaster console |
| How often you are cited in Copilot and Bing AI answers | Partly | Bing Webmaster Tools, sampled |
| Which AI bots fetched which pages | Yes, completely | Your own server logs |
One more that gets promised and cannot be delivered: your rank or position within an AI answer. Google assigns every link in an AI Overview the same position value, so even the one vendor that reports position deliberately does not differentiate placement inside the overview. There is no position three in an AI Overview.
What should an ecommerce store actually do about AI search?
Three lists. Do the first, budget for the second, decline the third.
What is real and cheap?
Audit your robots.txt against the bot table above and decide training and visibility separately. This is the highest-value hour in the whole subject because the downside of getting it wrong is total.
Check your templates for nosnippet and for a restrictive max-snippet, in the meta tag and in the X-Robots-Tag header.
Set up Bing Webmaster Tools and submit an XML sitemap with accurate lastmod. Turn on IndexNow if your platform supports it.
Start retaining server logs and filtering by the published bot user agents and IP lists.
Rewrite your returns, shipping and warranty pages as clear question and answer pairs with real headings. They are the most asked and least optimised pages you own.
Get your product feed accurate. Availability, price, identifiers, attributes, no disapprovals.
What is real and worth budget?
Buying guides and comparison content that actually answer the subquestions, with verifiable specifics, real numbers and cited sources rather than padding. This is the one content intervention with any research behind it, on a harness nothing like the engines you care about, which is why I offer it as good content practice rather than as a proven AI lever.
Making sure your important content is in text rather than in images, PDFs or collapsed tabs. On a lot of stores the sizing chart, the care instructions and the material spec are in a JPEG.
Fixing the eligibility basics on your product and collection templates, because indexed and snippet-eligible is the gate and stores fail it for ordinary technical reasons.
Structured data, for its documented purposes. Product, merchant listing, variants, review snippet, return policy, shipping policy, breadcrumb and Organization. Not because it is an AI lever, because it is not, but because it drives documented search appearances.
What has no evidence behind it?
llms.txt as a visibility lever. Nothing in the twenty vendor pages checked says any assistant reads it, and Google states you do not need to create new machine readable files or AI text files to appear in its AI features.
Schema markup sold as an AI Overviews lever. Google says there is no special structured data you need to add.
Any “GEO ranking factors” list. No vendor in the documentation checked here publishes ranking factors for AI answers. Google publishes eligibility conditions and best practices, OpenAI publishes four commerce criteria, and that is the extent of what I could find.
Writing in a more authoritative tone. The paper that coined GEO tested exactly that and found no significant improvement.
Keyword stuffing for AI, in any of its rebranded forms. The same paper found it performed worse than baseline.
Reporting that claims to show your AI Overviews traffic from Search Console. It is not a documented dimension.
What should I ask anyone selling me AI search work?
Five questions. The answers tell you everything.
Which of the bots are you going to allow and which are you going to block, and why. If they cannot distinguish OAI-SearchBot from GPTBot, stop there.
How will you measure success, and which part of that is sampled, which is inferred and which is complete.
What evidence is there for the specific tactic you are recommending, and can I see the source.
What are you going to stop doing in my ordinary SEO to pay for this.
What happens to this plan if the surface changes in six months.
How should this change your SEO and PPC budget?
Less than the noise suggests, and in a specific direction.
The honest position from everything above is that AI search is currently a small, poorly measurable traffic source with a directly measured click rate of about one percent on cited sources, on query shapes that mostly are not your product terms, with no published ranking factors and no reliable attribution. Meanwhile the things that make you visible in it are, almost item for item, the things that make you visible in ordinary search: be crawlable, be indexed, be snippet-eligible, have your important content in text, answer real questions properly, and keep your feed accurate.
So the budget answer is not a new line item. It is a set of checks against your existing work, plus a content emphasis on question-shaped pages, plus getting your feed right.
Where it does change things is in what you stop assuming. If your informational content loses clicks while your rankings hold, the AI summary may be part of it, and the response is to make that content worth clicking rather than to publish more of it. If your paid traffic is holding while organic informational traffic softens, that is a signal about where your acquisition is actually coming from, and it is worth reading properly rather than panicking. We do that analysis alongside the Google Ads and ecommerce PPC side rather than as a separate exercise, because the two budgets move together.
If you want the robots.txt and snippet audit above run on your store rather than explained, that is what we do, and the details are on our ecommerce AI search page.
Why work with Hustle Marketers on AI search for ecommerce?
Because most of what is being sold under this heading is an acronym attached to work nobody has evidence for, and the parts that do have evidence are unglamorous.
I am Ishant Sharma. I have run Google Ads, Microsoft Ads and ecommerce SEO since 2013, across Shopify, WooCommerce, Magento and BigCommerce stores.
Four things that are different about how we approach this.
We start with the blocking audit, because it is the only part of this subject where you can lose everything overnight from a single pasted line. The bot table in this post is the first hour of the engagement, and on more than one store it has been the whole finding.
We separate what is measurable from what is inferred, in writing, before we start. You will get a list of what we can prove, what we can estimate and what nobody can see, and that list will not change halfway through to suit a report.
We treat the feed as part of this, not as a separate channel. OpenAI’s published commerce criteria are availability, price, quality and who makes the product. Two of those live in your feed. Any AI commerce plan that does not look at your feed is not a plan.
We will tell you when the honest answer is that this is not your problem. If your product queries are one and two word head terms, your exposure to AI summaries is structurally low, and your money is better spent elsewhere. That is a real conclusion and we reach it regularly.
On evidence, I will be straight with you. I am not going to show you a screenshot of a brand being mentioned in ChatGPT, because there is no way for you to verify it, no way to know it is repeatable, and no way to connect it to revenue. What is published with real numbers is our pet ecommerce SEO case study, a Shopify store in Australia and New Zealand: organic clicks 6,517 to 10,189, impressions 421,365 to 579,659, average position 10.3 to 9.4, comparing the previous three months to the most recent three. That page also reports AI citations rising from about 111 to 171, but over a six month trend rather than the three month Search Console window above, and it does not name the tool, the prompt set or the sampling method behind it. No vendor console publishes a citation count for Google or ChatGPT, so it is not first-party data and I would not ask you to weigh it as evidence. The Search Console numbers are the evidence. For a second store with first-party numbers, our Judaica ecommerce SEO case study shows Search Console clicks up about 31 percent over 28 days, 10.59K free listing clicks in Merchant Center and AI assistant sessions in GA4, with the seasonality and GA4 channel caveats spelled out.
We publish how we report, in reporting integrity, for the same reason.
The platform work that sits underneath all of this is in Shopify SEO issues, WooCommerce technical SEO, Magento SEO and ecommerce SEO platforms, and the broader service is our SEO agency work.
Let me know if you want me to run the robots.txt and snippet directive audit on your store. It takes an hour and it is the one part of this where the downside is real.
AI SEO for ecommerce FAQs
Do I need to do anything special to appear in AI Overviews?
No. Google states there are no additional requirements to appear in AI Overviews or AI Mode and no other special optimisations necessary. The condition is that your page is indexed and eligible to be shown in Google Search with a snippet. Everything else Google lists as best practice is ordinary SEO.
Does schema markup help me appear in AI Overviews?
Google says there is no special schema.org structured data you need to add. Structured data is still worth doing for its documented purposes, which on a store means Product, merchant listing, variants, review snippet, return and shipping policy, breadcrumb and Organization markup, and Google does say your structured data should match the visible text. It is not an AI lever and anyone selling it as one should be asked for the source.
If I block GPTBot, do I disappear from ChatGPT?
No. GPTBot is OpenAI’s training crawler, and blocking it means your content should not be used in training foundation models. The bot that controls whether you appear in ChatGPT search answers is OAI-SearchBot, and OpenAI states that sites opted out of it will not be shown in ChatGPT search answers. Those are two different lines in your robots.txt.
Does blocking Google-Extended remove me from AI Overviews?
No. Google states that Google-Extended does not impact a site’s inclusion in Google Search and is not used as a ranking signal. Its documented scope is training future Gemini models for Gemini Apps and grounding on Vertex AI. AI Overviews and AI Mode are not in that scope, so blocking it changes nothing about either, in both directions.
What is the difference between AI Overviews and AI Mode?
AI Overviews is an AI-generated summary at the top of an ordinary Google results page, with the normal results still underneath. AI Mode is a separate conversational experience where a user can ask follow-up questions. Google describes query fan-out, splitting a question into subtopics and searching each at once, across its AI features rather than as the difference between the two. They share eligibility rules and both report into the Web search type in Search Console.
Could my SEO plugin be blocking me from AI answers?
Yes, through the snippet directives. Google states that nosnippet prevents content being used as a direct input for AI Overviews and AI Mode, and that max-snippet limits how much of the content may be used. A restrictive max-snippet set years ago for an unrelated reason is an active throttle today, and it can be set globally in a plugin or at server level in an X-Robots-Tag header.
Should I add an llms.txt file?
You can, and it is cheap, but do not buy it as a visibility service. This research checked around twenty primary pages across Google, OpenAI, Anthropic, Microsoft and Perplexity and found no statement from any of them that its assistant reads llms.txt when answering a query. The strongest thing that exists is a Chrome Lighthouse audit that calls the file an emerging convention and describes providing it as optional at the moment.
Where does the 40 percent GEO figure come from?
A November 2023 arXiv paper, revised in June 2024 and accepted to KDD 2024, that coined the term. The figure is a relative gain on Position-Adjusted Word Count, a text-share metric the authors defined, measured on their own harness wrapping gpt-3.5-turbo across a test set that is around 80 percent informational. On Perplexity, the one commercial engine tested, the gain was 9 percent on one of the two metrics over 200 samples. The paper states it did not evaluate effects on search rankings.
Does keyword stuffing work for AI search?
No. The paper that invented the term generative engine optimization tested it directly and it was the worst performing of nine methods, scoring below the baseline. The paper states that simple methods traditionally used in SEO do not perform well. Writing in a more authoritative tone also produced no significant improvement in the same study.
How much traffic does being cited in an AI Overview send?
On the best available direct measurement, very little. Pew Research Center’s metered study of 900 US adults across 68,879 Google searches found users clicked a source cited inside the AI summary on 1 percent of visits to search pages that had one. Clicks on standard results were 8 percent on pages with a summary against 15 percent without. The study is observational, so it does not establish causation.
Are my product pages at risk from AI Overviews?
Less than your buying guides are. In the Pew sample, one and two word searches produced an AI summary only 8 percent of the time, while question-format searches produced one 60 percent of the time. Most product and category queries are short head terms. Your informational content is structurally more exposed than your catalogue, in both directions.
How do I get my products into ChatGPT shopping results?
Through a feed, not through content. OpenAI runs a commerce pipeline that takes merchant product data and is documented as currently available to approved partners, and it names Shopify among its search partners. The ranking criteria OpenAI publishes are availability, price, quality and whether you are the maker or primary seller. Accurate feed data is the lever.
Can I see how much traffic AI Overviews sends me in Search Console?
No. Google counts AI Overviews and AI Mode clicks and impressions into the Web search type, and the documented dimensions are queries, pages, countries, devices, search appearance and dates. AI Overviews and AI Mode are not listed among the search appearance types and there is no AI search type. The clicks are in your total and cannot be separated out.
Is there any console that shows AI citations?
Bing Webmaster Tools, in public preview since February 2026. It reports total citations, average cited pages per day, grounding queries showing the phrases the AI used when retrieving your content, and citation counts for specific URLs, across Copilot and Bing AI summaries. Bing states the data is a sample of overall citation activity, so treat it as directional.
How do I know if AI bots are actually reading my site?
Your server logs, and this is the only complete measurement available. Four of the five vendors publish a machine-readable IP list: OpenAI one per bot, Anthropic one covering all three of its bots, Perplexity one each for its crawler and its user fetcher, and Common Crawl one for CCBot. Google documents its user agent strings in full. You can confirm which systems fetched which pages, and for the four with published IP lists you can verify the request was authentic rather than spoofed.
Should I move budget from SEO to AI search work?
On current evidence, no, because the two are mostly the same work. Being visible in AI answers requires being crawlable, indexed, snippet-eligible, having important content in text and answering real questions properly. What is worth adding is a blocking audit, Bing Webmaster Tools, feed accuracy and an emphasis on question-shaped content. What is not worth paying for is anything sold as a distinct ranking system, because no vendor in the documentation checked here publishes ranking factors for AI answers.
Sources and verification
Everything above was checked in September 2026 against vendor documentation and research papers that were fetched and read, not against summaries of them. Bot names, user agent behaviour, robots.txt tokens, directive definitions, published limits and ranking criteria come from the vendor that operates the system. Research figures come from the papers themselves.
Some things are deliberately absent. There is no Gartner figure here about search volume or organic traffic declining, because no Gartner primary source was fetched for this post, and I am not repeating a number I have not read at source. There is no statistic giving AI search a share of total search volume or of ecommerce traffic, because I could not reach a primary source for one. There is no bingbot user agent string, because Bing’s crawler help pages are client-rendered and returned nothing. There is no list of GEO ranking factors, because no vendor publishes ranking factors for AI answers. And there is no claim that any assistant reads llms.txt, because none of the twenty primary pages checked says so.
One limitation on the competitor observations. The search budget available for this research was exhausted before the SERP work could be done, so this post makes no claims about what currently ranks for these terms or about what other pages do or do not cover. Thirty-one third-party pages were read end to end from known URLs, and twelve more could not be reached and are excluded.
Google references used include the AI features documentation, Search Essentials, the robots meta tag and X-Robots-Tag documentation, the crawler overview, the common crawlers page where Google-Extended is documented, the user-triggered fetchers page, the structured data gallery, the Search Console performance report documentation, the data and metrics definitions, the dimensions documentation, the debugging traffic drops guide, Google’s consumer explanations of AI Overviews and AI Mode, and Google’s own AI Overviews and AI Mode explainer.
Other vendor references include OpenAI’s bot documentation, OpenAI on ChatGPT search, OpenAI on shopping in ChatGPT, OpenAI’s commerce best practices, Anthropic’s crawler documentation, Perplexity’s bot guide, Bing on data-nosnippet, Bing on sitemaps and freshness in AI-powered search, the Bing Webmaster Tools AI Performance announcement, Microsoft’s guidance on optimising content for AI search answers, Common Crawl’s CCBot page and the llms.txt specification.
Research references are GEO: Generative Engine Optimization, Evaluating Verifiability in Generative Search Engines, and Pew Research Center’s finding on clicks and AI summaries with its published methodology. The Chrome Lighthouse note on llms.txt is here.
For the platform-agnostic version, our guide to AI search optimization covers crawler access, extractable answers and why attribution is the weakest part of this whole subject.
If you are fixing this as part of a wider organic push, our ecommerce SEO guide covers where it sits against category pages, internal linking and product data.
Summarise this article with:









