How to Improve Your Brand's Visibility in AI Search Engines
Google, OpenAI, and Perplexity publish little beyond ordinary search fundamentals, and Google names specific tactics, including llms.txt, and states plainly they do nothing. Bing is the exception, with real published guidance. This guide covers what each documents, in its own words, plus how to check whether any of it results in an actual mention.
- You will end up with
- A site confirmed reachable by the major AI crawlers, marked up with accurate and current schema.org data (with any dead rich-result markup removed), and a repeatable routine for checking, across ChatGPT, Perplexity, Copilot, and Google's AI Overviews and AI Mode, whether any of it results in an actual mention. Not a guaranteed ranking, citation, or traffic increase: Google, OpenAI, and Bing each say so explicitly, in those words, on their own pages.
- Time
- About 2 to 3 hours for the technical and structured-data checks in the first seven steps, then about 15 to 20 minutes a month, ongoing, for the monitoring routine in the last three.
- Cost
- $0. Robots.txt is a plain text file you can read yourself, Google's Rich Results Test and Search Console are free, and every checker linked in this guide runs free with no signup. Paid AI-visibility monitoring platforms are named where relevant, but nothing here requires buying one.
Four companies publish documentation about getting into their AI answers: Google, OpenAI, Perplexity, and Microsoft's Bing. Three of them say, in their own words, that there's nothing to do beyond ordinary search fundamentals. Google goes further: it names specific tactics people are commonly told to do, including creating an llms.txt file, and states plainly that Google Search doesn't use them and doing so changes nothing. OpenAI publishes less than that: a crawler to allow, an opt-out, and "multiple factors" it doesn't name. Perplexity publishes the least: crawler access and a source-trust label, nothing about how citation selection works. Bing is the exception. Its webmaster guidelines name a second discipline, generative engine optimization, and give specific, dated, quotable guidance, still with an explicit no-guarantee caveat.
That asymmetry, not a list of ten equal tactics, is the honest structure of this subject, and it's why the steps below are organized by what's actually documented rather than presented as one uniform checklist. Nothing here promises a ranking, a citation, or a visitor. Google, OpenAI, and Bing each say so explicitly, in their own words, and this guide isn't going to claim more certainty than the platforms themselves do.
1. Confirm the crawlers that answer real-time questions can actually reach you
Every major AI answer engine operates more than one crawler. Broadly, each company runs one kind that gathers material to train its models, and a separate kind that fetches pages the moment someone asks a question the model wants to search the web to answer. Blocking the first doesn't, by itself, stop you from being cited in a live answer; that depends on the second, a differently named agent with its own line in robots.txt.
OpenAI names three agents, independently controlled: GPTBot, "used to make our generative AI foundation models more useful and safe"; OAI-SearchBot, "used to surface websites in search results in ChatGPT's search features," which OpenAI directly recommends allowing; and ChatGPT-User, which fires only on a live, user-triggered fetch and which OpenAI states "is not used to determine whether content may appear in Search." Blocking OAI-SearchBot doesn't fully remove you either: OpenAI documents a fallback where a disallowed page can still surface as a bare link and title if OpenAI has the URL from elsewhere; a noindex tag is the only way to opt out of that too. OpenAI also requires more than an open robots.txt file: a site must "confirm that the website host or content delivery network allows traffic from OpenAI's published searchbot IP addresses," so a firewall or CDN rule can still block the crawler even when robots.txt allows it.
Perplexity documents PerplexityBot, "designed to surface and link websites in search results on Perplexity" and explicitly "not used to crawl content for AI foundation models," plus a Perplexity-User agent that fires on a live, user-triggered fetch and "generally ignores robots.txt rules" by design. Anthropic's own crawler page, checked separately from this guide's main research pass, names ClaudeBot, Claude-User, and Claude-SearchBot as its current agents, following a comparable split. Googlebot, covered next, feeds both classic results and AI Overviews.
The action: open yourdomain.com/robots.txt in a browser and read it, or hand the job to a checker built for exactly this. Is My Brand In AI's bot checker is a free, no-signup way to get a fast, curated read on the major crawler families in one pass. AIclicks' checker validates against the same open-source parsing library Google itself publishes. Hyperleap's validator also flags a robots.txt file that's syntactically broken, which matters because one misplaced rule can silently block a crawler you meant to welcome. xSeek's checker tests against a wider 2026 crawler set and returns a single readiness score. A fuller side-by-side sits in our comparison of the free AI bot checkers.
One limit every checker in this category shares: each reads robots.txt and nothing else. A firewall rule or CDN preset can turn a crawler away before it reaches the file, and none of these tools can see that layer.
2. Know what Google, OpenAI, and Perplexity actually say beyond crawl access: not much
Google's current documentation states it plainly: "There are no additional requirements to appear in AI Overviews or AI Mode, nor other special optimizations necessary." Eligibility is simply that "a page must be indexed and eligible to be shown in Google Search with a snippet, fulfilling the Search technical requirements." Google goes further than silence. It names the tactics people are commonly told to do and rejects them: "You don't need to create new machine readable files, AI text files, markup, or Markdown to appear in Google Search (including its generative AI capabilities), as Google Search itself doesn't use them," and doing so "will neither harm nor help your site's visibility or rankings." That includes llms.txt by name, added to the page in a June 2026 update for exactly this reason.
OpenAI publishes less than that. Beyond the crawler mechanics in step 1, its search-help page states the entire eligibility position in two sentences: "ChatGPT ranks search results using multiple factors intended to help users find relevant, reliable information. Placement is not guaranteed." It doesn't name one of those factors on any OpenAI-owned page.
Perplexity publishes the least. Its own explainer describes source selection only as gathering "information from authoritative sources like articles, websites, and journals" and compiling "the most relevant insights," with no ranking factor named anywhere. Perplexity does publish one detailed page on how it labels domains, Government, Academic, or Trusted, based on correction practices and named authorship, but states plainly the label is a display badge set by its own review process, not a citation gate, and most domains carry no label at all.
None of this is a gap in this guide. It's what these three companies have chosen to publish, in their own words.
3. Three separate controls, three different effects
Most pages on this subject conflate these. Robots.txt, covered in step 1, controls crawling: whether a bot can fetch the page at all. Google-Extended is a separate token controlling training: whether your content can improve the standalone Gemini app and Vertex AI's generative models. It doesn't remove you from Google Search or AI Overviews, because both run on Googlebot's regular crawl, not on Google-Extended.
A third control governs a narrower thing: whether an already-crawled, already-indexed page can be quoted inside a generated answer. Google's current robots meta tag documentation states nosnippet "will also prevent the content from being used as a direct input for AI Overviews and AI Mode," and that max-snippet "will also limit how much of the content may be used as a direct input for AI Overviews and AI Mode." Both work on a page that's otherwise fully indexed; they're the tool for keeping specific text out of a generated answer while the page stays visible as an ordinary result. For the crawling lever specifically, the full walkthrough on writing and testing robots.txt rules is covered separately in our guide to controlling AI crawler access.
Apple publishes the same split, Applebot-Extended for training versus Applebot for crawling, and Common Crawl's CCBot feeds a corpus several smaller and open models train on. A Disallow: * written years ago for an unrelated reason can end up blocking training, crawling, and answering all at once, without anyone deciding to.
4. Mark up the page with accurate schema.org data, and skip the two that are dead
Schema.org is a shared vocabulary, not a Google product; it was founded jointly by Google, Microsoft, Yahoo, and Yandex in 2011. Google's policy is precise about what markup earns you: "Using structured data enables a feature to be present, it does not guarantee that it will be present," and extra schema.org properties beyond what Google reads carry no penalty and no benefit.
Two commonly-recommended types are worth naming, because a lot of published advice still recommends them and Google's Search Gallery no longer lists either. FAQPage rich results stopped showing for ordinary sites in 2023, restricted to "well-known, authoritative government and health websites," and Google's own changelog states the feature itself "will no longer appear in Google Search starting May 7, 2026," with the documentation page since removed entirely. HowTo went desktop-only in that same 2023 change and isn't live in any form we could find as of this read, though Google hasn't published an exact date for that final removal the way it did for FAQPage. Marking up either does no harm, but it earns nothing today.
The action, and the verification, are the same step: run the page through Google's Rich Results Test and through schema.org's own validator. Both are free, self-serve, and will tell you directly whether your markup parses.
5. Bing is the exception: nine documented items, plus named schema types
Steps 2 through 4 describe companies publishing little beyond "do ordinary search fundamentals well." Bing is different, and it's worth marking as Bing's own guidance rather than folding in as if it applied everywhere. Bing's Webmaster Guidelines name a second discipline directly: "Generative Engine Optimization (GEO) focuses on content eligibility for grounding and reference in AI responses," with sections dedicated to it, distinct from ordinary SEO.
Bing states a page is more likely to be selected for grounding and citation when content "stands on its own" and "can be verified independently," when entities are "defined clearly and consistently," when each URL "focuses on a single topic," and when key information is surfaced early. Its AI Performance launch post names five more, in its own words: strengthen depth and expertise; improve structure and clarity with headings, tables, and FAQ sections; support claims with evidence; keep content fresh and accurate; and reduce ambiguity by aligning text, image, and video around the same entities. A separate, named Microsoft blog post adds specifics no other platform here publishes: title, meta description, and H1 as signals "AI systems use," named schema types (product, review, FAQ, or event, in JSON-LD), and a formatting preference for lists, direct Q&A, and comparison tables.
Bing also documents AI-specific behavior for meta directives Google doesn't describe in these terms: NOARCHIVE "prevents content from being used in Copilot responses and grounding results," and NOCACHE limits Copilot "to using only the URL, title, and snippet, reducing citation depth." Bing recommends IndexNow by name for this purpose, "accurate and up to date content is important for inclusion and citation in AI-generated answers," plus registering local businesses with Bing Places for Business.
One caveat Bing states as directly as everything above: "GEO does not guarantee grounding or citations in AI experiences." More documentation is not a promise, even from the platform publishing the most of it.
6. Make who's behind the page unambiguous
None of the four platforms in this guide documents authorship or organizational clarity as a citation requirement. What's checkable regardless: a byline with a real name, a company or publisher page that plainly says who you are, and Organization or Person schema that matches both. Confirm every page making a claim worth citing has one, and that any Organization markup matches your current about page rather than an old version.
7. Keep the facts current
A page an AI system crawled six months ago can still be the version it quotes today. The sitemaps protocol, supported by Google and Bing, lets you list URLs with an accurate lastmod date so a crawler can prioritize what actually changed. Keep those dates honest: a sitemap claiming every page changed today, every day, trains nothing and helps nobody. IndexNow, Bing's specific recommendation for this, is covered in step 5.
The action: audit your sitemap's lastmod values against your actual publish history, and fix any page whose live facts, pricing, specs, dates, have drifted from what was true when it was last crawled.
Everything above clears preconditions. None of it tells you whether ChatGPT, Perplexity, Copilot, or Google's AI Overviews and AI Mode actually mention you. Measurement is where the honest answer gets uncomfortable: no platform in this guide, including Bing, publishes a free, first-party way to see whether your brand is mentioned by name without a link. That's the finding the next three steps work around, not a gap in what we looked for.
8. Check what each platform's own tool actually shows you, and notice what it can't
The free crawler-permission checkers from step 1, plus BrandCited's wider sweep of roughly 64 AI-related user-agents, all answer the same question: is a crawler allowed in. None of them, and no platform's own tool either, can show you whether your brand was actually mentioned without a link. Here's what each platform's own first-party tool actually shows, so the gap is concrete rather than implied.
Google Search Console's generative AI performance report gives impressions only, "the report does not include click data," and bundles AI Overviews and AI Mode into one number with no way to separate them. It shows impressions of your own already-indexed pages only, nothing about a competitor or an unlinked mention.
OpenAI's only documented path is narrower: publishers "can track referral traffic from ChatGPT using analytics platforms such as Google Analytics," reading the utm_source=chatgpt.com parameter ChatGPT adds automatically. That measures clicks that already landed on your site, nothing about a citation nobody clicked.
Perplexity publishes no general-purpose measurement tool at all; the only analytics it documents live inside an invite-only Publishers Program, run through a named third party, not a self-serve dashboard.
Bing comes closest. Its Citation Share shows "the percentage of citations attributed to your site out of all citations shown across all sites for that same grounding query," the only first-party metric here that compares you against a total. Microsoft states its limit just as directly: Citation Share "does not expose competitor domains, represent traffic share, or assign quality scores to content."
Every first-party tool these four platforms publish measures your own already-indexed, already-linked pages. If a paid AI-visibility tool implies it can show you an unlinked mention the platforms themselves can't, that's the specific claim to press it on.
9. Ask the engines about your own brand, on a schedule, and log what comes back
Because no platform's own tool shows an unlinked mention, asking directly is the closest thing to a real answer. Type a question a real customer might actually ask into ChatGPT, Perplexity, Copilot, and a Google query specific enough to trigger an AI Overview, plus a separate check in AI Mode. Note whether you're mentioned, what's cited if you are, and what's cited instead of you if you're not.
Log the date, the engine, the exact question, and the result on a fixed schedule, monthly is reasonable, and you have a record of whether anything changed after you made a change. It costs nothing, needs no vendor, and asks the systems themselves rather than a proxy metric.
10. Earn mentions on pages you don't control
Every system in this guide is built to browse and cite pages across the open web, not only the ones you publish yourself. A claim that lives solely on your own site can never be one of the outside sources an engine cites when deciding whether to trust or name you, since your own site is the thing being evaluated, not independent evidence about it.
One AI-visibility research platform, in material describing its own product, states it has found unlinked brand mentions track more closely with AI-answer citations than traditional backlinks do. We're citing that as one vendor's own stated finding about its own data, not an independently confirmed result, since neither that company nor any platform in this guide's research pass publishes the methodology behind it.
The action here isn't a tool; it's the slower work of being genuinely worth writing about: answering questions in the communities your customers already use, correcting a fact about you that's wrong in public, and making sure pages that already mention you state your current facts rather than a stale version.
What none of this can promise you
No step in this guide, done perfectly, guarantees a ranking, a citation, or a single visitor. Google says there are no special optimizations for its own systems; OpenAI states placement "is not guaranteed"; Bing states GEO "does not guarantee grounding or citations in AI experiences." The companies building these systems say the same thing this guide does, in their own words, on their own pages. Be skeptical of anyone, including a vendor whose product is AI-visibility software, who claims more certainty than that.
Verify it worked
Every check in this guide is one you can run yourself, without trusting this page.
- Fetch
yourdomain.com/robots.txtdirectly in a browser, then re-run one of the checkers from step 1, and confirm the crawlers you meant to allow show as allowed. - Run the page through Google's Rich Results Test and the schema.org validator, and confirm you haven't relied on FAQPage or HowTo as your only structured data.
- Use Search Console's URL Inspection tool to confirm the page is actually indexed, not just submitted.
- If you have Bing Webmaster Tools set up, read your own Citation Share, and remember it shows your share, not a competitor's.
- Ask ChatGPT, Perplexity, Copilot, a Google AI Overview, and Google AI Mode about your own brand by name, and log what each one says.
None of these require our word for anything. If a check fails, it tells you exactly which step to redo. If every check passes and you still aren't mentioned anywhere, that's a real, honest outcome this guide can't fix: getting cited is the one part of this subject even Bing, the platform documenting the most, will not guarantee.
<!-- Revision note: this draft was rewritten against Data/seo/briefs/ai-search-visibility-facts.md (compiled 2026-08-29, 30 CONFIRMED / 7 NOT DOCUMENTED findings, each fetched from a platform-owned URL). Quotes in the body are taken verbatim from that brief. The Google-Extended training/Gemini framing in step 3 is not itself covered by that brief (it contains no dedicated Google-Extended section); it is carried over from this guide's earlier draft and the coordinator's explicit restatement of it as accurate, not from a verbatim quote fetched in this pass. Anthropic/Claude agent names in step 1 are likewise outside the new brief's scope; see unconfirmedSteps. Primary sources used (verbatim quotes, all from the facts brief, all platform-owned URLs, read 2026-08-29): - developers.google.com/search/docs/appearance/ai-features - developers.google.com/search/docs/fundamentals/ai-optimization-guide - developers.google.com/search/docs/crawling-indexing/robots-meta-tag - developers.google.com/search/docs/appearance/structured-data/sd-policies - developers.google.com/search/blog/2023/08/howto-faq-changes - developers.google.com/search/updates (changelog) - developers.openai.com/api/docs/bots - help.openai.com/en/articles/9237897-chatgpt-search - help.openai.com/en/articles/12627856-publishers-and-developers-faq - docs.perplexity.ai/docs/resources/perplexity-crawlers - perplexity.ai/help-center (how Perplexity works; source labels) - bing.com/webmasters/help/webmaster-guidelines-30fba23a - blogs.bing.com/webmaster/February-2026/... (AI Performance launch) - blogs.bing.com/search/June-2026/... (Citation Share, Compare) - about.ads.microsoft.com/.../optimizing-your-content-for-inclusion-in-ai-search-answers - support.google.com/webmasters/answer/16984139 (generative AI performance report) - schema.org/docs/about.html · validator.schema.org · search.google.com/test/rich-results · sitemaps.org/protocol.html · indexnow.org · support.google.com/webmasters/answer/9012289 Sibling site reviews still used for tool-specific claims only (dateModified 2026-08-27): content/reviews/ismybrandinai.md, brandcited.md, aiclicks.md, hyperleap.md, xseek.md. -->Check it worked
Fetch your own robots.txt in a browser tab and re-run a free checker to confirm the crawlers you meant to allow show as allowed. Run the page through Google's Rich Results Test and the schema.org validator, and confirm you haven't relied on FAQPage or HowTo markup, both dead in Google Search. Use Search Console's URL Inspection tool to confirm the page is actually indexed. If you have Bing Webmaster Tools set up, read your own Citation Share. Then ask ChatGPT, Perplexity, Copilot, and Google, once through an AI Overview and once in AI Mode, about your own brand by name and read what comes back. Every check here is one you run yourself, on tools the vendors publish for free.
What we could not confirm
- The body text of Google's own blog post announcing the generative AI performance report (developers.google.com/search/blog/2026/06/gen-ai-performance-reports). The report's existence and behavior are confirmed directly from Google's Search Console Help page, but the announcement post itself returned only its archive/navigation shell during this research pass, not readable article text, so anything about that report beyond the Help Center page's own wording is not independently confirmed here.
- The exact date HowTo structured data was fully removed from Google Search. Google's 2023 post confirms HowTo was restricted to desktop-only that year, and the documentation page for it no longer resolves as of this read, consistent with full removal, but no Google Search Central changelog entry giving a specific final-removal date, the way Google dated FAQPage's removal to May 7, 2026, was found. Don't assert a specific HowTo removal date beyond what's stated here.
- Whether Perplexity's crawlers reliably honor robots.txt disallow rules in observed practice. Perplexity documents PerplexityBot as respecting robots.txt and Perplexity-User as generally ignoring it by design; independent reporting has previously disputed Perplexity's compliance in specific incidents, and this guide did not verify Perplexity's current practice first-hand.
- The precise functional split between Anthropic's Claude-User and Claude-SearchBot. This guide's primary research pass did not cover Anthropic at all; the agent names in step 1 rest on an earlier, separate check of Anthropic's crawler page that could not confirm which agent is scoped to live user fetches versus Claude's search feature specifically.
- The unlinked-mentions-versus-backlinks finding referenced in step 10 is a vendor's own stated claim about its own data. Neither that company nor any platform covered in this guide's research pass publishes the methodology behind it, and it is not treated as an independently verified result.
- Whether Google currently participates in IndexNow. Bing documents recommending it and Yandex is a known participant; Google's participation status was not addressed in this guide's research pass, so IndexNow is presented here as Bing's documented recommendation, not a way to reach Google specifically.