Articles

How Do I Get My Business Cited by ChatGPT and Perplexity?

Most published advice on this is wrong. Blocking GPTBot doesn't stop ChatGPT citing you. No engine reads llms.txt. And the only controlled study of schema markup found no benefit at all. Three things do work, in this order: let the right bots in, rank for the question people actually type, then put something on the page the engine can't say without you.

How Do I Get My Business Cited by ChatGPT and Perplexity?
The short answer

Three things, in this order. Let the right bots in. Rank for the question people actually type. Then put something on the page the engine can't say without you. The first two are a checklist and you can finish them this week. The third decides it, and no tool can do it for you.

Have you blocked GPTBot to keep ChatGPT out of your content? Published an llms.txt file? Added schema markup so AI would cite you?

None of those do what you think. Most of what’s published on this is wrong in specific, checkable ways, so three corrections come first. If you’ve acted on any of them, you spent that effort on nothing.

Blocking GPTBot does not stop ChatGPT citing you. OpenAI runs separate crawlers for separate jobs. GPTBot trains models. OAI-SearchBot builds the index behind ChatGPT’s search features. Publishers who blocked GPTBot to protect their content from training, then wondered why they still appeared in ChatGPT search, blocked the wrong bot. The reverse error is more common and more expensive: allowing GPTBot in the belief it earns citations, when OAI-SearchBot is what governs that.

Blocking Google-Extended does nothing to AI Overviews. Google-Extended governs Gemini app training and grounding. AI Overviews and AI Mode run on Googlebot, and are controlled with nosnippet, data-nosnippet, max-snippet or noindex. Google states this in its own documentation.

No engine reads llms.txt. It’s a September 2024 proposal. As of August 2026, no major engine has confirmed honoring it. Google’s optimization guide says explicitly that you don’t need to create AI text files, and John Mueller has called it speculative, pointing out the file has existed for years with no AI system using it.

Step 1: Check your CDN before you check your robots.txt

Check your CDN’s bot settings first. Then your robots.txt.

Why? Because firewall and bot rules are enforced before robots.txt is read. A CDN blocking AI crawlers overrides any Allow you write.

Cloudflare started asking new domains at sign-up whether to permit AI crawlers in July 2025. And from September 2026 it will block training and agent crawlers by default on ad-displaying pages for new domains.

Read their docs and it gets worse. Defaults are enforced by the most restrictive rule that applies. And blocking “training” can catch multi-purpose crawlers, including Googlebot and Bingbot.

So you can get the robots.txt perfect and still be invisible.

Step 2: Get the robots.txt right

Two mechanics that most published sample files get wrong.

Crawlers obey only the single most specific matching group. Per Google’s specification, user-agent-specific groups and the global * group are not combined. If you write User-agent: * with Disallow: /private/, then add a permissive group for OAI-SearchBot, that bot no longer inherits the /private/ rule. You have to restate it.

Noindex is not a valid robots.txt directive. It’s ignored. Use the meta robots tag.

The crawlers worth knowing, by job:

BotOperatorWhat it does
GPTBotOpenAITrains foundation models
OAI-SearchBotOpenAIIndexes for ChatGPT search. This is the citation gate
ChatGPT-UserOpenAIUser-initiated fetches, not automatic crawling
PerplexityBotPerplexityIndexes for Perplexity; not used for model training
Perplexity-UserPerplexityUser-initiated; documented as generally ignoring robots.txt
ClaudeBotAnthropicTrains models
Claude-SearchBotAnthropicSearch quality
Google-ExtendedGoogleGemini training and grounding. Not AI Overviews
GooglebotGoogleSearch, including AI Overviews and AI Mode

One caution about relying on any of this: Cloudflare reported in August 2025 that Perplexity used undeclared crawlers impersonating Chrome, with rotating IPs, to retrieve content from domains whose robots.txt disallowed all automated access. ChatGPT-User complied under identical testing. Perplexity disputed the finding. Cloudflare sells bot blocking, so read it with that in mind. The test used brand-new never-linked domains, which is hard to explain away. Robots.txt compliance is operator-dependent.

Step 3: Rank for the question

Before ChatGPT can name you, it has to pick your page as one of the handful it reads. Rank high in regular search and it usually picks you. Rank low and it doesn’t.

AirOps and Kevin Indig ran 16,851 queries. Pages sitting at position one got cited 58% of the time. Pages at position ten, 14%.

seoClarity looked at 362,000 queries that trigger an AI Overview. 94% cite at least one URL from the top 20 organic results.

But 44% of individual citations come from outside the top 20. So ranking is a strong gate, not a locked door.

For ChatGPT specifically, Seer Interactive found 87% of SearchGPT citations matched Bing’s top results, against 56% for Google. Small sample, roughly 100 queries. Treat it as directional.

Two corrections while we’re here.

Perplexity runs its own index. It documents over 200 billion URLs and says it has moved off commercial search APIs. So it’s no longer a Bing wrapper. Get PerplexityBot in, and don’t worry about Bing.

ChatGPT search runs on OpenAI’s own indexer, OAI-SearchBot. Bing results still correlate closely, but Bing isn’t what’s fetching you.

So old-fashioned SEO still does most of the work here. One thing to add: render your content on the server. OpenAI’s and Anthropic’s crawlers read raw HTML and don’t run JavaScript. Google renders, and Perplexity rendered in late-2025 tests, but why gamble?

Step 4: Put citable specifics on the page

The honest answer includes what doesn’t work. So let’s start there.

Schema markup: no measurable citation benefit. Ahrefs ran the only properly controlled study I could find — 1,885 pages that added JSON-LD, roughly 4,000 matched control pages, difference-in-differences. Result: −4.6% for Google AI Overviews (statistically significant), +2.4% AI Mode and +2.2% ChatGPT (both indistinguishable from zero). Separate live-fetch testing found no system extracted data that existed only in JSON-LD. Visible HTML worked. Google’s guide states no special markup is needed.

That null result is worth trusting partly because Ahrefs doesn’t sell schema tooling. The finding runs against the grain of the category it operates in. Keep schema for traditional rich results. Stop expecting it to earn AI citations.

Content specifics: the best-supported lever. Work presented at KDD 2024 benchmarked nine optimization methods across ~10,000 queries in 25 domains. The winners: adding statistics (+25.9%), quotations (+27.8%), and citations to sources (+24.9%). Keyword stuffing performed worst of the tested set and reduced visibility elsewhere.

Two caveats the paper’s own method demands. It was mostly tested against a built engine resembling Bing Chat rather than production systems, with a 200-sample check against live Perplexity. And it’s nearly two years old, which in this area is a long time.

What do those three winners have in common? A number, a quote, an attributed source. All things a model cannot generate and cannot work out from the category. An independent 2026 academic study of 21,143 citations found the same shape. Highly-cited pages are long, well-structured, on-topic, and rich in definitions and numbers.

Being talked about elsewhere. Ahrefs’ analysis of 75,000 brands found YouTube mentions correlating with AI visibility at roughly 0.74 and branded web mentions around 0.66, against backlinks at 0.25 to 0.30. Their own stated caveat, which I’ll repeat rather than bury: correlation isn’t causation, and improving these metrics won’t automatically lift visibility. It’s also vendor research measuring the thing the vendor sells.

Updating beats publishing. Seer’s analysis of 7,683 dated pages carrying 47,097 citations found 72% of cited pages look fresh by last-modified date, while only 42% were actually published recently. Over a quarter of “fresh” cited pages first went live two or more years ago. The freshness engines reward is largely manufactured by maintaining what already works.

Step 5: Give the engine something it can’t say without you

Everything above gets you eligible. This part decides it.

Google’s own guide to optimizing for generative AI features tells site owners to skip AI-specific files and markup. Then it says what to do instead: don’t just recycle what others have already said, or what could easily be produced by a generative AI model.

That’s the engine describing its own selection rule. A system assembling an answer already holds the category consensus. It has a reason to fetch and name a source when that source carries something the consensus doesn’t. A number nobody else has, a definition with an author, a named method, a position stated crisply enough to quote.

Which is why the mechanics have a ceiling. They make you retrievable and legible. They cannot make you worth retrieving. A perfectly optimized page saying what forty competitors say is a well-formatted redundancy, and the systems deciding these answers are built to skip redundancy. We make that argument at length in Being Generic Was a Brand Problem, and the underlying market condition in the sea of sameness.

The checklist

  1. CDN bot settings. Confirm AI crawlers aren’t blocked at the firewall. This overrides robots.txt.
  2. robots.txt. Allow OAI-SearchBot, PerplexityBot, Googlebot, Claude-SearchBot. Restate rules inside each named group, because they don’t inherit from *.
  3. Server-side render anything you want read.
  4. Rank for the question. Retrieval position dominates. Classic SEO is the gate.
  5. Put citable specifics on the page. Numbers, named sources, direct quotes, clear definitions.
  6. Earn mentions elsewhere. Third-party coverage correlates far better than backlinks.
  7. Maintain rather than multiply. Update the pages that already earn citations.
  8. Skip llms.txt. Skip schema-for-AI. Neither is supported by current evidence.
  9. Have something to say that the model can’t say without you. Steps one through eight multiply this. On its own, it’s what they multiply.

Step nine has a name. Your One Unforgettable Idea is the single idea a founder owns so completely that they stop being compared and become their own category. The other eight steps are plumbing for it.

Sources current as of August 2026. This area changes quickly, and several claims here contradict advice that was accurate in 2024. Where a finding rests on a single study or a small sample, I say so inline. Where something is unconfirmed, I say that too.

Sources

  1. OpenAI. Crawler documentation (GPTBot, OAI-SearchBot, ChatGPT-User). developers.openai.com
  2. Perplexity. Bot documentation (PerplexityBot, Perplexity-User). docs.perplexity.ai
  3. Anthropic. Crawler documentation (ClaudeBot, Claude-User, Claude-SearchBot).
  4. Google Search Central. "Guide to Optimizing for Generative AI Features," updated July 2026; AI features documentation, updated December 2025; robots.txt specification.
  5. Howard, Jeremy. llms.txt proposal, September 2024. Mueller comment reported by Search Engine Journal, June 2026.
  6. Linehan, Louise and Xibeijia Guan (Ahrefs). "Does Schema Markup Affect AI Citations?" May 2026. 1,885 treated pages, ~4,000 matched controls, difference-in-differences.
  7. searchVIU. "Schema Markup and AI in 2025." October 2025. Single test page across eight markup scenarios and five systems.
  8. Aggarwal, Pranjal et al. "GEO: Generative Engine Optimization." ACM SIGKDD 2024, arXiv:2311.09735. ~10,000 queries, 25 domains; metric is position-adjusted word count.
  9. Zhang, Kai, Xinyue He and Jingang Yao. geo-citation-lab, arXiv:2604.25707, April 2026. 602 queries, 21,143 citations, 18,151 pages.
  10. Ahrefs. "AI Brand Visibility Correlations," December 2025. 75,000 brands; Spearman correlations; vendor research measuring its own product's domain.
  11. seoClarity. AI Overview rankings overlap, October 2025. 362,000 US desktop queries.
  12. AirOps with Kevin Indig. "The Fan-Out Effect," April 2026. 16,851 queries, 815,484 scoring rows.
  13. Seer Interactive. SearchGPT/Bing citation overlap, April 2026 (~100 queries); content recency study, 2026 (7,683 pages, 47,097 citations).
  14. Perplexity. "Architecting and Evaluating an AI-First Search API," July 2026. 200B+ URL index.
  15. Cloudflare. "Content Independence Day," July 2025; "Your site, your rules," July 2026; stealth crawler report, August 2025. Perplexity disputes the last of these.
  16. Vercel. "The Rise of the AI Crawler," December 2024. JavaScript execution by crawler.

Questions people ask

Does blocking GPTBot stop ChatGPT from citing my site?
No. OpenAI runs separate crawlers for separate jobs. GPTBot trains foundation models. OAI-SearchBot builds the index that surfaces websites in ChatGPT's search features. Blocking GPTBot has no effect on whether ChatGPT search cites you. Blocking OAI-SearchBot does. These two are constantly conflated in published advice, in both directions.

Does adding schema markup get you cited by AI?
The available controlled evidence says no. Ahrefs studied 1,885 pages that added JSON-LD against roughly 4,000 matched control pages and found no citation benefit. Slightly negative for Google AI Overviews, statistically indistinguishable from zero for AI Mode and ChatGPT. Google's own optimization guide states there is no special schema.org markup needed for generative AI search. Schema remains worthwhile for traditional rich results; the AI-citation claim is unsupported.

Does llms.txt work?
There is no evidence any major engine reads it. llms.txt is a proposal published by Jeremy Howard in September 2024. As of August 2026, OpenAI, Google, Anthropic and Perplexity have not confirmed honoring it, Google's optimization guide says you don't need to create AI text files, and Google's John Mueller has described it as speculative, noting the file has existed for years without AI systems using it.

Does traditional SEO still matter for AI citations?
Yes, as a gate rather than a guarantee. Analysis of 362,000 AI Overview-triggering queries found that 94% cite at least one URL from the top 20 organic results, though 44% of individual citations come from outside the top 20. Being retrieved is the binding constraint: one study of 16,851 queries found a 58% citation rate at retrieval position one versus 14% at position ten. Ranking gets you considered. It does not decide who gets named.

The fastest way to find out what only you can say: the free Viral Genius Profile — 12 questions, about 15 minutes, spoken out loud.

Start talking — free, 12 questions, ~15 min

Know a founder sitting on a gold mine he isn't exploiting? Send him this article.