How to Get Your Website Cited by ChatGPT and Perplexity
ChatGPT answered this query from memory with zero citations. This 2026 checklist makes your site the source it quotes next.
TL;DR: Allow OAI-SearchBot and PerplexityBot in robots.txt, get indexed in Bing, and lead each page with a direct answer plus FAQPage schema.
ChatGPT returned zero grounded citations for this exact query and answered from memory, which means no single guide owns it yet. The gap is practical: most sites either block the wrong bots or bury the answer where retrieval cannot lift it. The checklist below closes both gaps in one pass, with the same audit tooling AutomateLab uses on client sites.
How does citation by ChatGPT and Perplexity actually work?
Answer engines retrieve passages and lift the most quotable span into the response. ChatGPT Search builds its candidate set from the Bing index using OAI-SearchBot, while Perplexity retrieves live web content at query time using PerplexityBot. A page missing from Bing is invisible to ChatGPT even with perfect prose, and a page blocking PerplexityBot never enters the Perplexity candidate set.
Only 11% of domains are cited by both engines, per the citation-overlap analysis quoted across 2026 GEO guides, so measure each engine separately. Perplexity typically surfaces a well-structured new page within days, ChatGPT Search within one to three weeks, and Google AI Overviews within four to eight weeks. The 13 signals AI assistants use to decide what to cite breaks down why structure beats authority at this stage: tables are extracted at 81% versus 23% for the same content in prose, FAQ sections are cited at 2.3 times the rate of pages without them, and 44.2% of all citations come from the first 30% of page text.

| ChatGPT Search | Perplexity | |
|---|---|---|
| Index bot | OAI-SearchBot | PerplexityBot |
| Index source | Bing index plus OAI-SearchBot crawl | Live retrieval plus PerplexityBot crawl |
| Typical pickup | One to three weeks | Days |
| Citation style | Fewer sources, conservative domain mix | More sources per answer, footnote links |
| Biggest blocker | Missing from Bing, OAI-SearchBot disallowed | PerplexityBot disallowed, stale date |
How to allow OAI-SearchBot and PerplexityBot in robots.txt?
Training crawlers and search crawlers are separate agents, and confusing them is the most common citation killer. OpenAI documents OAI-SearchBot as the agent that surfaces sites in ChatGPT search results and GPTBot as the training crawler, and Perplexity documents PerplexityBot as the agent that surfaces sites in Perplexity results. Allowing the search bots while disallowing the training bots keeps citation eligibility without opting content into model training.
User-agent: OAI-SearchBot
Allow: /
User-agent: PerplexityBot
Allow: /
User-agent: Claude-SearchBot
Allow: /
User-agent: GPTBot
Disallow: /Three details decide whether this actually works. First, check for a wildcard block above these lines, because a broad Disallow under User-agent star can override specific allows on some stacks. Second, check the edge layer, because Cloudflare AI-bot presets and WAF bot rules enforce before robots.txt is consulted and silently drop OAI-SearchBot and PerplexityBot even when the file looks correct. Third, submit the XML sitemap in Bing Webmaster Tools and confirm index coverage, because ChatGPT retrieves through the Bing index and an unindexed page gives it nothing to cite. The companion technical GEO checklist shows the full per-bot template this starter block comes from.
How to structure pages so answers can be extracted?
Retrieval lifts short self-contained spans, not pages. Open the page with a two to four sentence direct answer inside the first 100 words, then open every major section with a 40 to 60 word answer to that section question before any background. Phrase H2 and H3 headings as the questions readers type, name entities explicitly instead of leaning on pronouns, and keep each section readable without the sections around it.
Evidence density is the cheapest lever on the page. The Princeton GEO study found that adding statistics and adding cited sources each raised AI visibility by roughly 30 to 40%, and pages with sourced statistics are cited three to five times more often than pages with vague claims. Replace each generic claim with a dated number linked to its primary source on the noun phrase, render every comparison as an HTML table, render every procedure as a numbered list, and close with an FAQ block of three to six long-tail questions. Run the result through the free AI SEO checker, which scores answer placement, heading shape, and extractability across 39 checks in about 20 seconds.

How to add llms.txt, schema, and entity signals?
Three small files do outsized work once the prose is extractable. Publish an llms.txt index at the domain root listing the pages AI agents should read first, following the llms.txt specification, and keep it curated rather than auto-dumping the sitemap. Add JSON-LD for Article with author plus datePublished and dateModified, FAQPage built from the FAQ block, HowTo on procedural posts, and Organization with sameAs links to LinkedIn, GitHub, and Crunchbase so the entity resolves to one real organization. Keep a visible publish date and a genuine last-updated date in the body, because Perplexity weights recency and prefers refreshed pages over stale ones.
If this feels like a lot of moving parts, that is exactly what the service funnel is for. The AI-SEO MCP audit_page tool scores eight citation signals automatically and returns per-signal fixes, the AI SEO checker verifies the same signals from a browser with no signup, and the GPTBot check confirms crawler access in one click. When the audit flags structural work across dozens of pages, the AI SEO service ships robots.txt, llms.txt, schema, heading rewrites, and entity links as a fixed-scope engagement.
How do you test all seven checks in 20 minutes?
- Curl robots.txt and confirm OAI-SearchBot, PerplexityBot, and Claude-SearchBot return Allow while GPTBot stays disallowed.
- Confirm the page is indexed in Bing Webmaster Tools with no exclusion warnings.
- Read the first 100 words in isolation and confirm a stranger could answer the query from them.
- Count question-shaped H2s and confirm each section opens with a 40 to 60 word direct answer.
- Confirm one HTML table, one numbered list, dated statistics with primary-source links, and a visible update date.
- Validate FAQPage, Article, and Organization JSON-LD with the Rich Results Test before deploying.
- Ask ChatGPT and Perplexity five target queries weekly and log whether the domain appears as a cited source.
FAQ
How long until a fixed page gets cited by ChatGPT and Perplexity?
Perplexity typically picks up a well-structured page within days, ChatGPT Search within one to three weeks once Bing indexes it, and Google AI Overviews within four to eight weeks. Robots.txt changes take about 24 hours to propagate on each vendor side.
Do I need to rank on Google before AI engines cite me?
No. Only about 12% of cited URLs sit in the top 10 for the same query, and Perplexity is the most willing to cite smaller domains. Ranking helps discovery but extraction quality decides the citation.
Should I block GPTBot to stop AI training use?
Blocking GPTBot stops training use and does not remove the page from ChatGPT search answers, because search eligibility is governed by OAI-SearchBot. Block GPTBot, allow OAI-SearchBot, and keep the two decisions independent.
Do I need an llms.txt file to get cited?
No page requires it, but it is the fastest way to hand agents a curated reading list instead of letting them guess from navigation. Publish a short curated file, link it from key pages, and keep it updated when the site map changes.
How do I check which AI bots actually reach my site?
Filter 30 days of server logs for OAI-SearchBot, ChatGPT-User, PerplexityBot, Perplexity-User, ClaudeBot, and Claude-SearchBot, verify Bing index coverage, and run the free AI SEO checker plus the GPTBot check for a second opinion.