Getting cited by ChatGPT, Gemini and Perplexity
AI assistants such as ChatGPT, Gemini and Perplexity now answer questions with links to the pages they used. Each has its own crawlers and its own rules. This article covers how they find pages, which crawlers you should allow, and what makes a page worth citing, using only what each company documents.
On this page — 5 sections
How do AI assistants find web pages?
Quick answer
Through their own crawlers, search indexes they build or license, and live fetches when a user asks. ChatGPT search uses OpenAI's crawler plus third-party search providers. Gemini grounds answers in Google Search. Perplexity runs its own crawler and index.
When you ask an assistant a question that needs fresh information, it searches the web, reads a handful of pages and writes an answer with citations. The pages it can read depend on its sources. OpenAI says ChatGPT search draws on its own crawler and third-party search providers. Google grounds Gemini answers in Google Search. Perplexity documents its own crawler, PerplexityBot.
The practical result: a page that is well indexed in Google and Bing, and not blocked to the assistants' crawlers, is a candidate everywhere. A real estate agency in Dubai that blocks every unfamiliar bot "to be safe" may be invisible to half of these answers.
Which crawlers should you allow in robots.txt?
Quick answer
Separate search crawlers from training crawlers. OAI-SearchBot lets pages appear in ChatGPT search; GPTBot is for model training. PerplexityBot indexes for Perplexity answers. Google-Extended controls Gemini training and grounding, not Google Search. Allow the search ones if you want citations.
Each company documents several user agents with different jobs. Blocking a training crawler does not remove you from search answers, and the reverse is also true. Decide each one on purpose:
| User agent | Company | Documented purpose |
|---|---|---|
| OAI-SearchBot | OpenAI | Surfacing and linking pages in ChatGPT search. OpenAI says sites that block it will not be shown in ChatGPT search answers |
| GPTBot | OpenAI | Collecting content that may be used to train models |
| ChatGPT-User | OpenAI | Fetching a page when a user asks ChatGPT to visit it |
| PerplexityBot | Perplexity | Indexing pages to surface and link in Perplexity answers |
| Google-Extended | Use of content for Gemini training and grounding; does not affect Google Search |
A sensible default for a business that wants to be cited: allow the search crawlers, then decide separately on training. Also check your CDN or firewall. Some security settings block AI crawlers by default, whatever robots.txt says.
What kind of page gets cited?
Quick answer
Pages that answer one question directly and specifically, near the top, with facts the assistant can quote. Assistants cite the passage that supports a sentence in their answer, so a clear definition, a number, or a step list is easier to cite than a long story.
An assistant writes its answer, then attaches sources to the claims. The easiest page to attach is one where a single passage supports a single claim. "Tinting car windows in Saudi Arabia is allowed up to a set percentage on side windows" is citable if your page states the rule and its source plainly. A page that buries it in paragraph six is not.
- Question as heading, answer first: the passage under the heading should stand on its own.
- Specific facts: numbers, dates, names and conditions, with sources for anything official.
- Information gain: something other pages do not have, such as your own prices, cases or test results.
- Clean, readable HTML: text in the page, not only in images or behind scripts that crawlers may not run.
Do assistants cite Arabic pages?
Quick answer
Yes, when the question is asked in Arabic and good Arabic pages exist. In many Gulf and Egyptian niches, few Arabic pages answer questions clearly, so a well-structured Arabic page faces less competition for citations than its English equivalent.
Assistants generally answer in the user's language and prefer sources in that language when they are good enough. This is an opportunity. Ask an assistant in Arabic about end-of-service pay in Saudi Arabia or rental contracts in Egypt, and you often see the same few sources, sometimes translated from English.
Write the Arabic page natively rather than translating it. Use the words people actually search with, including Gulf or Egyptian terms where they differ, and keep Modern Standard Arabic for the explanation. An assistant can only cite what is on the page.
How do you track whether assistants cite you?
Quick answer
There is no complete report. Check referral traffic from chatgpt.com, perplexity.ai and gemini.google.com in your analytics, run a fixed list of real questions in each assistant every month, and record which pages are cited. Watch the trend, not single answers.
Assistant answers vary from one run to the next, so a single check proves little. A simple routine works better: pick 20 to 30 questions your customers really ask, run them in each assistant on the same day each month, and note which sources appear. Pair that with referral data in your analytics tool.
This article is part of the Semantic SEO series — Writing for meaning: entities, definitions, attribute-value facts, semantic distance and pages that actually answer.
About the author
Mohamed Youns
Semantic SEO Engineer · Author & system developer
Mohamed Youns writes about how search engines understand content — the same standards he applies when building semantic systems at Nut Hub. nut-hub.org