How a site gets into ChatGPT, Claude and Perplexity answers
What to check so AI search reads your site and names it as a source: robots.txt, text without JavaScript, owner markup.
Checked against sources: 30 September 2026

- 4 botsGPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot
- HTMLtext must be in the markup, no scripts
- 1 filerobots.txt decides who gets in
Let the bots in
AI services run their own crawlers: GPTBot and OAI-SearchBot at OpenAI, ClaudeBot at Anthropic, PerplexityBot at Perplexity. A blanket disallow in robots.txt means you will not be read or cited.
Training and search can be separated: block GPTBot but allow OAI-SearchBot. That is your choice, and Awe Check reports it as it stands.
Text must be in the HTML
Many bots do not run JavaScript. If a script paints your article, the bot sees an empty page. The test is simple: open View Source and look for your paragraph in the markup.
Say who you are
Models cite what can be verified. A recognizable owner makes a site that kind of source.
- Organization markup with a name, site URL and logo.
- An About page and contacts: email and legal details.
- Articles show an author plus published and updated dates.
About llms.txt
The llms.txt file is optional. Google Search does not use it and it guarantees no citations. Add one if you want to hint at your site structure to models, but do not treat it as a substitute for good text.
Frequently asked
Can I block training and still appear in answers?
Yes. Block training crawlers and leave search crawlers allowed. The companies publish bot names and purposes themselves.
How do I know I am being cited?
Ask ChatGPT, Perplexity and Claude a question on your topic with search on and look at the sources. Once a month is enough.
