Tools Free · SEO and AI search

Decide who crawls your site.

Build a robots.txt that lets AI search engines cite you while keeping training bots out, if you want. Then turn a list of URLs into a valid sitemap.xml.

Optional. Google ignores it.

AI crawlers

  • GPTBot OpenAI Training

    Collects pages to train OpenAI models.

  • OAI-SearchBot OpenAI Search

    Indexes pages for ChatGPT search results and citations.

  • ChatGPT-User OpenAI User fetch

    Opens a page when a ChatGPT user asks about it.

  • ClaudeBot Anthropic Training

    Collects pages to train Claude models.

  • Claude-User Anthropic User fetch

    Opens a page when a Claude user asks about it.

  • Claude-SearchBot Anthropic Search

    Indexes pages for Claude search answers.

  • PerplexityBot Perplexity Search

    Indexes pages for Perplexity answers and citations.

  • Perplexity-User Perplexity User fetch

    Opens a page when a Perplexity user asks about it.

  • Google-Extended Google Training

    Controls Gemini training. Does not affect Google Search or AI Overviews.

  • Applebot-Extended Apple Training

    Controls Apple AI training. Siri and Spotlight search still work.

  • CCBot Common Crawl Training

    Open web archive used in many AI training sets.

  • Bytespider ByteDance Training

    Collects pages for ByteDance models.

  • Meta-ExternalAgent Meta Training

    Collects pages to train Meta AI models.

  • Amazonbot Amazon Training

    Crawls for Alexa answers and Amazon AI models.

robots.txt

User-agent: *
Disallow: /admin/
Disallow: /api/

# AI crawlers allowed
User-agent: OAI-SearchBot
User-agent: ChatGPT-User
User-agent: Claude-User
User-agent: Claude-SearchBot
User-agent: PerplexityBot
User-agent: Perplexity-User
Allow: /
Disallow: /admin/
Disallow: /api/

# AI crawlers blocked
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: CCBot
User-agent: Bytespider
User-agent: Meta-ExternalAgent
User-agent: Amazonbot
Disallow: /

Sitemap: https://example.com/sitemap.xml

Upload to your site root so it loads at /robots.txt. Blocking search bots removes you from ChatGPT, Claude and Perplexity citations.

Get new tools first

Optional. New free tools and growth notes, roughly once a month.

  1. Set your default policy and the paths to keep private.
  2. Pick a preset or toggle each AI crawler on or off.
  3. Copy or download robots.txt, then switch tabs to build sitemap.xml.

Want this done for you? I run PR and growth for AI and Web3 founders. Book a free 30-minute teardown →

Questions founders ask

Should I block GPTBot in robots.txt?

GPTBot only collects training data. Blocking it does not remove you from ChatGPT search, which uses OAI-SearchBot and ChatGPT-User. Many brands block GPTBot and allow the search bots so they still get cited.

Does Google-Extended affect my Google rankings?

No. Google-Extended only controls whether your content trains Gemini models. Google Search, including AI Overviews, still uses Googlebot, so blocking Google-Extended has no effect on rankings.

How many URLs can a sitemap hold?

One sitemap file can hold up to 50,000 URLs and 50 MB uncompressed. Larger sites split URLs across several sitemaps and list them in a sitemap index file. Submit the sitemap in Google Search Console and reference it in robots.txt.