Skip to content

Robots.txt Checker & Validator

Short answer

Check your robots.txt: its rules, the sitemap line, whether a page may be crawled, and which AI crawlers may read it.

Free · No signup · Fetched by our server
We fetch the URL on our server to run the check. Results aren't saved.

What It Does

Fetches your `/robots.txt` and parses every directive — User-agent groups, Allow/Disallow rules, Sitemap and Crawl-delay. Checks whether the URL you entered is allowed, using the same longest-match precedence real crawlers use, and flags a blanket `Disallow: /`, a missing sitemap directive and an oversized file. It also lists the main AI crawlers and control tokens, shows whether each may fetch that URL and explains what each one is for: training AI models, AI search answers, or opening a page when a user asks.

Why It Matters

A single stray `Disallow: /` can stop search engines from crawling a whole site, and it is easy to ship by accident from a staging setup. AI crawlers add a second question: blocking a training crawler does not remove you from AI search answers, but blocking a search crawler can — so it pays to know which token does what.

How It Works

  1. Enter the URL you want to test

  2. We fetch `/robots.txt` and parse every directive

  3. Check whether that URL is allowed for crawlers in general and for each AI crawler, using the precedence rules real crawlers use

  4. Flag a full-site block, a blocked URL and a missing sitemap directive, and explain each AI crawler

Sample input + output

INPUT
url: https://rankproof.eu/tools
OUTPUT
https://rankproof.eu/robots.txt: 200 OK
file size: 412 B · user-agent groups: 2 · crawl-delay: not set
this page is crawlable                     OK
sitemap directive present                  OK

AI crawlers — all may read this page:
  GPTBot            Allowed   OpenAI: collects pages to train AI models.
  OAI-SearchBot     Allowed   OpenAI: indexes pages so its AI search can show and cite them.
  ChatGPT-User      Allowed   OpenAI: opens a page when a user asks the assistant about it.
  Google-Extended   Allowed   Google: not a separate crawler — a rule for AI use of pages it crawls.
  …

Who Uses This

  • Technical SEO

    Before launch, check the production robots.txt so a staging "block everything" rule never goes live.

  • Site Owner

    Decide which AI crawlers to allow: keep AI search crawlers open while choosing whether your pages may train AI models.

  • Content Manager

    When a section disappears from Google Search Console, first check whether robots.txt blocks it.

Frequently Asked Questions

What is robots.txt?

A plain-text file at your domain root that tells crawlers which URLs they may or may not request. It is standardized in RFC 9309 (2022).

Can I block AI crawlers separately?

Yes — each has its own user-agent token, such as `GPTBot` or `ClaudeBot`. Several AI companies run separate crawlers for model training and for search, so you can block training while staying visible in their AI search answers.

What is Google-Extended?

A control token, not a separate crawler. Google crawls with its usual crawlers either way; disallowing `Google-Extended` tells Google not to use your content to train Gemini models or to ground answers in Gemini apps. It does not affect Google Search or rankings.

Do all AI crawlers follow robots.txt?

The major AI companies say their crawlers do. Some fetchers that act on a user's request, such as ChatGPT-User and Perplexity-User, may not apply robots.txt rules, and the report marks them.

What is the difference between robots.txt and noindex?

robots.txt prevents crawling. `noindex` allows crawling but prevents indexing. If you block a page in robots.txt, Google cannot see its noindex tag — a common mistake.

Should I include a sitemap directive?

Yes — `Sitemap: https://example.com/sitemap.xml` helps every crawler find your sitemap, including Bing and other search engines.

Does Google always respect robots.txt?

For crawling, yes. But a blocked URL can still appear in search results if other pages link to it — without a snippet. Use noindex to keep a page out.