Skip to content
ToolShelf

Robots.txt tester

Paste your robots.txt, give it a URL, and find out whether a crawler may fetch it, which group it reads, which rule decided it, and what to add if the answer is wrong.

Your robots.txt

Paste the file rather than a link to it. A page in your browser cannot read another site’s robots.txt, so a tool that offered to fetch it would be sending your URLs through somebody’s server. Yours is at your-site.com/robots.txt.

Your robots.txt and the URL you test stays in this tab. The work is done by JavaScript in your browser. None of it is uploaded, logged or saved, and the tool keeps working with the network off.

What to test

A full address, or just the path from the slash onwards.

Crawls pages for Google Search. Checked against Google’s own documentation on 29 September 2026.

What you want to happen

Your robots.txt and the URL stays in this tab. The work is done by JavaScript in your browser. None of it is uploaded, logged or saved, and the tool keeps working with the network off.

Result

—

Paste your robots.txt and enter a URL.

A crawler reads one group, not the whole file

This is the rule that catches nearly everybody, and almost every “why is my page still blocked” question comes back to it.

A crawler picks exactly one group out of your robots.txt: the one whose User-agent name is the longest that matches it. Having chosen, it ignores every other group in the file completely.

So a file with a User-agent: * group and a User-agent: Googlebot group does not apply both to Googlebot. It applies the Googlebot group and nothing else. Every Disallow in the catch-all is invisible to Googlebot, which is occasionally what somebody wanted and much more often a site accidentally opened up.

The same rule is why adding a group for a crawler is riskier than it looks: the moment User-agent: GPTBot appears with one line under it, GPTBot stops obeying everything it was obeying in the catch-all. That is why the suggested fix on this page repeats those rules rather than only adding the new one.

Then the longest rule wins

Within that one group, every Allow and Disallow whose pattern matches the path is a candidate, and the one with the longest pattern wins. Length is counted as characters written, so /folder/public/ beats /folder/ and the page is allowed.

When two matching rules are exactly the same length, the less restrictive one wins, which means Allow. Two characters matter in the patterns: * stands for any run of characters, and $ at the end pins the pattern to the end of the path, so /*.pdf$ matches a PDF but not report.pdf.html.

Anything no rule matches is allowed. robots.txt is a list of exceptions to “yes”, not a list of permissions, which is why an empty file and a missing file mean the same thing.

Blocking is not hiding

Disallow stops a page being fetched. It does not stop the address being listed. Google can and does show a URL it has never crawled, with no description under it, when other pages link to it.

To keep a page out of search results you need a noindex instruction on the page itself, and here is the trap: Google can only see that by fetching the page. Block the page in robots.txt and the noindex is never read, so the page stays listed indefinitely. The two tools pull against each other, and using both is the way to get neither.

For anything that genuinely must not be read, robots.txt is the wrong instrument entirely. It is a public file that politely asks, and a crawler that means you harm will read it as a list of the interesting directories. Put a password on it.

The AI crawlers, and which rules actually bind

The names have multiplied, and they are not interchangeable. Each vendor now runs separate crawlers for separate jobs, and blocking one does not block the others.

Training, search and user requests are three decisions. OpenAI splits them into GPTBot, OAI-SearchBot and ChatGPT-User; Anthropic into ClaudeBot, Claude-SearchBot and Claude-User. Refusing to have your content used for training is a different choice from taking your site out of an assistant’s search results, and a single Disallow: / aimed at the wrong token does one when you meant the other.

Some names are not crawlers. Google-Extended and Applebot-Extended never fetch anything and will never appear in a server log. They are permission tokens: disallowing one withholds consent for a use of content that a different crawler already collected. Google states that Google-Extended does not affect how a page ranks in Search.

Some rules are requests rather than answers. The fetchers that run because a person asked about a specific page behave differently between vendors, and they say so themselves: Perplexity documents that Perplexity-User generally ignores robots.txt, while Anthropic states that Claude-User honours it. This tool says which, beside the result, because “blocked” means different things for different names.

Every token in the list here was taken from the vendor’s own documentation, with the date it was read, because a name is matched literally and one that is slightly wrong fails silently, exactly like a typo.

Questions

Why is my page still blocked when robots.txt clearly says Allow?
Almost always because the crawler is reading a different group from the one you edited. A crawler reads exactly one group: the one whose User-agent name is the longest that matches it. If your file has a Googlebot group and a * group, Googlebot reads its own group and ignores the * group completely, including every Allow in it. Paste your file here and the result says which group the crawler actually read.
Which rule wins when two of them match?
The longest pattern, counted as characters. Disallow: /folder/ and Allow: /folder/public/ both match /folder/public/page, and the Allow wins because it is longer. When two matching rules are exactly the same length, the less restrictive one wins, which means Allow. That is why the suggested fix sometimes adds a $ to the end: it makes the pattern one character longer and settles it.
Does robots.txt stop a page appearing in Google?
No, and this is the most expensive misunderstanding in the file. Disallow stops a page being crawled, not being indexed. Google can still list a URL it has never fetched if other pages link to it, showing the address with no description. To keep a page out of results you need a noindex meta tag on the page, which Google can only see if the page is crawlable. Blocking it in robots.txt guarantees the noindex is never read.
What is the difference between GPTBot, OAI-SearchBot and ChatGPT-User?
They do different jobs and take separate rules. GPTBot collects content that may be used to train OpenAI's models. OAI-SearchBot indexes pages so they can appear in ChatGPT's search results. ChatGPT-User fetches a page because a person asked about it. Blocking GPTBot refuses training and leaves you in ChatGPT's search results; blocking OAI-SearchBot takes you out of them. Anthropic splits its crawlers the same three ways, with ClaudeBot, Claude-SearchBot and Claude-User.
Why does the tool say Google-Extended is not a crawler?
Because it is not one. Nothing is ever fetched by Google-Extended and it will never appear in a server log. It is a permission token: disallowing it tells Google that content its other crawlers already collected may not be used to train Gemini or to ground answers. Google states it does not affect how a page ranks in Search. Applebot-Extended works the same way, sitting alongside Applebot, which does the actual fetching.
Will every crawler obey my robots.txt?
The well-behaved ones will, and it is a request rather than a lock in every case. Two kinds are worth knowing about. Vendors' user-initiated fetchers, which run because a person asked about a specific page, often ignore the file: Perplexity documents that Perplexity-User generally does, while Anthropic states that Claude-User honours it. And a crawler that is not well-behaved simply will not read the file at all. For anything that genuinely must not be fetched, use authentication rather than robots.txt.
Why can I not just give it my website address?
Because a page running in your browser is not allowed to read another site's files. That is the same-origin policy, and it is a security rule rather than a limitation worth working around. The way around it is to send your address to a server that fetches the file on your behalf, which would mean this tool quietly collecting the URLs you test. Opening your-site.com/robots.txt and pasting is one extra step and keeps everything here.
Does a blank line separate groups?
No. A new group starts the moment a User-agent line appears after a rule, whether or not there is a blank line in between. Blank lines are purely for readability. Consecutive User-agent lines with no rules between them share one group, which is how you write one set of rules for several crawlers at once.
What happens to a line with a typo?
A crawler ignores it and carries on, which is what makes typos here so expensive: the file looks right and the rule is simply not in force. Disallow /admin with no colon is the classic one. This tool lists every line no crawler would read, with the reason, rather than quietly parsing around them.

More tools