Robots.txt tester
Paste your robots.txt, give it a URL, and find out whether a crawler may fetch it, which group it reads, which rule decided it, and what to add if the answer is wrong.
Your robots.txt
Paste the file rather than a link to it. A page in your browser cannot read another site’s robots.txt, so a tool that offered to fetch it would be sending your URLs through somebody’s server. Yours is at your-site.com/robots.txt.
Your robots.txt and the URL you test stays in this tab. The work is done by JavaScript in your browser. None of it is uploaded, logged or saved, and the tool keeps working with the network off.
What to test
A full address, or just the path from the slash onwards.
Crawls pages for Google Search. Checked against Google’s own documentation on 29 September 2026.
Your robots.txt and the URL stays in this tab. The work is done by JavaScript in your browser. None of it is uploaded, logged or saved, and the tool keeps working with the network off.
Result
—
Paste your robots.txt and enter a URL.
A crawler reads one group, not the whole file
This is the rule that catches nearly everybody, and almost every “why is my page still blocked” question comes back to it.
A crawler picks exactly one group out of your robots.txt: the one whose User-agent name is the longest that matches it. Having chosen, it ignores every other group in the file completely.
So a file with a User-agent: * group and a User-agent: Googlebot group does not apply both to Googlebot. It applies the Googlebot group and nothing else. Every Disallow in the catch-all is invisible to Googlebot, which is occasionally what somebody wanted and much more often a site accidentally opened up.
The same rule is why adding a group for a crawler is riskier than it looks: the moment User-agent: GPTBot appears with one line under it, GPTBot stops obeying everything it was obeying in the catch-all. That is why the suggested fix on this page repeats those rules rather than only adding the new one.
Then the longest rule wins
Within that one group, every Allow and Disallow whose pattern matches the path is a candidate, and the one with the longest pattern wins. Length is counted as characters written, so /folder/public/ beats /folder/ and the page is allowed.
When two matching rules are exactly the same length, the less restrictive one wins, which means Allow. Two characters matter in the patterns: * stands for any run of characters, and $ at the end pins the pattern to the end of the path, so /*.pdf$ matches a PDF but not report.pdf.html.
Anything no rule matches is allowed. robots.txt is a list of exceptions to “yes”, not a list of permissions, which is why an empty file and a missing file mean the same thing.
Blocking is not hiding
Disallow stops a page being fetched. It does not stop the address being listed. Google can and does show a URL it has never crawled, with no description under it, when other pages link to it.
To keep a page out of search results you need a noindex instruction on the page itself, and here is the trap: Google can only see that by fetching the page. Block the page in robots.txt and the noindex is never read, so the page stays listed indefinitely. The two tools pull against each other, and using both is the way to get neither.
For anything that genuinely must not be read, robots.txt is the wrong instrument entirely. It is a public file that politely asks, and a crawler that means you harm will read it as a list of the interesting directories. Put a password on it.
The AI crawlers, and which rules actually bind
The names have multiplied, and they are not interchangeable. Each vendor now runs separate crawlers for separate jobs, and blocking one does not block the others.
Training, search and user requests are three decisions. OpenAI splits them into GPTBot, OAI-SearchBot and ChatGPT-User; Anthropic into ClaudeBot, Claude-SearchBot and Claude-User. Refusing to have your content used for training is a different choice from taking your site out of an assistant’s search results, and a single Disallow: / aimed at the wrong token does one when you meant the other.
Some names are not crawlers. Google-Extended and Applebot-Extended never fetch anything and will never appear in a server log. They are permission tokens: disallowing one withholds consent for a use of content that a different crawler already collected. Google states that Google-Extended does not affect how a page ranks in Search.
Some rules are requests rather than answers. The fetchers that run because a person asked about a specific page behave differently between vendors, and they say so themselves: Perplexity documents that Perplexity-User generally ignores robots.txt, while Anthropic states that Claude-User honours it. This tool says which, beside the result, because “blocked” means different things for different names.
Every token in the list here was taken from the vendor’s own documentation, with the date it was read, because a name is matched literally and one that is slightly wrong fails silently, exactly like a typo.
Questions
Why is my page still blocked when robots.txt clearly says Allow?
Which rule wins when two of them match?
Does robots.txt stop a page appearing in Google?
What is the difference between GPTBot, OAI-SearchBot and ChatGPT-User?
Why does the tool say Google-Extended is not a crawler?
Will every crawler obey my robots.txt?
Why can I not just give it my website address?
Does a blank line separate groups?
What happens to a line with a typo?
More tools
JSON formatter & validator
Indent it, shrink it, or find out what is wrong with it
Base64 encoder / decoder
Both directions, both alphabets, Unicode included
URL encoder / decoder
Percent-encoding, with the component and whole-URL rules apart