Robots.txt Checker for AI Crawlers
Validate your robots.txt syntax and see which group each AI crawler actually obeys: GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-User, PerplexityBot and more. Paste a draft or fetch the live file, and catch the named-group mistake that opens your admin pages to every crawler you welcomed.
Paste your robots.txt or fetch it by URL to validate + check AI crawler access.
A robots.txt checker parses your file into User-agent groups, tests each Allow and Disallow line for syntax, confirms a Sitemap directive and reports which crawlers can reach your site. This one also resolves the rules for 14 AI agents, including GPTBot, OAI-SearchBot, ClaudeBot and PerplexityBot, so you see who is blocked before an answer engine does.
How the robots.txt checker works
The tool is a parser and a rule resolver. It reads the file the way a crawler does, builds one group per User-agent block, then asks the question that matters for AI visibility: for each named AI agent, which single group applies, and does it open or close the site? Matching follows the precedence Google documents: an exact, case-insensitive name match wins, the wildcard group applies otherwise, and with neither the agent is allowed everywhere. Consecutive User-agent lines share one group; any other directive between them starts a new one.
- 1
Paste the file or fetch it by URL
Paste mode runs in your browser and nothing leaves your machine, so drafts are safe. Fetch mode takes a domain, appends /robots.txt and retrieves the public file through our server, since a browser cannot read another domain’s file directly. Prefer the fetch: what the CDN serves is what crawlers read.
- 2
Click Validate
The checker strips comments, splits the file into User-agent groups, records every Allow, Disallow, Crawl-delay and Sitemap line, and flags anything it cannot parse.
- 3
Read the syntax verdict first
Only three things are errors: an empty file, no User-agent group, and a Sitemap that is not an absolute URL. Everything else is a warning or a note, and the file still works.
- 4
Read the AI crawler panel
Each of 14 agents gets ALLOWED, PARTIAL or BLOCKED, with a note naming the group that decided: an explicit rule for that agent, or the wildcard. PARTIAL means some paths are closed but the root is open, the normal shape for a well-run site.
- 5
Fix the group, redeploy, re-check the live URL
Edit the group named in the note, not the wildcard, because a named group replaces the wildcard for its agent. After deploy, fetch the live file again in case a CDN prepended rules, then confirm Googlebot’s view in Search Console’s robots.txt report.
Every check this robots.txt validator runs
The exact checks, with the severity each carries. Any error fails the file. Warnings and notes never change the verdict, but each is a real gap a crawler will hit.
| Check | Severity | Why it matters | How to fix |
|---|---|---|---|
| File is empty | Error | Legal, and means allow everything, but nobody pastes an empty file on purpose. Usually the host is serving a blank response at /robots.txt. | Open the live URL and fix the hosting rule before editing contents. |
| Starts with a UTF-8 byte-order mark | Warning | Three invisible bytes before the first character. Some parsers discard the first directive, so the first User-agent line is lost. | Save as UTF-8 without BOM. Windows editors and some CMS exports add it silently. |
| Line has no ":" separator | Warning | Every directive is field, colon, value. A line without a colon is skipped, so whatever you meant it to do does not happen. | Add the colon. Rules pasted from chat or documents often lose it. |
| No User-agent group at all | Error | Without a User-agent line no rule belongs to anyone. Leading Allow and Disallow lines are rescued into an implicit wildcard group, so this only fires when there are no rules of either kind. | Start the file with "User-agent: *" and put rules under it. |
| Sitemap value is not an absolute URL | Error | A relative path such as /sitemap.xml is not resolved against your domain by every crawler and is dropped by the ones that do not. | Write "Sitemap: https://yourdomain.com/sitemap.xml", one line per sitemap. |
| No Sitemap directive | Warning | robots.txt is the one place every crawler looks, including AI agents with no Search Console equivalent to submit a sitemap to. | Add a Sitemap line at the end of the file, where it cannot split a group. |
| Noindex directive present | Warning | Google stopped honouring Noindex in robots.txt in 2019. The line does nothing and the pages stay indexable. | Use a meta robots noindex tag or an X-Robots-Tag header, and delete the line. |
| Unknown directive | Warning | The checker recognises User-agent, Allow, Disallow, Crawl-delay, Sitemap, Host, Clean-param, Request-rate, Noindex and Content-Signal. Anything else is flagged because most crawlers ignore what they do not implement. | Check the spelling. If you added it deliberately for a crawler that documents it, ignore the warning. |
| Per-agent access status | Allowed / Partial / Blocked | For each of 14 AI agents the checker picks the one group that agent obeys. Disallow: / without Allow: / is BLOCKED. Any narrower Disallow is PARTIAL. No Disallow, or no matching group, is ALLOWED. | Edit the group the note names. A named group replaces the wildcard for its agent rather than adding to it. |
| One or more AI crawlers blocked | Note | A summary line whenever any agent is BLOCKED. It does not affect validity; many sites block training crawlers on purpose. | Decide agent by agent with the reference table below. |
The named-group bug: how one Allow: / line exposes your admin pages
This is the most common serious mistake in files that name AI crawlers, and it comes from one rule almost every tutorial skips: a crawler obeys exactly one group.
When Googlebot, GPTBot or ClaudeBot reads your file, it finds the group whose User-agent matches it most specifically, reads that group’s rules, and ignores every other group in the file, including "User-agent: *". Groups do not stack and nothing is inherited from the wildcard. The named group is the whole policy for that agent.
Now consider the file most people write when they decide to welcome AI crawlers. The wildcard group carries the real work: Disallow lines for /login, /admin/, /dashboard/ and /api/. Below it sits a tidy list of named groups, one per AI agent, each containing a single "Allow: /". The intent is "you are welcome too". The effect is "ignore every Disallow above and crawl the admin panel".
That was the shape of this site’s own robots.txt until it was rebuilt. The fix is not clever: every group that names an agent must repeat the full Disallow block, because that group is the only thing the agent reads. The corrected file is long and repetitive, and that is the correct shape; the header comment in the live file at asvaai.com/robots.txt records the bug for the next person tempted to tidy it.
Broken: named group discards every Disallow
User-agent: * Allow: / Disallow: /login Disallow: /admin/ Disallow: /dashboard/ Disallow: /api/ User-agent: GPTBot Allow: / User-agent: ClaudeBot Allow: / Sitemap: https://example.com/sitemap.xml
Fixed: every group carries the full policy
User-agent: * Allow: / Disallow: /login Disallow: /admin/ Disallow: /dashboard/ Disallow: /api/ User-agent: GPTBot User-agent: ClaudeBot Allow: / Disallow: /login Disallow: /admin/ Disallow: /dashboard/ Disallow: /api/ Sitemap: https://example.com/sitemap.xml
In the broken file GPTBot and ClaudeBot match their named groups, read one Allow: /, and crawl /admin/ and /dashboard/. In the fixed file the two agents share a group through consecutive User-agent lines, and that group repeats the Disallow block. Different policies per agent need separate groups with full rule sets.
AI crawler reference: who each agent is and what blocking it costs
The checker resolves 14 agents in two classes that deserve different decisions. Retrieval agents fetch pages at query time so an answer engine can read and cite them; block one and you vanish from that engine’s answers. Training crawlers collect pages for future models; block one and you are excluded from training with no effect on live answers. Two entries are not crawlers at all.
| User-agent | Vendor | Class | What blocking it costs you |
|---|---|---|---|
| OAI-SearchBot | OpenAI | Retrieval (search index) | You leave ChatGPT search results and stop being cited with links. Documented as separate from training. |
| ChatGPT-User | OpenAI | Retrieval (on demand) | ChatGPT cannot open your page when a user or a GPT action points at it. Not used for training. |
| GPTBot | OpenAI | Training | Excluded from OpenAI’s future model training. Does not remove you from ChatGPT search; that is OAI-SearchBot. |
| PerplexityBot | Perplexity | Retrieval (search index) | You drop out of Perplexity’s index and cited sources. Documented as not used for training. |
| Perplexity-User | Perplexity | Retrieval (on demand) | Perplexity documents this user-initiated fetcher as generally not following robots.txt, so a Disallow here is advisory. |
| ClaudeBot | Anthropic | Training and general crawl | Excluded from Anthropic’s crawl for model development. Claude’s live fetches use Claude-User. |
| Claude-User | Anthropic | Retrieval (on demand) | Claude cannot open your page for a user, and you stop appearing as a cited source in its web answers. |
| Google-Extended | Opt-out token, not a crawler | Controls whether Googlebot’s crawl is used for Gemini training and grounding. No effect on Google Search or AI Overviews, which are part of Search. Never appears as a User-agent header. | |
| Applebot-Extended | Apple | Opt-out token, not a crawler | Controls whether Applebot’s crawl trains Apple’s foundation models. Applebot still crawls for Siri and Spotlight. Never fetches anything. |
| Amazonbot | Amazon | Crawl and index | Documented as improving services such as Alexa answers. Blocking removes you as a source for Alexa and Amazon’s shopping assistants. |
| Meta-ExternalAgent | Meta | Training and indexing | Documented as crawling to train models and index content for Meta products. Blocking excludes you from both. |
| CCBot | Common Crawl | Training (indirect) | Common Crawl’s open archive feeds many model trainers. Blocking removes you from the archive; no effect on any live engine. |
| Bingbot | Microsoft | Search index and Copilot | Copilot’s web answers draw on Bing’s index. Blocking removes you from Bing and Copilot citations together. |
| DuckAssistBot | DuckDuckGo | Retrieval (on demand) | DuckAssist answers cannot cite you. Documented as not used for training. |
| Bytespider | ByteDance | Training. Not in the checker | Feeds ByteDance’s models. Costs nothing in ChatGPT, Claude, Perplexity or Google. Falls under whatever your wildcard group says. |
The practical split for a business that wants citations: allow every retrieval agent, then decide the training crawlers on their own merits. A site that blocks GPTBot but allows OAI-SearchBot keeps its ChatGPT visibility; a site that blocks OAI-SearchBot by accident loses it while its owner believes the site is open. This site allows all of them and says so in the file; that is a policy choice for a company that sells AI visibility, not the right answer for every publisher.
CDN-managed robots.txt: when Cloudflare writes rules above yours
Cloudflare offers a managed robots.txt that blocks AI crawlers on your behalf. When it is on, the response at /robots.txt is your file with Cloudflare-written directives prepended above it. The file in your repository never changes, which is why fetching the live URL matters.
The managed block opens with a comment containing "BEGIN Cloudflare Managed content" and closes with a matching END comment. Between them sit User-agent groups for AI crawlers with Disallow rules and, depending on settings, Content-Signal lines.
The problem arrives when your own file also names those agents. You now have two GPTBot groups: Cloudflare’s at the top saying Disallow: /, yours below saying Allow: /. What a crawler does with duplicate groups is parser-dependent. RFC 9309 says matching groups should be combined, and Google’s parser does that, after which Allow: / and Disallow: / tie on length and the Allow wins. A first-match parser, which is what this checker and several crawlers use, stops at Cloudflare’s block and blocks the agent. Two crawlers, two answers, one file.
Ambiguous: two groups for the same agent
# BEGIN Cloudflare Managed content User-agent: GPTBot Disallow: / User-agent: ClaudeBot Disallow: / # END Cloudflare Managed Content User-agent: * Disallow: /admin/ User-agent: GPTBot Allow: / Disallow: /admin/
Unambiguous: one source of truth
User-agent: * Disallow: /admin/ User-agent: GPTBot Allow: / Disallow: /admin/ User-agent: ClaudeBot Allow: / Disallow: /admin/ Sitemap: https://example.com/sitemap.xml
Google combines the two GPTBot groups and allows the crawl; a first-match parser blocks it, and the checker reports BLOCKED because Cloudflare’s group comes first. Remove the duplication: turn the managed block off and own the policy in your file, or keep it and delete your own groups for the agents it covers.
✓Fetch, do not paste
Run Fetch by URL on the live domain. A Cloudflare Managed comment you did not write means the CDN is editing your file.
✓Pick one owner
Either the CDN decides AI crawler policy or your file does. Two owners means the answer depends on which crawler is asking.
✓Re-check after CDN changes
Toggling a bot-management option rewrites the served file without a deploy. Make the fetch part of the checklist.
Common robots.txt failures and how to fix them
Five patterns account for most files that pass a casual read and fail a crawler. Failing shape on the left, fixed shape on the right.
Fails: relative Sitemap, placed inside a group
User-agent: Googlebot Sitemap: /sitemap.xml User-agent: Bingbot Disallow: /admin/
Passes
User-agent: Googlebot User-agent: Bingbot Disallow: /admin/ Sitemap: https://example.com/sitemap.xml
Two faults in four lines. The relative URL is an error. And the Sitemap line between the two User-agent lines ends the first run, so Googlebot gets a group with no rules, meaning allow everything, and only Bingbot gets the Disallow. Sitemap is group-independent; put it at the end.
Silent: Host does nothing for most crawlers
User-agent: * Disallow: /admin/ Host: www.example.com
Fixed
User-agent: * Disallow: /admin/ Sitemap: https://www.example.com/sitemap.xml
Host was a Yandex directive for choosing a preferred domain; Google, Bing and every AI crawler ignore it. The checker accepts it without a warning because it is a known directive, but it sets a canonical for nobody. Use a 301 to the preferred host and canonical tags. This site dropped its own Host line for that reason.
Wrong: regex syntax in a path
User-agent: * Disallow: /products/.* Disallow: /*.pdf
Fixed
User-agent: * Disallow: /products/ Disallow: /*.pdf$
Paths support two special characters: * matches any run of characters and $ anchors the end. There is no regex. "/products/.*" literally matches paths beginning "/products/." and nothing you meant. "/*.pdf" without the anchor also matches "/report.pdf?download=1"; add $ to stop at the extension.
Over-blocks: /api/ serves an embed
User-agent: * Disallow: /api/
Fixed
User-agent: * Disallow: /api/ Allow: /api/embed/ Allow: /api/og-image/
Blocking /api/ is usually right, until a page depends on a script, image or JSON payload served under it. A crawler that cannot fetch the embed renders the page without it, and a preview image behind /api/ disappears from social cards and AI answers. Allow the specific paths; under the longest-match rule an Allow for /api/embed/ beats a Disallow for /api/.
Misplaced: trailing Allow joins the last group
User-agent: * Disallow: /admin/ User-agent: GPTBot Disallow: / Allow: /blog/
Fixed
User-agent: * Disallow: /admin/ Allow: /blog/ User-agent: GPTBot Disallow: /
A blank line does not end a group; only the next User-agent line does. The trailing "Allow: /blog/" therefore belongs to GPTBot, opening your blog to the one crawler you meant to block and doing nothing for the wildcard group it was written for.
Checker vs validator vs Google’s robots.txt report
A validator asks whether the file is well-formed: directives spelled right, every line parseable, Sitemap absolute. A checker asks what the file does: which crawlers it lets in, which it keeps out, and by which rule. Google’s robots.txt report in Search Console asks a narrower question from one crawler’s seat: what Googlebot fetched at /robots.txt on each host, with what status, and which lines it could not parse. This page is validator and checker together, pointed at AI agents because that is where the expensive mistakes now sit. It does not replace Google’s report, which is the place to confirm a deploy landed; Google does not show whether OAI-SearchBot or PerplexityBot can reach you, because Google does not run those crawlers.
| Question | This checker | Search Console robots.txt report |
|---|---|---|
| Is the syntax valid? | Yes, with line numbers for parse failures, BOM, unknown directives and Noindex. | Yes, for lines Googlebot could not parse. |
| Which AI crawlers can reach the site? | Yes, 14 agents resolved to the group each obeys. | No. Googlebot only. |
| Is a specific URL allowed? | No. The checker judges the site root. | Not in the report; Google’s open-source parser answers this offline. |
| What is the CDN actually serving? | Yes, via Fetch by URL. | Yes, from Googlebot’s own fetch, with version history. |
| Works on a draft or staging file? | Yes, in paste mode, without uploading. | No. Live hosts only. |
Who uses a robots.txt checker for AI crawlers
Three roles, three reasons.
✓SEO and AEO leads
Add the fetch to the launch checklist beside the sitemap check. The question is whether the retrieval agents behind the engines you report on can reach the pages you want cited, and whether a well-meant Allow: / opened the admin area to them.
✓Developers and platform engineers
Paste the file before it merges. The named-group bug, the trailing rule on the wrong group and the regex-style path all look fine in a diff and fail at runtime. Then fetch the live URL after deploy to see what the CDN did.
✓Agencies and publishers
Run every domain through Fetch by URL and compare the panels. Keep the training decision (GPTBot, ClaudeBot, CCBot) separate from the visibility decision (OAI-SearchBot, Claude-User, PerplexityBot). Files that conflate the two lose citations they never meant to give up.
Frequently asked questions
What does this robots.txt checker check?+
Syntax and access. On syntax it flags an empty file, a UTF-8 BOM, lines without a colon, a missing User-agent group, a Sitemap that is not an absolute URL, a missing Sitemap directive, a Noindex line and any unrecognised directive. On access it resolves 14 AI agents, including GPTBot, OAI-SearchBot, ClaudeBot, Claude-User, PerplexityBot and Bingbot, to the one group each obeys and reports ALLOWED, PARTIAL or BLOCKED with the deciding rule.
Is a robots.txt checker the same as a robots.txt validator?+
A validator tests whether the file is well-formed; a checker tests what the file does to specific crawlers. This tool is both: the syntax checks are the validator and the AI crawler panel is the checker. Google’s robots.txt report in Search Console is a third thing, a log of what Googlebot fetched, and it says nothing about AI agents.
Why does the checker say an AI crawler is allowed when my wildcard group blocks it?+
Because that crawler has its own named group, and a crawler obeys only the most specific group that matches it. If "User-agent: GPTBot" is followed by "Allow: /" and nothing else, GPTBot reads that one line and ignores every Disallow under "User-agent: *". Repeat the Disallow block inside the named group, or stack the User-agent lines so the agents share one group.
What is the difference between PARTIAL and BLOCKED?+
BLOCKED means the group that agent obeys contains Disallow: / and no Allow: /, so the whole site is closed. PARTIAL means the group disallows a narrower path such as /admin/ while the root stays open, the healthy state for most sites. ALLOWED means no Disallow at all: fine for a wildcard group with nothing to hide, suspicious for a named AI group that was supposed to inherit your private paths.
Does blocking GPTBot remove my site from ChatGPT?+
No. GPTBot is OpenAI’s training crawler. ChatGPT search results come from OAI-SearchBot and on-demand page reads from ChatGPT-User, both documented as separate from training. Block GPTBot and you are excluded from future training; block OAI-SearchBot and you disappear from ChatGPT search citations.
Does Google-Extended affect my Google rankings or AI Overviews?+
No. Google-Extended is an opt-out token read from Googlebot’s crawl, not a separate crawler, and it governs whether your pages are used for Gemini training and grounding. Google documents that it has no effect on inclusion or ranking in Search, and AI Overviews are part of Search. Applebot-Extended works the same way for Apple’s models.
Why does the live file contain rules I never wrote?+
Almost always a CDN or a plugin. Cloudflare’s managed robots.txt prepends a block marked with a "BEGIN Cloudflare Managed content" comment that disallows AI crawlers, and some SEO plugins inject rules of their own. Your repository file is unchanged; the served file is not. Use Fetch by URL to see the served version, then decide whether the CDN or your file owns the policy.
Does the checker test whether a specific URL is allowed?+
No. It judges each agent against the site root: Disallow: / is blocked, a narrower Disallow is partial, no Disallow is allowed. It does not evaluate a path against wildcard patterns, and it reads Disallow: / literally, so Disallow: /* reports as partial. For per-URL answers, Google’s open-source robots.txt parser runs the exact matching logic offline.
Is my file uploaded when I use this tool?+
In paste mode, no; the parser runs in your browser and the text stays on your machine, so an unreleased file is safe to test. In Fetch by URL mode our server retrieves the public robots.txt at the domain you enter, because a browser cannot read another site’s file directly, and returns it to your browser for checking.
Does the checker understand Content-Signal?+
Yes. Content-Signal is a proposal for stating search, AI-input and AI-training preferences inside robots.txt, and this site’s own file carries one. The checker parses it as a known directive and does not warn on it. It does not enforce it, because the signal is a request to crawlers rather than an access rule, and support varies by vendor. Treat it as a statement of policy, and keep using Allow and Disallow for actual access control.