// open source

answer-crawler-check

Install
npx answer-crawler-check
Runs on
Node.js 20 or later
Licence
MIT, free to use

answer-crawler-check fails your build when robots.txt blocks an AI answer crawler from any page in your sitemap.

npx answer-crawler-check https://example.com/sitemap-index.xml

Why it exists

Each AI engine reads robots.txt through its own crawler names. OpenAI documents OAI-SearchBot for search and GPTBot for training, each with its own robots.txt entry. Block the first by mistake and ChatGPT search can’t fetch the page it would have cited.

We saw the gap on this site. In a review of the build, someone added User-agent: Claude-SearchBot and Disallow: / to the built file, and every other check passed. Since then the build tests every page for every answer crawler we allow. This tool is that test, for any site.

A check of your home page, or of one URL you paste in, misses a rule like Disallow: /docs/ in one crawler’s group, or /*.html in another. Those block some pages and leave the home page open. This tool reads every URL in the sitemap.

What you get back

https://example.com: robots.txt from dist/robots.txt, local file, 3 page(s)
ok      OAI-SearchBot  (no group of its own, uses *)
BLOCKED Claude-SearchBot  1 URL(s)
          https://example.com/docs/setup/  Disallow: /docs/ (line 9)
...
answer-crawler-check: FAILED, 2 problem(s)

Each blocked page comes with the rule and line that blocked it.

Which crawlers it checks

By default it checks 11 crawlers. They fetch pages while an answer is written, or build the index the answers cite. They are OAI-SearchBot, ChatGPT-User, Claude-SearchBot, Claude-User, PerplexityBot, Perplexity-User, DuckAssistBot, MistralAI-User, Applebot, Googlebot and Bingbot.

Training crawlers such as GPTBot and CCBot are left out. Blocking them is a fair choice, and it doesn’t take you out of search answers. Name them with --agents if you want them checked too.

The crawler names come from ai.robots.txt, the community list of AI crawlers, used under its MIT licence. A copy ships with the tool, and --list latest fetches the current one.

Check before you deploy

npm run build
npx answer-crawler-check dist/sitemap-index.xml --site-dir dist

With --site-dir, the sitemaps and robots.txt come from your build folder, so a bad rule fails the build before it reaches the live site.

Running it in GitHub Actions

- uses: synapsereality/answer-crawler-check@v0.1.0
  with:
    sitemap: dist/sitemap-index.xml
    site-dir: dist

What it can’t see

robots.txt is one of several ways to turn a crawler away. A firewall rule, a bot challenge or a noindex header can block a crawler this tool calls allowed. Our free instant audit reads the HTML a crawler receives for one page.

< All open source