Free Rainproxy tool

Robots.txt Tester

Fetch any site's robots.txt and check whether a path is crawlable by Googlebot or another user-agent.
GooglebotBingbotSitemaps
Live preview
robots.txtFetched
GooglebotAllowed
SitemapFound
READY
1. Fetch robots.txt

Crawling sites that fight back?

Even an allowed path can get rate-limited, fingerprinted or geo-blocked. Rainproxy routes your crawler through real residential and mobile IPs that pass Cloudflare, Akamai and DataDome cleanly.

See residential plans
How it works

Know what's crawlable before you crawl it

We fetch robots.txt, parse every group exactly the way Google does, and tell you whether your target path is fair game.

Step 1

Fetch robots.txt

Server-side request grabs the live robots file, no CORS, no caching surprises, no spoofed responses.

Step 2

Parse groups + rules

User-agent groups, allow/disallow precedence and sitemaps are read using Google's longest-match logic.

Step 3

Verdict on your path

Tells you Allow vs Disallow for Googlebot, Bingbot or any custom UA, with the exact matching rule.

Pre-crawl audits

Run before launching a scraper or SEO crawler to avoid wasted budget on disallowed paths.

Sitemap discovery

Surfaces every Sitemap: directive so you can point your indexer or scraper at the right URLs.

Ethical by default

We always show what the site asks crawlers to do. Respecting robots.txt is good citizenship and good SEO.

How to test robots.txt rules

This free Rainproxy tool fetches a site's robots.txt, parses the rules and tells you whether a given URL is allowed or blocked for the crawler you choose. It shows which line produced the verdict, so you can see exactly why a page is or is not crawlable.

How the rules are matched

Crawlers pick the most specific matching user-agent group, then apply the longest matching Allow or Disallow path within it. Longer wins, and if an Allow and a Disallow of equal length both match, the Allow takes precedence. Wildcards (*) and end-of-URL anchors ($) are supported. Anything not matched by a Disallow is allowed.

Blocked in robots.txt is not the same as noindex

Disallow stops a crawler fetching a page; it does not remove it from search results. A URL blocked in robots.txt can still appear in Google, described only by its anchor text, because Google can see links to it but cannot read the page. To keep a page out of the index, allow crawling and serve a noindex meta tag or header instead.

AI crawlers and scraping

robots.txt is also where you accept or refuse AI training and answer engines: GPTBot, ClaudeBot, PerplexityBot, Google-Extended and others read the same file. Test each of those user-agents separately, because a rule written for * may not be the one they follow. On the scraping side, remember robots.txt is a stated preference, not an access control, and respecting it is the baseline for compliant crawling.

Frequently asked questions