Robots.txt Analyzer
Enter a domain to fetch and parse its robots.txt file. The tool breaks down the rules by user agent, showing allowed and disallowed paths, crawl delay settings, and sitemap references. Useful for SEO auditing and verifying that search engines can access the pages you want indexed.
Without https://
returns
- data.user_agents
- Rules per user agent with allow, disallow and crawl_delay.
- data.sitemaps
- Sitemap URLs listed in the file.
- data.found
- False when the site doesn't have a robots.txt.
- data.raw
- The file exactly as served.
- data.status_code
- HTTP status of the robots.txt request.
Fill in the fields and send a request. The response lands here.
Run it from your code
A temporary key takes one request and lasts seven days at 20 calls a day. Register and it becomes 75 a day, still free.
$ curl -s "https://apixies.io/api/v1/robots-txt?domain=github.com" \ -H "X-API-Key: $APIXIES_KEY" | jq '.data.user_agents' { "*": { "allow": [], "disallow": ["/cdn-cgi/"], "crawl_delay": null } }
questions
Does robots.txt actually block crawlers?
It's an advisory standard, not an enforcement mechanism. Well-behaved bots (Googlebot, Bingbot) respect it. Malicious bots ignore it. Don't use robots.txt to hide sensitive content — use authentication instead.
Should I block /admin in robots.txt?
Ironically, listing /admin in robots.txt tells everyone where your admin panel is. If your admin area requires authentication (it should), there's no need to list it. Robots.txt is publicly readable.