Skip to content
GET /api/v1/robots-txt

Robots.txt Analyzer

Enter a domain to fetch and parse its robots.txt file. The tool breaks down the rules by user agent, showing allowed and disallowed paths, crawl delay settings, and sitemap references. Useful for SEO auditing and verifying that search engines can access the pages you want indexed.

string required

Without https://

X-Sandbox-Remaining: - get an api key

returns

data.user_agents
Rules per user agent with allow, disallow and crawl_delay.
data.sitemaps
Sitemap URLs listed in the file.
data.found
False when the site doesn't have a robots.txt.
data.raw
The file exactly as served.
data.status_code
HTTP status of the robots.txt request.
response awaiting request sending
Fill in the fields and send a request. The response lands here.

Run it from your code

A temporary key takes one request and lasts seven days at 20 calls a day. Register and it becomes 75 a day, still free.

$ curl -s "https://apixies.io/api/v1/robots-txt?domain=github.com" \
    -H "X-API-Key: $APIXIES_KEY" | jq '.data.user_agents'

{ "*": { "allow": [], "disallow": ["/cdn-cgi/"], "crawl_delay": null } }

questions

Does robots.txt actually block crawlers?

It's an advisory standard, not an enforcement mechanism. Well-behaved bots (Googlebot, Bingbot) respect it. Malicious bots ignore it. Don't use robots.txt to hide sensitive content — use authentication instead.

Should I block /admin in robots.txt?

Ironically, listing /admin in robots.txt tells everyone where your admin panel is. If your admin area requires authentication (it should), there's no need to list it. Robots.txt is publicly readable.

cookies

We use analytics cookies to see how the site gets used. Nothing loads until you accept. Privacy policy