You want to know which of your visitors are bots, which bots they are, and whether the one calling itself Googlebot really is. The User Agent Inspector API answers the first two from the User-Agent header. The third needs a DNS check, and that's further down.
What the API tells you
Send it a User-Agent string:
curl -G "https://apixies.io/api/v1/inspect-user-agent" \
-H "X-API-Key: YOUR_API_KEY" \
--data-urlencode "user_agent=Mozilla/5.0 (compatible; AhrefsBot/7.0; +http://ahrefs.com/robot/)"
{
"status": "success",
"data": {
"is_bot": true,
"device": { "family": "Spider", "model": "Desktop", "brand": "Spider" },
"os": { "family": "Other" },
"browser": { "family": "AhrefsBot", "major": "7", "minor": "0" }
}
}
Two fields do the work. is_bot says bot or not, and browser.family says which one. I ran a handful of real strings through it:
| User-Agent contains | is_bot |
browser.family |
|---|---|---|
Googlebot/2.1 (desktop and smartphone) |
true | Googlebot |
bingbot/2.0 |
true | bingbot |
AhrefsBot/7.0 |
true | AhrefsBot |
GPTBot/1.2 |
true | GPTBot |
ClaudeBot/1.0 |
true | ClaudeBot |
curl/8.5.0 |
true | curl |
python-requests/2.32.3 |
true | Python Requests |
| Chrome 128 on Windows | false | Chrome |
Note the lowercase bingbot. The family is whatever the bot calls itself, so compare names without caring about case.
Scripts count as bots too. curl and python-requests come back with is_bot: true, which is what you want for analytics and maybe not what you want for an API that people are supposed to call from scripts.
Don't look up every request
It's tempting to put this in a middleware and call it on each request. Don't. It adds a network round trip to every page, and the free tier is 75 requests a day.
You don't need to. A site sees the same few hundred User-Agent strings over and over. Look each one up once and keep the answer:
const seen = new Map();
async function classify(userAgent) {
if (seen.has(userAgent)) return seen.get(userAgent);
const params = new URLSearchParams({ user_agent: userAgent });
const res = await fetch(`https://apixies.io/api/v1/inspect-user-agent?${params}`, {
headers: { "X-API-Key": process.env.APIXIES_API_KEY },
});
const body = await res.json();
// If the lookup fails, treat the visitor as human. Never block on an outage.
const verdict = body.status === "success"
? { isBot: body.data.is_bot, name: body.data.browser.family }
: { isBot: false, name: null };
if (body.status === "success") seen.set(userAgent, verdict);
return verdict;
}
In Express that becomes:
app.use(async (req, res, next) => {
const verdict = await classify(req.headers["user-agent"] || "");
req.isBot = verdict.isBot;
req.botName = verdict.name;
next();
});
With more than one server, put the map in Redis. Even better for big sites: run the lookups offline over yesterday's access log and skip the middleware completely. Analytics doesn't need the answer in real time.
Welcome some, turn others away
Once you have the name, the policy is yours. Search engines in, SEO crawlers out, AI crawlers wherever you stand on that:
import os
import requests
ALLOWED = {"googlebot", "bingbot", "duckduckbot", "applebot"}
BLOCKED = {"ahrefsbot", "semrushbot", "mj12bot", "dotbot"}
def policy(user_agent):
res = requests.get(
"https://apixies.io/api/v1/inspect-user-agent",
params={"user_agent": user_agent},
headers={"X-API-Key": os.environ["APIXIES_API_KEY"]},
timeout=10,
)
data = res.json()["data"]
if not data["is_bot"]:
return "human"
name = data["browser"]["family"].lower()
if name in BLOCKED:
return "blocked"
if name in ALLOWED:
return "allowed"
return "unknown_bot"
For the well-behaved ones, robots.txt is the better tool. Ahrefs and Semrush both obey it. A block list in code is for the ones that don't.
Is it really Googlebot?
Here's the hole in everything above. The User-Agent is a header, and the client writes it. This request claims to be Googlebot:
curl -A "Mozilla/5.0 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)" https://your-site.example/
So does every scraper that read a tutorial. If your allow list waves Googlebot past the rate limiter, you've just told them how to get past it.
Google's own answer is a DNS check in two steps. Look up the name for the IP address, then look up the address for that name, and see if you land where you started. Only Google can make both directions agree.
The DNS Lookup API takes an IP address directly and does the reverse lookup. Here's a real Googlebot address:
curl -G "https://apixies.io/api/v1/dns-lookup" \
-H "X-API-Key: YOUR_API_KEY" \
--data-urlencode "domain=66.249.66.1"
{
"status": "success",
"data": {
"ip": "66.249.66.1",
"domain": "1.66.249.66.in-addr.arpa",
"records": [
{ "type": "PTR", "ttl": 7102, "target": "crawl-66-249-66-1.googlebot.com" }
]
}
}
The name ends in googlebot.com. Good, but not proof yet, because whoever owns an IP range can put any name they like in its reverse DNS. So go back the other way:
curl -G "https://apixies.io/api/v1/dns-lookup" \
-H "X-API-Key: YOUR_API_KEY" \
--data-urlencode "domain=crawl-66-249-66-1.googlebot.com" \
--data-urlencode "type=A"
{
"records": [
{ "type": "A", "host": "crawl-66-249-66-1.googlebot.com", "ttl": 7102, "ip": "66.249.66.1" }
]
}
Same address. That one is Google.
Now an address that isn't. I took a Tor exit node and imagined it sending the Googlebot header:
{
"ip": "185.220.101.34",
"records": [
{ "type": "PTR", "target": "tor-exit-34.for-privacy.net" }
]
}
Wrong name, so you can stop after one call. An address with no PTR record at all fails the same way.
The rules, all of them:
- The PTR name has to end in
.googlebot.com,.google.comor.googleusercontent.com. Check the ending with the dot.evilgooglebot.comends ingooglebot.comtoo. - The A record (AAAA for IPv6) of that name has to contain the address you started with.
- Both or nothing.
It works for IPv6 as well. 2001:4860:4801:10::1 comes back as crawl-2001-4860-4801-0010-0000-0000-0000-0001.googlebot.com.
In code
JavaScript
const GOOGLE = [".googlebot.com", ".google.com", ".googleusercontent.com"];
async function dns(params) {
const res = await fetch(`https://apixies.io/api/v1/dns-lookup?${new URLSearchParams(params)}`, {
headers: { "X-API-Key": process.env.APIXIES_API_KEY },
});
const body = await res.json();
return body.status === "success" ? body.data.records : [];
}
async function isRealGooglebot(ip) {
const ptr = (await dns({ domain: ip })).find((r) => r.type === "PTR");
if (!ptr || !GOOGLE.some((suffix) => ptr.target.endsWith(suffix))) return false;
const type = ip.includes(":") ? "AAAA" : "A";
const forward = await dns({ domain: ptr.target, type });
return forward.some((r) => r.ip === ip || r.ipv6 === ip);
}
Python
import os
import requests
GOOGLE = (".googlebot.com", ".google.com", ".googleusercontent.com")
def dns(**params):
res = requests.get(
"https://apixies.io/api/v1/dns-lookup",
params=params,
headers={"X-API-Key": os.environ["APIXIES_API_KEY"]},
timeout=15,
)
body = res.json()
return body["data"]["records"] if body["status"] == "success" else []
def is_real_googlebot(ip):
ptr = next((r for r in dns(domain=ip) if r["type"] == "PTR"), None)
if not ptr or not ptr["target"].endswith(GOOGLE):
return False
record_type = "AAAA" if ":" in ip else "A"
forward = dns(domain=ptr["target"], type=record_type)
return any(ip in (r.get("ip"), r.get("ipv6")) for r in forward)
PHP
function dns(array $params): array
{
$context = stream_context_create(['http' => [
'header' => 'X-API-Key: ' . getenv('APIXIES_API_KEY'),
'ignore_errors' => true,
]]);
$body = json_decode(file_get_contents(
'https://apixies.io/api/v1/dns-lookup?' . http_build_query($params), false, $context
), true);
return ($body['status'] ?? '') === 'success' ? $body['data']['records'] : [];
}
function isRealGooglebot(string $ip): bool
{
$ptr = current(array_filter(dns(['domain' => $ip]), fn ($r) => $r['type'] === 'PTR'));
if (! $ptr || ! preg_match('/\.(googlebot|google|googleusercontent)\.com$/', $ptr['target'])) {
return false;
}
$type = str_contains($ip, ':') ? 'AAAA' : 'A';
foreach (dns(['domain' => $ptr['target'], 'type' => $type]) as $record) {
if (in_array($ip, [$record['ip'] ?? null, $record['ipv6'] ?? null], true)) {
return true;
}
}
return false;
}
That's two API calls per address, so cache the verdict per IP for a day. Google's crawlers come from a small pool of addresses, and you'll see the same ones all week.
Only run it when it matters. A visitor who says Chrome doesn't need checking. One who says Googlebot and is about to skip your rate limit does.
The same check for other crawlers
Bing works the same way, and its names end in .search.msn.com. Apple's end in .applebot.apple.com. Most AI crawlers don't offer a DNS check. They publish IP lists instead.
Google publishes lists too, as JSON, at developers.google.com/static/crawling/ipranges/common-crawlers.json. Matching an address against those ranges needs no DNS at all. It's the better choice if you check a lot of traffic, as long as you refresh the file every day.
Keep bots out of your numbers
The quiet win is analytics. Tag the request, and leave bots out when you count visitors:
app.use((req, res, next) => {
if (req.isBot) {
botLog.write({ bot: req.botName, path: req.path, ip: req.ip, at: new Date() });
}
next();
});
After a week that log answers things you couldn't see before. How often Google really crawls you. Which pages the scrapers go for. Whether the traffic spike on Tuesday was people.
Next steps
- User Agent Inspector API reference: every field in the response
- DNS Lookup API reference: record types and reverse lookups
- User Agent Parser API tutorial: the basics, in four languages
- Browser analytics with the User Agent Parser API: a traffic breakdown from your logs
- All guides