Skip to content

You want to know which of your visitors are bots, which bots they are, and whether the one calling itself Googlebot really is. The User Agent Inspector API answers the first two from the User-Agent header. The third needs a DNS check, and that's further down.

What the API tells you

Send it a User-Agent string:

curl -G "https://apixies.io/api/v1/inspect-user-agent" \
     -H "X-API-Key: YOUR_API_KEY" \
     --data-urlencode "user_agent=Mozilla/5.0 (compatible; AhrefsBot/7.0; +http://ahrefs.com/robot/)"
{
  "status": "success",
  "data": {
    "is_bot": true,
    "device": { "family": "Spider", "model": "Desktop", "brand": "Spider" },
    "os": { "family": "Other" },
    "browser": { "family": "AhrefsBot", "major": "7", "minor": "0" }
  }
}

Two fields do the work. is_bot says bot or not, and browser.family says which one. I ran a handful of real strings through it:

User-Agent contains is_bot browser.family
Googlebot/2.1 (desktop and smartphone) true Googlebot
bingbot/2.0 true bingbot
AhrefsBot/7.0 true AhrefsBot
GPTBot/1.2 true GPTBot
ClaudeBot/1.0 true ClaudeBot
curl/8.5.0 true curl
python-requests/2.32.3 true Python Requests
Chrome 128 on Windows false Chrome

Note the lowercase bingbot. The family is whatever the bot calls itself, so compare names without caring about case.

Scripts count as bots too. curl and python-requests come back with is_bot: true, which is what you want for analytics and maybe not what you want for an API that people are supposed to call from scripts.

Don't look up every request

It's tempting to put this in a middleware and call it on each request. Don't. It adds a network round trip to every page, and the free tier is 75 requests a day.

You don't need to. A site sees the same few hundred User-Agent strings over and over. Look each one up once and keep the answer:

const seen = new Map();

async function classify(userAgent) {
  if (seen.has(userAgent)) return seen.get(userAgent);

  const params = new URLSearchParams({ user_agent: userAgent });
  const res = await fetch(`https://apixies.io/api/v1/inspect-user-agent?${params}`, {
    headers: { "X-API-Key": process.env.APIXIES_API_KEY },
  });
  const body = await res.json();

  // If the lookup fails, treat the visitor as human. Never block on an outage.
  const verdict = body.status === "success"
    ? { isBot: body.data.is_bot, name: body.data.browser.family }
    : { isBot: false, name: null };

  if (body.status === "success") seen.set(userAgent, verdict);
  return verdict;
}

In Express that becomes:

app.use(async (req, res, next) => {
  const verdict = await classify(req.headers["user-agent"] || "");
  req.isBot = verdict.isBot;
  req.botName = verdict.name;
  next();
});

With more than one server, put the map in Redis. Even better for big sites: run the lookups offline over yesterday's access log and skip the middleware completely. Analytics doesn't need the answer in real time.

Welcome some, turn others away

Once you have the name, the policy is yours. Search engines in, SEO crawlers out, AI crawlers wherever you stand on that:

import os
import requests

ALLOWED = {"googlebot", "bingbot", "duckduckbot", "applebot"}
BLOCKED = {"ahrefsbot", "semrushbot", "mj12bot", "dotbot"}

def policy(user_agent):
    res = requests.get(
        "https://apixies.io/api/v1/inspect-user-agent",
        params={"user_agent": user_agent},
        headers={"X-API-Key": os.environ["APIXIES_API_KEY"]},
        timeout=10,
    )
    data = res.json()["data"]

    if not data["is_bot"]:
        return "human"

    name = data["browser"]["family"].lower()
    if name in BLOCKED:
        return "blocked"
    if name in ALLOWED:
        return "allowed"
    return "unknown_bot"

For the well-behaved ones, robots.txt is the better tool. Ahrefs and Semrush both obey it. A block list in code is for the ones that don't.

Is it really Googlebot?

Here's the hole in everything above. The User-Agent is a header, and the client writes it. This request claims to be Googlebot:

curl -A "Mozilla/5.0 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)" https://your-site.example/

So does every scraper that read a tutorial. If your allow list waves Googlebot past the rate limiter, you've just told them how to get past it.

Google's own answer is a DNS check in two steps. Look up the name for the IP address, then look up the address for that name, and see if you land where you started. Only Google can make both directions agree.

The DNS Lookup API takes an IP address directly and does the reverse lookup. Here's a real Googlebot address:

curl -G "https://apixies.io/api/v1/dns-lookup" \
     -H "X-API-Key: YOUR_API_KEY" \
     --data-urlencode "domain=66.249.66.1"
{
  "status": "success",
  "data": {
    "ip": "66.249.66.1",
    "domain": "1.66.249.66.in-addr.arpa",
    "records": [
      { "type": "PTR", "ttl": 7102, "target": "crawl-66-249-66-1.googlebot.com" }
    ]
  }
}

The name ends in googlebot.com. Good, but not proof yet, because whoever owns an IP range can put any name they like in its reverse DNS. So go back the other way:

curl -G "https://apixies.io/api/v1/dns-lookup" \
     -H "X-API-Key: YOUR_API_KEY" \
     --data-urlencode "domain=crawl-66-249-66-1.googlebot.com" \
     --data-urlencode "type=A"
{
  "records": [
    { "type": "A", "host": "crawl-66-249-66-1.googlebot.com", "ttl": 7102, "ip": "66.249.66.1" }
  ]
}

Same address. That one is Google.

Now an address that isn't. I took a Tor exit node and imagined it sending the Googlebot header:

{
  "ip": "185.220.101.34",
  "records": [
    { "type": "PTR", "target": "tor-exit-34.for-privacy.net" }
  ]
}

Wrong name, so you can stop after one call. An address with no PTR record at all fails the same way.

The rules, all of them:

  1. The PTR name has to end in .googlebot.com, .google.com or .googleusercontent.com. Check the ending with the dot. evilgooglebot.com ends in googlebot.com too.
  2. The A record (AAAA for IPv6) of that name has to contain the address you started with.
  3. Both or nothing.

It works for IPv6 as well. 2001:4860:4801:10::1 comes back as crawl-2001-4860-4801-0010-0000-0000-0000-0001.googlebot.com.

In code

JavaScript

const GOOGLE = [".googlebot.com", ".google.com", ".googleusercontent.com"];

async function dns(params) {
  const res = await fetch(`https://apixies.io/api/v1/dns-lookup?${new URLSearchParams(params)}`, {
    headers: { "X-API-Key": process.env.APIXIES_API_KEY },
  });
  const body = await res.json();
  return body.status === "success" ? body.data.records : [];
}

async function isRealGooglebot(ip) {
  const ptr = (await dns({ domain: ip })).find((r) => r.type === "PTR");
  if (!ptr || !GOOGLE.some((suffix) => ptr.target.endsWith(suffix))) return false;

  const type = ip.includes(":") ? "AAAA" : "A";
  const forward = await dns({ domain: ptr.target, type });
  return forward.some((r) => r.ip === ip || r.ipv6 === ip);
}

Python

import os
import requests

GOOGLE = (".googlebot.com", ".google.com", ".googleusercontent.com")

def dns(**params):
    res = requests.get(
        "https://apixies.io/api/v1/dns-lookup",
        params=params,
        headers={"X-API-Key": os.environ["APIXIES_API_KEY"]},
        timeout=15,
    )
    body = res.json()
    return body["data"]["records"] if body["status"] == "success" else []

def is_real_googlebot(ip):
    ptr = next((r for r in dns(domain=ip) if r["type"] == "PTR"), None)
    if not ptr or not ptr["target"].endswith(GOOGLE):
        return False

    record_type = "AAAA" if ":" in ip else "A"
    forward = dns(domain=ptr["target"], type=record_type)
    return any(ip in (r.get("ip"), r.get("ipv6")) for r in forward)

PHP

function dns(array $params): array
{
    $context = stream_context_create(['http' => [
        'header' => 'X-API-Key: ' . getenv('APIXIES_API_KEY'),
        'ignore_errors' => true,
    ]]);
    $body = json_decode(file_get_contents(
        'https://apixies.io/api/v1/dns-lookup?' . http_build_query($params), false, $context
    ), true);

    return ($body['status'] ?? '') === 'success' ? $body['data']['records'] : [];
}

function isRealGooglebot(string $ip): bool
{
    $ptr = current(array_filter(dns(['domain' => $ip]), fn ($r) => $r['type'] === 'PTR'));
    if (! $ptr || ! preg_match('/\.(googlebot|google|googleusercontent)\.com$/', $ptr['target'])) {
        return false;
    }

    $type = str_contains($ip, ':') ? 'AAAA' : 'A';
    foreach (dns(['domain' => $ptr['target'], 'type' => $type]) as $record) {
        if (in_array($ip, [$record['ip'] ?? null, $record['ipv6'] ?? null], true)) {
            return true;
        }
    }

    return false;
}

That's two API calls per address, so cache the verdict per IP for a day. Google's crawlers come from a small pool of addresses, and you'll see the same ones all week.

Only run it when it matters. A visitor who says Chrome doesn't need checking. One who says Googlebot and is about to skip your rate limit does.

The same check for other crawlers

Bing works the same way, and its names end in .search.msn.com. Apple's end in .applebot.apple.com. Most AI crawlers don't offer a DNS check. They publish IP lists instead.

Google publishes lists too, as JSON, at developers.google.com/static/crawling/ipranges/common-crawlers.json. Matching an address against those ranges needs no DNS at all. It's the better choice if you check a lot of traffic, as long as you refresh the file every day.

Keep bots out of your numbers

The quiet win is analytics. Tag the request, and leave bots out when you count visitors:

app.use((req, res, next) => {
  if (req.isBot) {
    botLog.write({ bot: req.botName, path: req.path, ip: req.ip, at: new Date() });
  }
  next();
});

After a week that log answers things you couldn't see before. How often Google really crawls you. Which pages the scrapers go for. Whether the traffic spike on Tuesday was people.

Next steps

Try the User Agent Inspector API

Free tier is for development & small projects. 75 requests/day with a registered account.

cookies

We use analytics cookies to see how the site gets used. Nothing loads until you accept. Privacy policy