Skip to content

Your access log already knows which browsers and devices visit you. It has every request, including the ones an analytics script never sees: people with ad blockers, API clients, bots. This guide turns that log into a browser report with the User Agent Inspector API, using a couple of dozen API calls for a whole day of traffic.

Count first, look up second

The obvious version is a middleware that parses the User-Agent of every request. Don't build that one. It puts a network call in front of every page, and the free tier is 75 requests a day.

You don't need it, because traffic repeats itself. I made a sample log of 1,400 requests to try this out. It has 22 different User-Agent strings in it. So the plan is: count the distinct strings, look each one up once, and multiply.

Counting is one line. The User-Agent is the sixth quoted field in nginx's and Apache's combined log format:

awk -F'"' '{print $6}' /var/log/nginx/access.log | sort | uniq -c | sort -rn > ua_counts.txt
    262 Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/140.0.0.0 Safari/537.36
    228 Mozilla/5.0 (iPhone; CPU iPhone OS 18_5 like Mac OS X) AppleWebKit/605.1.15 (KHTML, like Gecko) Version/18.5 Mobile/15E148 Safari/604.1
    182 Mozilla/5.0 (Linux; Android 10; K) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/140.0.0.0 Mobile Safari/537.36
    123 Mozilla/5.0 (compatible; AhrefsBot/7.0; +http://ahrefs.com/robot/)
    ...

A real site has more than 22, with a long tail of old versions and odd bots at a hit or two each. That's why the list is sorted by count. Whatever fits in today's quota is the part that matters most, and the tail can wait for tomorrow.

Look up and report

This script reads ua_counts.txt, asks the API about every string it hasn't seen before, and keeps the answers in ua_cache.json. Then it adds things up, weighted by hits.

import json
import os
import re
from collections import Counter

import requests

CACHE_FILE = "ua_cache.json"
cache = json.load(open(CACHE_FILE)) if os.path.exists(CACHE_FILE) else {}

def lookup(user_agent):
    if user_agent not in cache:
        res = requests.get(
            "https://apixies.io/api/v1/inspect-user-agent",
            params={"user_agent": user_agent},
            headers={"X-API-Key": os.environ["APIXIES_API_KEY"]},
            timeout=10,
        )
        body = res.json()
        if body["status"] != "success":
            # Out of quota or a bad string. Leave it for the next run.
            print(f"skipped ({body['code']}): {user_agent[:60]}")
            return None
        cache[user_agent] = body["data"]
    return cache[user_agent]

def device_class(data):
    if data["device"]["family"] == "iPad":
        return "tablet"
    if data["os"]["family"] in ("iOS", "Android") or "Mobile" in data["browser"]["family"]:
        return "mobile"
    return "desktop"

browsers, systems, devices, bots = Counter(), Counter(), Counter(), Counter()
unknown = 0

for line in open("ua_counts.txt"):
    hits, user_agent = re.match(r"\s*(\d+) (.*)", line).groups()
    hits = int(hits)

    data = lookup(user_agent) if len(user_agent) > 1 else None
    if data is None:
        unknown += hits
    elif data["is_bot"]:
        bots[data["browser"]["family"]] += hits
    else:
        browsers[f"{data['browser']['family']} {data['browser']['major']}"] += hits
        systems[data["os"]["family"]] += hits
        devices[device_class(data)] += hits

json.dump(cache, open(CACHE_FILE, "w"))

people = sum(browsers.values())
total = people + sum(bots.values()) + unknown
print(f"{total} requests: {people} from people, {sum(bots.values())} from bots, {unknown} unknown\n")

for title, counter in (("Browsers", browsers), ("Systems", systems), ("Devices", devices)):
    print(title)
    for name, hits in counter.most_common(8):
        print(f"  {name:<24}{hits:>6}  {hits / people:6.1%}")
    print()

print("Bots")
for name, hits in bots.most_common(8):
    print(f"  {name:<24}{hits:>6}")

On my sample log it made 21 API calls and printed this:

1400 requests: 1113 from people, 265 from bots, 22 unknown

Browsers
  Chrome 140                 358   32.2%
  Mobile Safari 18           228   20.5%
  Chrome Mobile 140          182   16.4%
  Firefox 143                 63    5.7%
  Safari 18                   60    5.4%
  Edge 140                    49    4.4%
  Instagram 340               39    3.5%
  Mobile Safari 17            38    3.4%

Systems
  Windows                    382   34.3%
  iOS                        305   27.4%
  Android                    256   23.0%
  Mac OS X                   156   14.0%
  Linux                       14    1.3%

Devices
  desktop                    552   49.6%
  mobile                     523   47.0%
  tablet                      38    3.4%

Bots
  AhrefsBot                  123
  Googlebot                   86
  bingbot                     33
  Other                       12
  okhttp                       7
  curl                         4

The second run made no calls at all and took under a tenth of a second, because everything was in the cache. That's the point of the cache file. Run the report from cron every night and you only pay for strings that are new.

The 22 unknown requests had - as their User-Agent. That's how nginx logs a missing header, and there's nothing to look up.

If the rest of your tooling is Node, here's the lookup step on its own. It fills the same cache file, so you can build the report in whatever you like:

import { existsSync, readFileSync, writeFileSync } from "node:fs";

const CACHE_FILE = "ua_cache.json";
const cache = existsSync(CACHE_FILE) ? JSON.parse(readFileSync(CACHE_FILE, "utf8")) : {};

const lines = readFileSync("ua_counts.txt", "utf8").split("\n").filter(Boolean);
let looked = 0;

for (const line of lines) {
  const userAgent = line.trim().replace(/^\d+ /, "");
  if (userAgent.length < 2 || cache[userAgent]) continue;

  const params = new URLSearchParams({ user_agent: userAgent });
  const res = await fetch(`https://apixies.io/api/v1/inspect-user-agent?${params}`, {
    headers: { "X-API-Key": process.env.APIXIES_API_KEY },
  });
  const body = await res.json();

  if (body.status !== "success") {
    console.error(`stopped at ${body.code}, run it again tomorrow`);
    break;
  }
  cache[userAgent] = body.data;
  looked++;
}

writeFileSync(CACHE_FILE, JSON.stringify(cache));
console.log(`${looked} new lookups, ${Object.keys(cache).length} strings known`);

Reading the report honestly

A few things in that output aren't what they look like.

The API has no desktop or mobile field. device.family is a device name: iPhone, Samsung SM-S921B, Mac, or Other for a Windows or Linux PC. The device_class function in the script is my own rule on top of it. Android tablets land under mobile, because their strings don't say tablet.

Don't report OS versions. The script counts OS families only, and that's on purpose. Chrome froze the OS part of its User-Agent, so every desktop Chrome says Windows 10 or macOS 10.15.7, and Chrome on Android says Android 10; K for every phone. Safari and Firefox pin the macOS version too. A chart of "Windows 10 vs Windows 11" built from User-Agents would be fiction. Browser family and major version are still real, and those are what you need for a support decision.

In-app browsers are their own row. Instagram shows up as Instagram 340, not as Safari. Keep it that way. If 3.5% of your visitors are inside Instagram's webview, that's a browser you've probably never tested in.

Bots are almost a fifth of this log, and a header is just a header. is_bot catches crawlers and HTTP libraries that say who they are. A scraper that sends a Chrome string is counted with the people. The bot detection guide goes into that.

Other under bots is a client the crawler list knows and the browser parser doesn't. In my log it was Postman.

What it's good for

The questions this answers are the dull, useful ones. Can you drop support for a browser version? Look for it in the list. Is mobile half your traffic or a tenth? Which in-app browsers should be on the test plan? How much of yesterday's spike was AhrefsBot?

It won't replace a real analytics tool. There are no sessions, no unique visitors, no referrers here, only requests. What you get is a browser and device breakdown that counts everybody, from data you already have, with nothing added to your pages.

Next steps

Try the User Agent Inspector API

Free tier is for development & small projects. 75 requests/day with a registered account.

cookies

We use analytics cookies to see how the site gets used. Nothing loads until you accept. Privacy policy