Your access log already knows which browsers and devices visit you. It has every request, including the ones an analytics script never sees: people with ad blockers, API clients, bots. This guide turns that log into a browser report with the User Agent Inspector API, using a couple of dozen API calls for a whole day of traffic.
Count first, look up second
The obvious version is a middleware that parses the User-Agent of every request. Don't build that one. It puts a network call in front of every page, and the free tier is 75 requests a day.
You don't need it, because traffic repeats itself. I made a sample log of 1,400 requests to try this out. It has 22 different User-Agent strings in it. So the plan is: count the distinct strings, look each one up once, and multiply.
Counting is one line. The User-Agent is the sixth quoted field in nginx's and Apache's combined log format:
awk -F'"' '{print $6}' /var/log/nginx/access.log | sort | uniq -c | sort -rn > ua_counts.txt
262 Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/140.0.0.0 Safari/537.36
228 Mozilla/5.0 (iPhone; CPU iPhone OS 18_5 like Mac OS X) AppleWebKit/605.1.15 (KHTML, like Gecko) Version/18.5 Mobile/15E148 Safari/604.1
182 Mozilla/5.0 (Linux; Android 10; K) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/140.0.0.0 Mobile Safari/537.36
123 Mozilla/5.0 (compatible; AhrefsBot/7.0; +http://ahrefs.com/robot/)
...
A real site has more than 22, with a long tail of old versions and odd bots at a hit or two each. That's why the list is sorted by count. Whatever fits in today's quota is the part that matters most, and the tail can wait for tomorrow.
Look up and report
This script reads ua_counts.txt, asks the API about every string it hasn't seen before, and keeps the answers in ua_cache.json. Then it adds things up, weighted by hits.
import json
import os
import re
from collections import Counter
import requests
CACHE_FILE = "ua_cache.json"
cache = json.load(open(CACHE_FILE)) if os.path.exists(CACHE_FILE) else {}
def lookup(user_agent):
if user_agent not in cache:
res = requests.get(
"https://apixies.io/api/v1/inspect-user-agent",
params={"user_agent": user_agent},
headers={"X-API-Key": os.environ["APIXIES_API_KEY"]},
timeout=10,
)
body = res.json()
if body["status"] != "success":
# Out of quota or a bad string. Leave it for the next run.
print(f"skipped ({body['code']}): {user_agent[:60]}")
return None
cache[user_agent] = body["data"]
return cache[user_agent]
def device_class(data):
if data["device"]["family"] == "iPad":
return "tablet"
if data["os"]["family"] in ("iOS", "Android") or "Mobile" in data["browser"]["family"]:
return "mobile"
return "desktop"
browsers, systems, devices, bots = Counter(), Counter(), Counter(), Counter()
unknown = 0
for line in open("ua_counts.txt"):
hits, user_agent = re.match(r"\s*(\d+) (.*)", line).groups()
hits = int(hits)
data = lookup(user_agent) if len(user_agent) > 1 else None
if data is None:
unknown += hits
elif data["is_bot"]:
bots[data["browser"]["family"]] += hits
else:
browsers[f"{data['browser']['family']} {data['browser']['major']}"] += hits
systems[data["os"]["family"]] += hits
devices[device_class(data)] += hits
json.dump(cache, open(CACHE_FILE, "w"))
people = sum(browsers.values())
total = people + sum(bots.values()) + unknown
print(f"{total} requests: {people} from people, {sum(bots.values())} from bots, {unknown} unknown\n")
for title, counter in (("Browsers", browsers), ("Systems", systems), ("Devices", devices)):
print(title)
for name, hits in counter.most_common(8):
print(f" {name:<24}{hits:>6} {hits / people:6.1%}")
print()
print("Bots")
for name, hits in bots.most_common(8):
print(f" {name:<24}{hits:>6}")
On my sample log it made 21 API calls and printed this:
1400 requests: 1113 from people, 265 from bots, 22 unknown
Browsers
Chrome 140 358 32.2%
Mobile Safari 18 228 20.5%
Chrome Mobile 140 182 16.4%
Firefox 143 63 5.7%
Safari 18 60 5.4%
Edge 140 49 4.4%
Instagram 340 39 3.5%
Mobile Safari 17 38 3.4%
Systems
Windows 382 34.3%
iOS 305 27.4%
Android 256 23.0%
Mac OS X 156 14.0%
Linux 14 1.3%
Devices
desktop 552 49.6%
mobile 523 47.0%
tablet 38 3.4%
Bots
AhrefsBot 123
Googlebot 86
bingbot 33
Other 12
okhttp 7
curl 4
The second run made no calls at all and took under a tenth of a second, because everything was in the cache. That's the point of the cache file. Run the report from cron every night and you only pay for strings that are new.
The 22 unknown requests had - as their User-Agent. That's how nginx logs a missing header, and there's nothing to look up.
If the rest of your tooling is Node, here's the lookup step on its own. It fills the same cache file, so you can build the report in whatever you like:
import { existsSync, readFileSync, writeFileSync } from "node:fs";
const CACHE_FILE = "ua_cache.json";
const cache = existsSync(CACHE_FILE) ? JSON.parse(readFileSync(CACHE_FILE, "utf8")) : {};
const lines = readFileSync("ua_counts.txt", "utf8").split("\n").filter(Boolean);
let looked = 0;
for (const line of lines) {
const userAgent = line.trim().replace(/^\d+ /, "");
if (userAgent.length < 2 || cache[userAgent]) continue;
const params = new URLSearchParams({ user_agent: userAgent });
const res = await fetch(`https://apixies.io/api/v1/inspect-user-agent?${params}`, {
headers: { "X-API-Key": process.env.APIXIES_API_KEY },
});
const body = await res.json();
if (body.status !== "success") {
console.error(`stopped at ${body.code}, run it again tomorrow`);
break;
}
cache[userAgent] = body.data;
looked++;
}
writeFileSync(CACHE_FILE, JSON.stringify(cache));
console.log(`${looked} new lookups, ${Object.keys(cache).length} strings known`);
Reading the report honestly
A few things in that output aren't what they look like.
The API has no desktop or mobile field. device.family is a device name: iPhone, Samsung SM-S921B, Mac, or Other for a Windows or Linux PC. The device_class function in the script is my own rule on top of it. Android tablets land under mobile, because their strings don't say tablet.
Don't report OS versions. The script counts OS families only, and that's on purpose. Chrome froze the OS part of its User-Agent, so every desktop Chrome says Windows 10 or macOS 10.15.7, and Chrome on Android says Android 10; K for every phone. Safari and Firefox pin the macOS version too. A chart of "Windows 10 vs Windows 11" built from User-Agents would be fiction. Browser family and major version are still real, and those are what you need for a support decision.
In-app browsers are their own row. Instagram shows up as Instagram 340, not as Safari. Keep it that way. If 3.5% of your visitors are inside Instagram's webview, that's a browser you've probably never tested in.
Bots are almost a fifth of this log, and a header is just a header. is_bot catches crawlers and HTTP libraries that say who they are. A scraper that sends a Chrome string is counted with the people. The bot detection guide goes into that.
Other under bots is a client the crawler list knows and the browser parser doesn't. In my log it was Postman.
What it's good for
The questions this answers are the dull, useful ones. Can you drop support for a browser version? Look for it in the list. Is mobile half your traffic or a tenth? Which in-app browsers should be on the test plan? How much of yesterday's spike was AhrefsBot?
It won't replace a real analytics tool. There are no sessions, no unique visitors, no referrers here, only requests. What you get is a browser and device breakdown that counts everybody, from data you already have, with nothing added to your pages.
Next steps
- User Agent Inspector API reference: the parameter and every response field
- User Agent Parser API tutorial: what the API returns for real strings, and where it stops
- Bot detection with the User Agent API: sort the welcome bots from the rest
- All guides