Skip to content

Someone pastes a link into your app and you want a picture next to it. The usual source is the page's og:image, and plenty of pages don't have one. A screenshot always exists. This guide builds a small preview generator on the Screenshot API: one capture per link, stored on your side, with a check for the captures that come out wrong.

The capture

curl -G "https://apixies.io/api/v1/screenshot" \
     -H "X-API-Key: YOUR_API_KEY" \
     --data-urlencode "url=https://news.ycombinator.com" \
     -d width=1200 -d height=630 -d format=jpeg \
     -o hn-preview.jpg

1200 x 630 is the size Facebook and LinkedIn recommend for og:image, and it crops well into most card layouts. Hacker News has no og:image at all, so this is the only preview you'll get for it.

Use JPEG. I captured three homepages at that size in both formats:

Page PNG JPEG (quality 80)
github.com 115 KB 60 KB
news.ycombinator.com 146 KB 91 KB
laravel.com 200 KB 92 KB

Roughly half, and at card size nobody can tell. If that's still too heavy, quality=40 took a 1280 x 800 GitHub capture from 57 KB to 36 KB.

Don't shrink the viewport to get a thumbnail

It's tempting to ask for width=640&height=400 when the card is small. Don't. width is the browser window, so the site answers with its tablet or phone layout. GitHub at 640 wide gives you a hamburger menu and a headline that fills the frame. It doesn't look like the site people know.

Capture at 1200 x 630 and resize the file yourself (sharp, Pillow, GD, or just CSS). You get the desktop page, and one stored image serves every card size.

Capture once, keep the file

A capture takes a few seconds (2.3 for Hacker News from my machine, over 4 for Laravel), and a free account has 75 requests a day. Both facts say the same thing: never call the API while someone is waiting for a page, and never call it twice for the same link.

The pattern:

  1. A link gets saved. Queue a job.
  2. The job captures the page and stores the JPEG under a name made from the URL.
  3. Your pages only ever serve the stored file.
  4. A nightly job refreshes the oldest previews with whatever quota is left.

75 a day goes further than it sounds. It covers 75 new links every day, or a weekly refresh of a library of about 500.

JavaScript

import { createHash } from "node:crypto";
import { access, mkdir, writeFile } from "node:fs/promises";

const DIR = "previews";

async function preview(url) {
  const file = `${DIR}/${createHash("sha256").update(url).digest("hex").slice(0, 16)}.jpg`;

  // Already have it? Done. No request spent.
  try { await access(file); return file; } catch {}

  const params = new URLSearchParams({ url, width: "1200", height: "630", format: "jpeg" });
  const res = await fetch(`https://apixies.io/api/v1/screenshot?${params}`, {
    headers: { "X-API-Key": process.env.APIXIES_API_KEY },
  });

  if (!res.headers.get("content-type")?.startsWith("image/")) {
    const body = await res.json();
    throw new Error(`${body.code}: ${body.message}`);
  }

  await mkdir(DIR, { recursive: true });
  await writeFile(file, Buffer.from(await res.arrayBuffer()));
  return file;
}

Python

import hashlib
import os
import requests

DIR = "previews"

def preview(url):
    file = os.path.join(DIR, hashlib.sha256(url.encode()).hexdigest()[:16] + ".jpg")
    if os.path.exists(file):
        return file

    res = requests.get(
        "https://apixies.io/api/v1/screenshot",
        params={"url": url, "width": 1200, "height": 630, "format": "jpeg"},
        headers={"X-API-Key": os.environ["APIXIES_API_KEY"]},
        timeout=90,
    )
    if not res.headers.get("Content-Type", "").startswith("image/"):
        body = res.json()
        raise RuntimeError(f"{body['code']}: {body['message']}")

    os.makedirs(DIR, exist_ok=True)
    with open(file, "wb") as f:
        f.write(res.content)
    return file

PHP

function preview(string $url, string $dir = 'previews'): string
{
    $file = $dir . '/' . substr(hash('sha256', $url), 0, 16) . '.jpg';
    if (is_file($file)) {
        return $file;
    }

    $query = http_build_query(['url' => $url, 'width' => 1200, 'height' => 630, 'format' => 'jpeg']);
    $ch = curl_init("https://apixies.io/api/v1/screenshot?$query");
    curl_setopt_array($ch, [
        CURLOPT_HTTPHEADER => ['X-API-Key: ' . getenv('APIXIES_API_KEY')],
        CURLOPT_RETURNTRANSFER => true,
        CURLOPT_TIMEOUT => 90,
    ]);
    $body = curl_exec($ch);
    $type = (string) curl_getinfo($ch, CURLINFO_CONTENT_TYPE);
    curl_close($ch);

    if (! str_starts_with($type, 'image/')) {
        $error = json_decode((string) $body, true);
        throw new RuntimeException(($error['code'] ?? 'ERROR') . ': ' . ($error['message'] ?? 'no response'));
    }

    if (! is_dir($dir)) {
        mkdir($dir, 0775, true);
    }
    file_put_contents($file, $body);

    return $file;
}

When the call throws, show a placeholder and try again tomorrow. A missing thumbnail is fine. A broken page isn't.

The previews you don't want

This is the part nobody mentions. The API returns 200 and a valid JPEG whenever the browser managed to load something. What it loaded is another matter. Four real ones from my test run:

  • theguardian.com: a "Personalised advertising, it's your choice" dialog sits in the middle of the page.
  • spiegel.de: a consent dialog covers nearly everything.
  • stackoverflow.com: not the site at all. A Cloudflare page saying "Performing security verification".
  • g2.com: "Access is temporarily restricted".

There's no parameter to click a banner away or hide an element, so you can't fix these from your side. You can catch some of them, though.

The bot walls are the easy half. The Meta Tag Extractor fetches the same page without a browser, and for both Stack Overflow and G2 it came back with FETCH_FAILED. When that happens the screenshot is almost certainly a block page, so skip it:

curl -G "https://apixies.io/api/v1/extract-meta" \
     -H "X-API-Key: YOUR_API_KEY" \
     --data-urlencode "url=https://stackoverflow.com/questions"

That call pays for itself in a second way. Its response has data.open_graph.image. GitHub has a designed one, and using it costs you no screenshot. So the full order is: read the meta tags, take og:image if there is one, take a screenshot if there isn't, use a placeholder if the fetch failed.

Be a little careful with og:image too. Der Spiegel's is just its logo, and the Guardian's international front page doesn't set one.

Consent dialogs have no such signal. They're mostly a news-site problem, and mostly European. If your links lean that way, keep a short list of domains where you prefer og:image or a placeholder over a screenshot.

Using the image as your own og:image

If the preview is for a page of yours (a bookmark page, a directory entry), point the tags at the stored file:

<meta property="og:image" content="https://yoursite.example/previews/3f9a1c2b7d4e5f60.jpg" />
<meta property="og:image:width" content="1200" />
<meta property="og:image:height" content="630" />
<meta name="twitter:card" content="summary_large_image" />

Next steps

Try the URL Screenshot Generator API

Free tier is for development & small projects. 75 requests/day with a registered account.

cookies

We use analytics cookies to see how the site gets used. Nothing loads until you accept. Privacy policy