Puppeteer on Serverless Keeps Breaking: When to Use a Screenshot API

It's 2am, PagerDuty fires, and the invoice PDF job on Lambda has died again (TimeoutError: Waiting for navigation ... 30000ms exceeded), on every third…

It's 2am, PagerDuty fires, and the invoice PDF job on Lambda has died again (TimeoutError: Waiting for navigation ... 30000ms exceeded), on every third request, for no reason your local repro can show. If you've shipped Puppeteer to serverless, you know this night. It's not your bug. Headless Chrome on serverless fails in predictable, compounding ways, and at some point the fix is deleting the browser from your infrastructure entirely.

Failure mode #1: You can't even fit Chrome in the box

As the accreditly team lays out, AWS Lambda has a 250MB unzipped deployment limit, while a full Chromium binary runs around 170MB compressed. The DocuPotion team notes the standard puppeteer package bundles Chrome for Testing at roughly 170–280MB, blowing past both Lambda's 50MB zipped direct-upload cap and the unzipped S3 limit. Vercel fails the same way, capping compressed function bundles at about 50MB [1].

So you reach for puppeteer-core plus @sparticuz/chromium, the recommended workaround for Node 18+ [1]. Except even the latest @sparticuz/chromium (v143.0.4) zipped still exceeds Lambda's 50MB upload limit, as the DocuPotion team found when they tried to ship it. And even if it fits, DocuPotion found Lambda's minimal Linux lacks the shared libraries Chrome needs, like libnss3.so and fonts, so launches fail before you've rendered a single pixel.

Failure mode #2: Cold starts tax every render

The accreditly team measured that booting headless Chrome inside a Lambda container takes two to four seconds even on warm infrastructure, longer when the binary has to be extracted from a layer first. On Lambda with Chromium, cold starts run 3–5 seconds minimum, which is why toolkitonline concluded a screenshot service simply can't run on serverless and needs persistent containers.

Vercel is no kinder. The first request can be seconds slower than warm ones because the container, Node, and Chromium all have to boot [1]. Teams fight back with warming pings every ~10 minutes and Vercel's Fluid compute, just to keep it tolerable [1]. With scale-to-zero, every cold start lands on a real user waiting for their screenshot.

Failure mode #3: Memory math that doesn't survive traffic

The Custodia team found each Puppeteer browser process holds roughly 150MB base plus page overhead. Ten concurrent screenshot requests exhaust a small instance, and the tenth gets killed. Worse, accreditly hit a page with a complex JS-rendered chart that briefly spiked Chrome to 700MB, OOM-killing a 1024MB Lambda on specific URLs only. Intermittent, URL-dependent, and invisible locally: the worst kind of bug.

Even in the clean case, Puppeteer leaks. The devforth team measured a single browser instance with pages created and closed in a loop: dramatic memory jumps plus a constant slow climb, and their belief is the leak may never receive a true fix [2]. Production then needs a browser pool on top of everything else (infrastructure to manage your infrastructure).

Failure mode #4: Zombies and timeouts that only happen in prod

One developer who spent six months running Puppeteer screenshot infrastructure in production found that if your Node process crashes mid-render, an unhandled rejection or an OOM kill, the Chromium it launched keeps running as a zombie, eating memory until the container dies. Serverless makes this nastier because containers recycle unpredictably, so the zombie's lifetime is out of your hands.

Then there's the classic Custodia describes: a screenshot that takes 2 seconds locally fails in production with TimeoutError: Waiting for navigation ... Timeout 30000ms exceeded on every third request. Same code, same URL. The difference is constrained CPU, shared memory, and cold containers, things you can't reproduce with node script.js on your laptop.

Failure mode #5: Chrome version drift: the slow rot

The accreditly team points out Chrome ships a new major version roughly every four weeks, chrome-aws-lambda famously stopped receiving updates, and its successor @sparticuz/chromium requires a matching puppeteer-core version. Someone has to babysit that pairing forever.

The accreditly team lived it: after eight months of stable operation, a Chrome major-version update left their pinned Chromium layer stale, causing missing fonts and cold starts creeping from two seconds to nearly nine. The same author later hit version drift where the same URL rendered differently locally (Chrome 120) versus on Lambda (pinned to 114), half a day lost to a collapsing flexbox. Their verdict after rebuilding the layer three times in two years: it works, but it doesn't stay working without active maintenance.

The bill nobody puts in the estimate

The compute itself is almost embarrassingly cheap. DocuPotion's worked example: 100,000 PDFs per month at five seconds each with 2048MB memory costs about $10/month on Lambda after the free tier. That's the number that tricks you into shipping it.

The real cost is engineering. The Custodia team describes one: three months after adding Puppeteer for screenshots and PDFs, a team's Docker image was 800MB heavier, CI took twice as long, and they had debugged three separate production memory-leak incidents. Add the operational gotchas DocuPotion lists: 2048–3072MB is the optimal memory range for most Puppeteer tasks, and Lambda's 6MB synchronous response limit forces you into S3 plus presigned URLs for anything larger. Your "one function" now has a bucket, a signing step, and a memory-tuning spreadsheet.

The decision framework: keep Puppeteer or switch?

Keep Puppeteer if you genuinely need browser automation, clicking, form filling, multi-step authenticated flows. Rendering a string of HTML or a URL to an image or PDF is just a function call, and a function call doesn't need a browser you own.

Even the serverless optimists draw this line: andreas_a advises skipping Vercel Functions for screenshots if you need consistently sub-second latency, high concurrency, or long-running sessions, and suggests a browser pool on Fly.io, Railway, or EC2, or a managed screenshot service [1]. If you stay self-hosted, persistent containers with a pool beat scale-to-zero serverless every time. If your job is launch β†’ goto β†’ screenshot/pdf, hosted APIs win on cold starts because they run pre-warmed Chrome instances. The cold start is someone else's problem, permanently.

The migration checklist: from Puppeteer block to one API call

When accreditly finally swapped their Puppeteer block for a fetch to a rendering API, their Lambda package dropped about 80%, down to roughly 40KB zipped, and the cold start vanished with it. Their Puppeteer version had been 180MB of dependencies with a three-second cold start; the fetch-based version had no browser to maintain at all. Do the same in six steps, ending on MarkupGo.

1. Delete the Chromium layer and dependency. Uninstall puppeteer-core, @sparticuz/chromium, the Lambda layer, the --no-sandbox, --disable-gpu, and --disable-dev-shm-usage flags. None of it is your problem anymore.

2. Map your options. Puppeteer's viewport maps to MarkupGo's width/height (defaults 800Γ—600, formats png/jpeg/webp, default quality 80) [3]. waitForTimeout becomes waitDelay, page.waitForSelector becomes waitForExpression, and the image API docs cover the full options table, including emulatedMediaType, extraHttpHeaders, and failOnHttpStatusCodes.

3. Fix the silent defaults Puppeteer never told you about. Puppeteer defaults to an 800Γ—600 viewport, which makes screenshots look like they were taken on a monitor from 2001. Set your real dimensions explicitly.

4. Swap the code. browser.launch() + page.goto() + page.screenshot() becomes:

const markupgo = new MarkupGo(API_KEY);

// Screenshot a URL
markupgo.image.fromUrl("https://example.com").json();

// Invoice PDF
markupgo.pdf.fromUrl(url).buffer();

5. Verify output handling and expiry. Check whether your flow wants a JSON task object or a buffer, and set expiration (expiresInSeconds from 60 seconds up to 90 days) so anything sensitive auto-deletes. MarkupGo processes document inputs in memory and never stores them, which makes the security conversation with your customers a short one.

6. Test the URLs that broke you. Run the chart-heavy page that OOM'd your Lambda against the API and watch it just work.

Frequently asked questions

How much memory does Puppeteer need on AWS Lambda?

Plan for 2048–3072MB as the optimal range for most Puppeteer tasks. Below that, complex pages can spike Chrome to 700MB and get OOM-killed on a 1024MB allocation. And remember the deployment side: a full Chromium binary is around 170MB compressed against Lambda's 250MB unzipped limit, so Docker images (up to 10GB uncompressed) are often the only path.

What's the best way to run Chromium on serverless in 2026?

The accepted setup is puppeteer-core plus @sparticuz/chromium for Node 18+, with chrome-aws-lambda as the Node 16 fallback [1]. But even the latest @sparticuz/chromium zipped exceeds Lambda's 50MB upload limit, so you'll be on Docker images. The honest answer from people who've done it: persistent containers with a browser pool, or a managed rendering API [1]. Accreditly rebuilt their Lambda layer three times in two years before concluding it doesn't stay working without active maintenance.

How much does a screenshot API cost compared to self-hosting?

The Lambda compute itself is cheap, about $10/month for 100,000 PDFs at five seconds each with 2048MB. A rendering API is comparable on volume: MarkupGo's Lite plan is $29/month for 2,000 credits [4], Plus is $49/month for 10,000 [4], and Pro is $99/month for 20,000 [4]. What you're actually buying is the engineering hours the self-hosted route silently eats: the 800MB image, the doubled CI, the three memory-leak incidents in three months.

Why does Puppeteer leak memory in production?

Two known issues compound. First, if your Node process crashes mid-render, the Chromium it launched keeps running as a zombie process eating memory until the container dies. Second, even with a single browser and pages created and closed in a loop, Puppeteer shows both dramatic memory jumps and a constant slow climb that may never receive a true fix [2]. Long-running processes need browser pool management to survive production.

Is it safe to send my HTML to a screenshot API?

With MarkupGo, document inputs are processed in memory only and never persisted to storage [5]. Generated outputs can be set to auto-delete anywhere between 60 seconds and 90 days [5], and the full GDPR/DPA terms are published so legal doesn't need a discovery call. It's worth reading the DPA before sending customer invoices through any provider. That's exactly what it's for.


The decision rule: if your Puppeteer code is just launch β†’ goto β†’ screenshot/pdf, you're paying the Chromium tax to run a rendering function. Delete the Chromium layer this week, swap the block for fromUrl()/fromHtml() (Node client, 100 free credits every month, no credit card), and keep Puppeteer only for the flows that genuinely need to click and scroll. Your 2am timeout tickets go with it.

Sources

  1. dev.to
  2. devforth.io
  3. markupgo.com
  4. markupgo.com
  5. markupgo.com