HTML to PDF in Node.js: Puppeteer vs a Hosted Conversion API

Your Puppeteer PDF code works perfectly on your laptop. Then you deploy it to Lambda: the third deploy breaks the build, every emoji turns into a hollow…

Your Puppeteer PDF code works perfectly on your laptop. Then you deploy it to Lambda: the third deploy breaks the build, every emoji turns into a hollow rectangle, and a 40-page report occasionally kills the function outright. This post walks through the page.pdf() code that does work, documents exactly why it falls apart on serverless, and shows the same conversion as a single hosted API call, so you can decide which trade-off your traffic can afford.

The Puppeteer baseline: page.pdf() in 15 lines

First, the code that works. On a long-running server you control, this is a good solution and you may not need anything else.

npm install puppeteer
import puppeteer from "puppeteer";

const browser = await puppeteer.launch();
const page = await browser.newPage();

await page.setContent(html, { waitUntil: "networkidle0" });

const pdf = await page.pdf({
  format: "A4",            // also: Letter, Legal, Tabloid
  landscape: false,
  scale: 1,                // 0.1 to 2
  printBackground: true,   // without this, your CSS backgrounds vanish
  margin: { top: "20mm", bottom: "20mm", left: "15mm", right: "15mm" },
  displayHeaderFooter: true,
  headerTemplate: `<div style="width:100%;text-align:center;">
    Invoice</div>`,
  footerTemplate: `<div style="width:100%;text-align:center;">
    <span class="pageNumber"></span> / <span class="totalPages"></span></div>`,
});

await browser.close();

Browserless's own Puppeteer PDF guide covers the option set well: format accepts A4, Legal, Letter, and Tabloid, and you can set landscape, custom width/height, and scale from 0.1 to 2. The options that make invoices actually look right are printBackground: true (otherwise every background color and image is stripped), the margins, and the displayHeaderFooter templates with pageNumber and totalPages classes.

One gotcha to know before you write any CSS: page.pdf() re-lays the page out against a physical page size in print media, so your screen styles don't carry over. page.pdf() switches the page to print media, re-lays out against the page dimensions, embeds fonts, and serialises content as vectors (dev.to). Use @media print rules, or you'll debug phantom layout differences between what you see in the browser and what lands in the PDF.

Why the same code breaks on serverless: the Chromium tax

The moment this code moves into a Lambda function, you inherit four problems that have nothing to do with PDFs.

Bundling. Lambda's unzipped deployment package limit is 250 MB including layers, and a Chromium build alone consumes most of that. The classic workaround, chrome-aws-lambda, stopped receiving updates; its successor is @sparticuz/chromium, which decompresses a couple hundred megabytes into /tmp on first launch. That decompression happens on the cold start your users wait through.

Version churn. Chrome ships a new major version roughly every four weeks, and puppeteer-core and @sparticuz/chromium have to be upgraded in lockstep or renders break. This is the "it worked last deploy" failure: nothing in your code changed, but the pinned Chromium binary and the protocol version drifted apart. If you don't own an upgrade cadence for this pair, you now do.

Fonts. The Amazon Linux image underlying Lambda ships with almost no fonts, so unresolvable glyphs render as tofu, hollow rectangles, with emoji and CJK text as the classic casualties. The fix is shipping font files in your deployment package, which eats more of the 250 MB you don't have.

Container images. Switching to a container image raises Lambda's deployment limit to 10 GB, which solves bundling, but worsens cold starts. You're trading one tax for another, not eliminating it.

Cold starts, timeouts, and the memory ceiling

Booting headless Chrome inside a Lambda function takes two to four seconds, plus extra time to decompress the binary on cold start. Then the render itself: a multi-page report with webfonts can take another five to ten seconds to paginate. Per APITemplate's serverless guide, cold starts on Chromium-based functions can add 5–15 seconds to total response time.

That number matters because of what sits in front of your function. API Gateway's default integration timeout is 29 seconds, which a cold start plus font fetch plus a long PDF document can exceed, producing 504s and retry amplification. The retry is the nasty part: your client retries, Lambda renders the same slow PDF twice, and your render load doubles at exactly the moment your infrastructure is struggling.

Why is PDF so much slower than the screenshots you may already be generating? page.pdf() is the heavy path: print media, re-layout, font embedding, vector serialisation. A long report that screenshots in two seconds can take eight or ten seconds to paginate into a PDF.

Provisioned concurrency mitigates the cold start, but it costs money continuously and undermines the cost benefits that got you into serverless in the first place.

The failure modes nobody puts in the tutorial

Memory kills. Chrome's print pipeline holds the whole paginated document in memory. Pages that fit comfortably in a 1024 MB function can spike past it on long tables or chart-heavy reports, and Lambda kills the function on roughly one render in fifty. One in fifty sounds tolerable until it's an invoice for your biggest customer.

Leaked browsers. If your code throws after launch() but before browser.close(), the Chrome process survives into the next invocation. Leaked browsers eventually poison the container with Protocol error: Target closed launch failures. The fix is try/finally around everything, plus a wrapper that force-kills strays, and it's the first thing every team forgets.

Crashed pages. Very long PDFs generated via page.pdf() can fail with a Page crashed! error. Browserless's guide notes that page.createPDFStream() avoids this by sending the PDF in chunks rather than building it entirely in memory. Worth switching to, but it's another non-obvious detail you discover in production.

The scaling math. Lambda allocates CPU linearly with memory: at 1,792 MB a function gets the equivalent of one full vCPU (What Does Lambda's Big Memory Increase Enable?). AWS raised the maximum memory from 3,008 MB to 10,240 MB with a linear CPU increase, at unchanged pricing of $0.0000166667 per GB-second. Translation: PDF rendering speed is something you literally pay per gigabyte. Slow render? Pay for more memory. That's a legitimate lever, but it means your unit cost is coupled to how badly Chrome behaves on your documents.

The same conversion as one API call

Now the same job with a hosted conversion API. MarkupGo provides a Node.js client, markupgo-node, that generates images and PDFs from templates, URLs, or HTML (markupgo.com).

npm install markupgo-node
import MarkupGo from "markupgo-node";
import fs from "fs";

const markupgo = new MarkupGo({
  API_KEY: process.env.MARKUPGO_API_KEY,
});

const html = `<html><body><h1>Invoice #1042</h1>...</body></html>`;

const pdfOptions = {
  properties: {
    printBackground: true,
    landscape: false,
    // see MarkupGo's API reference for the full option set:
    // headers, footers, margins, format
  },
};

markupgo.pdf.fromHtml(html, pdfOptions).buffer()
  .then((buffer) => {
    fs.writeFileSync("invoice-1042.pdf", Buffer.from(buffer));
  });

The full method list and option reference live in the Node client docs (https://markupgo.com/docs/node-client).

The options map one-to-one onto what you were passing to page.pdf(): the PDF conversion supports printBackground, landscape, and a singlePage option that renders all content on one page. Custom headers, footers, and margins are supported, along with SVG and emoji support and accessibility options (markupgo.com). That last part is the tofu fix from section 2, handled server-side: emoji render correctly without you shipping font packs in a deployment bundle.

What disappears from your codebase: no Chromium binary, no lockstep upgrades every four weeks, no /tmp decompression on cold start, no font bundles, no browser lifecycle management, no try/finally around browser.close().

The honest cost framing: you pay per conversion instead of per GB-second, and you take on an external dependency and network latency. If your PDFs contain data you can't send to a third party, that's a real constraint, not a footnote. And every conversion is a network round trip, which matters for latency-sensitive paths.

Puppeteer vs hosted API: the trade-off table

CriterionSelf-hosted PuppeteerHosted API (e.g. MarkupGo)
Setup & maintenanceYou own Chromium bundling, upgrades, and browser lifecycleInstall a client, manage an API key
Cold-start latency2–4 s Chrome boot plus decompression; 5–15 s total on LambdaNo cold start of your own; network round trip only
Memory headroomChrome holds the paginated doc in memory; spikes kill functionsNot your problem
Version churnLockstep upgrades of puppeteer-core + Chromium every ~4 weeksHandled by the vendor
Fonts, emoji, CJKShip font packs yourself or get tofuEmoji and SVG supported server-side
Per-render costPay per GB-second; engineering hours on topPer-credit pricing
Control & offlineFull control, works air-gapped, data stays in-houseExternal dependency, data leaves your network

On per-render cost: APITemplate's guide puts a 2048 MB Lambda running 10 seconds at roughly $0.00033 per invocation, varying by region. Cheap per call. The real cost is the engineering hours for upgrades and debugging that the table's first three rows describe.

On hosted pricing: MarkupGo plans start at $29/month for the Lite plan with 2,000 credits, Plus at $49/month for 10,000 credits, and Pro at $99/month for 20,000 credits with unlimited templates (markupgo.com). At Lite that's roughly 1.5¢ per credit, and the free trial grants 100 credits per month, enough to test with real documents before paying anything.

Where each wins, plainly:

  • Puppeteer wins on control, on cost at high steady volume on servers you already run, and on offline or air-gapped requirements. If your PDFs contain regulated data that can't leave your infrastructure, self-hosting isn't the expensive option, it's the only option.
  • A hosted API wins on serverless, on spiky traffic, on small teams without ops bandwidth, and on anything with emoji, CJK text, or long documents. The entire failure-mode list from sections 2 through 4 stops being your list.

Which should you choose? By situation

Long-running server with steady volume. You run an EC2 box, a container, or an existing Node app where the process stays warm. Stay with Puppeteer. The code in section 1 is fine, you avoid a new dependency, and there's no version churn problem because you control the Chrome version.

Lambda, serverless, or spiky traffic. A hosted API call removes the entire failure-mode list from sections 2–4: the 250 MB squeeze, the decompression, the lockstep upgrades, the leaked browsers, the memory kills. If your team has no ops bandwidth, this is the deciding criterion, not price.

Recurring invoice and report layouts. Re-sending full HTML on every conversion is wasteful when the layout never changes. Templates let you define the layout once and pass dynamic data per render. MarkupGo's Magic Template URL goes further: generate a URL for a template once, then produce images or PDFs from it with no API call at all, passing dynamic data as query parameters. Once a URL is generated it can be reused multiple times without additional credits, with one credit per conversion. Markdown-to-PDF is also supported in the Node client, and MarkupGo's Office conversion covers over 130 document formats including .docx, .xlsx, and .pptx, with merging into a single PDF, so a customer uploads a Word file and you hand back a PDF without touching Chromium at all.

The hybrid pattern worth naming. Prototype with Puppeteer locally, where it's fast and free, then switch the render call to a hosted endpoint in production behind a single renderPdf(html, options) interface. One function to swap, and you keep local iteration speed.

Frequently asked questions

Why does Puppeteer keep breaking on AWS Lambda?

Lambda's unzipped deployment limit is 250 MB including layers, and a Chromium build alone consumes most of it. On top of that, Chrome ships a major version roughly every four weeks and puppeteer-core and @sparticuz/chromium must be upgraded in lockstep or renders break. Per APITemplate's serverless guide, the old workaround package, chrome-aws-lambda, is deprecated in favor of @sparticuz/chromium, which shifts the problem rather than solving it.

How much memory does Puppeteer need to generate a PDF in Lambda?

Per APITemplate's serverless guide, cold starts on Chromium-based Lambda functions can add 5–15 seconds to response time, and provisioned concurrency mitigates this at a continuous cost. Memory-wise, Chrome's print pipeline holds the entire paginated document in memory, so a document that renders fine at 1024 MB can spike past it on long tables or chart-heavy reports. Since Lambda allocates CPU linearly with memory, bumping memory also buys rendering speed.

How do I fix the 'Page crashed!' error when generating long PDFs with Puppeteer?

Switch from page.pdf() to page.createPDFStream(). Very long PDFs built entirely in memory can fail with Page crashed!, and streaming sends the PDF in chunks instead, which avoids the crash. Also audit your error handling: any throw between launch() and browser.close() leaks a Chrome process that eventually poisons the container.

Why do emojis and Chinese characters show as boxes in my Lambda-generated PDFs?

The Amazon Linux image underlying Lambda ships with almost no fonts, so unresolvable glyphs render as tofu, hollow rectangles, with emoji and CJK text as the classic casualties. The fix is bundling font files into your deployment package, which consumes more of the 250 MB you're already fighting for. Hosted APIs handle this server-side; MarkupGo's PDF output supports emoji and SVG natively.

Is there a free HTML to PDF API I can test before paying?

MarkupGo's free trial is perpetual and grants 100 credits per month with no credit card required. That's enough to convert your actual invoice or report template, not a hello-world, and compare the output against your Puppeteer build side by side. Paid plans start at $29/month for 2,000 credits if the trial convinces you.

The decision rule

If your PDFs render on a long-running server you already operate, keep Puppeteer. It's the right tool, the code in section 1 is all you need, and adding a hosted dependency there buys you nothing.

If they render inside Lambda or any ephemeral environment, budget for the Chromium tax: lockstep upgrades, /tmp decompression, memory headroom, font packs. Or replace it with one API call and spend the reclaimed engineering hours on something else.

Either way, test with real documents. A 40-page report with webfonts, not a hello-world, because that's where both approaches actually differ. If you want to try the hosted path without committing, MarkupGo's free trial gives you 100 credits a month, enough to convert your real invoice template and compare output side by side with your Puppeteer build (https://markupgo.com/pricing).