DOCX to PDF in Production: Why Headless LibreOffice Hangs, and When to Use an API

You drop a DOCX into your upload handler, spawn libreoffice --headless --convert-to pdf, and wait. On your laptop it works. In production it hangs, exits with…

You drop a DOCX into your upload handler, spawn libreoffice --headless --convert-to pdf, and wait. On your laptop it works. In production it hangs, exits with no output, or silently skips half your concurrent requests. You Google the symptom, find a Stack Overflow thread from 2017, apply the accepted answer, and the problem moves somewhere else. This is the headless LibreOffice trap: the default answer to "convert Office files to PDF server-side" is not production-grade. Let's walk through the failure modes one by one, then look at what the same job looks like as a single API call.

The Hang That Never Returns

A corrupted or malformed DOCX can send headless LibreOffice into an infinite loop with 100% CPU, blocking the calling thread forever because --convert-to has no built-in timeout [1]. One user reported an Ubuntu web server hanging until reboot when unoconv's doc2pdf hit a bad DOCX [2]. Another saw LibreOffice 7.4.1.2 spin at 100% on certain DOCX files [3]. Even major open-source projects ship this blind spot: docling issue #3819 documents that its LibreOffice conversion path calls soffice with no timeout, so a hung process blocks forever [1].

Converterer's workaround is wrapping every invocation in timeout 120 soffice ..., which means you're already in the business of killing processes you spawned. That's not a conversion pipeline. That's process management with extra steps.

Silent Failures: No Output, No Error, No Clue

Headless soffice often exits 0 or 1 with no console output and no PDF file, even when a running LibreOffice instance or Quickstarter module is present [4]. Converterer's cheatsheet notes you may see the terse Error: no export filter, meaning the import type has no default mapping to PDF. Converterer's production guidance is to not trust exit codes and instead verify the output file exists and is non-empty — an admission the CLI isn't production-grade.

Mixed page formats in a single DOCX can be flattened to one orientation, silently breaking layout expectations [5]. Your document has a portrait cover page followed by landscape tables? Headless LibreOffice tends to flatten that to one size for the whole export [5]. You only discover this when a customer complains their report looks wrong.

The Concurrency Lock: Why Parallel Requests Randomly Vanish

Converterer documents that two soffice processes cannot share a user profile; the second exits 1 and produces no output file with no meaningful error message. Converterer's fix is passing -env:UserInstallation=file:///tmp/lo_profile_$$ per instance, but parallelism means multiple instances, not threads, because conversions are sequential within one instance. Docling #3819 proves this bites real projects: it doesn't pass -env:UserInstallation, so concurrent conversions silently fail due to the exclusive profile lock, with images or charts just skipped [1]. Running N instances with separate profiles risks lock contention and corrupted profiles if anything is shared [5].

You're not building a document converter anymore. You're building a process orchestrator with profile isolation, lock handling, and cleanup logic.

Speed and Deployment Traps: Bloat, Fonts, and Cold Starts

Converterer's benchmarks put a single conversion at roughly 2-3 seconds on a small VPS, dominated by startup. Converterer's own benchmarks show batching four files into one invocation drops total time from 13.52s to 5.96s, which fights against per-request architecture. A minimal Debian install for headless conversion has no official support; trial-and-error package selection still pulls in fonts, X, and desktop packages, and Docker images exceed 1GB [6]. Converterer notes missing server fonts silently shift layouts and page breaks as substitutes kick in; the fix is OS-level font archaeology. AWS Lambda packages like @shelf/aws-lambda-libreoffice add 250MB layers, cold starts, and orphaned soffice process leaks [7].

Fidelity and Security: PPTX Breakage and Untrusted Uploads

Headless PPTX-to-PDF has a reputation for breakage with charts and complex layouts, reflected in Reddit threads asking if the feature is broken outright [8]. DOCX fidelity issues include field updates, TOC formatting, header/footer alignment, and unsupported features like alt chunks [9]. Converterer's hardening guidance for parsing user-uploaded files is sandboxing, unprivileged users, capped CPU/memory, and no outbound network.

The Same Job as One API Call

MarkupGo's Office to PDF API converts 130+ formats including DOCX, XLSX, and PPTX, currently in BETA [10]. Upload as multipart/form-data, get back a task object, or append /buffer for a PDF buffer directly [10]. Set page properties like landscape orientation and page ranges in the request body; merge multiple documents into one PDF; control image quality with DPI options [10]. Each file is limited to 5MB and 10 files per request, with larger limits available on contact [10].

The markupgo-node client gives you office.convert with files as paths or streams, returning JSON task or buffer output [11]:

import MarkupGo from "markupgo-node";
import fs from "fs";

const markupgo = new MarkupGo({
  API_KEY: process.env.MARKUPGO_API_KEY,
});

const files = [
  "./invoice.docx",
  {
    data: fs.readFileSync("./report.xlsx"),
    ext: "xlsx",
  }
];

// Single PDF, merged
markupgo.office.convert(files, { merge: true, properties: { landscape: true } })
  .buffer()
  .then(buffer => {
    fs.writeFileSync("merged.pdf", Buffer.from(buffer));
  });

No process management. No profile isolation. No font archaeology. No hung processes to kill.

When Self-Hosting Still Makes Sense

Air-gapped environments where no external API is permissible. Extreme volume where per-request API costs exceed dedicated hardware. Situations requiring full control over the conversion binary and its sandbox. For everyone else, the operational burden of headless LibreOffice is a tax you don't need to pay.

Sources

  1. github.com
  2. Converting docx in headless mode hangs (ask.libreoffice.org)
  3. LibreOffice version 7.4.1 hangs while converting docx to pdf (ask.libreoffice.org)
  4. stackoverflow.com
  5. dev.to
  6. unix.stackexchange.com
  7. Convert DOCX to PDF Programmatically: AWS Lambda & LibreOffice (dev.to)
  8. Is headless PPTX -> PDF conversion broken? (reddit.com)
  9. Lack of good libraries doing DOCX to PDF (reddit.com)
  10. markupgo.com
  11. markupgo.com