Wellform

Background

Why headless Chrome makes untagged PDFs

Last updated 30 August 2026

The usual way to turn HTML into a PDF is to drive a browser: Puppeteer or Playwright over headless Chrome, or the older wkhtmltopdf. It works, it is nearly free, and the output looks right.

Then someone runs the file through an accessibility checker and it fails, often on every document the pipeline has ever produced. This page is about why that happens, because the failure is structural rather than a matter of the wrong flag, and teams lose weeks looking for the flag.

The print path was never a document model

A browser's PDF export exists to reproduce what you would get from a printer. It walks the render tree — the boxes and glyphs after layout — and paints them onto pages. By the time that code runs, the information accessibility depends on has already been discarded.

The browser knows perfectly well that an element is a heading; it has a full accessibility tree, which is how a screen reader works on the live page. But that tree lives on the DOM side, and the print path does not consult it. Your <h2> arrives at the PDF as text at a size, in a weight, at a position. The fact that it was a heading is not encoded anywhere in the output.

This is why no combination of CSS produces a conformant PDF. You are styling the thing that gets painted. Conformance is about the structure tree that runs alongside it, and CSS has no vocabulary for that.

It has been asked for upstream, and declined

You do not have to take our word for any of this. The history is public, and reading it is the fastest way to understand that this is a design boundary rather than a backlog item:

That last one is the page in miniature. The flags exist, they were set correctly, and the output was still not accessible. There is no flag that fixes this, and the two reports that asked for it to be fixed were closed as not planned.

Chrome has gained partial tagged-PDF output over the years, which is genuine progress and also why this problem is so easy to misdiagnose: a file can come back with some structure and still fail conformance comprehensively. Partial tagging produces a checker report full of individual failures rather than a clean "untagged" verdict, and that reads like a configuration problem when it is not one.

The regression is the part worth internalising. Where tagging is a side effect of the print path rather than a supported contract, it can come and go with a patch release — and nothing in your pipeline will tell you. Your documents get quietly worse and your tests stay green.

wkhtmltopdf is a separate problem

If you are on wkhtmltopdf, tagging is not the most urgent thing on your list. The project's own status page is unusually candid: Qt 4, which it uses, has not been supported since 2015, and the WebKit inside it has not been updated since 2012. The same page carries the instruction "Do not use wkhtmltopdf with any untrusted HTML", and notes that doing so can lead to complete takeover of the server it runs on.

If your HTML is assembled from customer-supplied data — names, addresses, line items, anything a user typed — that warning is about you. It produces no structure tree either, but that is the second finding.

What actually has to happen

Conformant output requires something that holds the semantics of the source all the way through layout and writes them into the PDF deliberately:

That is a different job from printing a page, which is why it tends to need a different tool rather than a different flag.

How Wellform does it

Wellform is not a headless browser. Layout and text shaping are done by WeasyPrint and Pango, which keep the source structure available through layout; the structure tree and document metadata are written with pikepdf. The pipeline is deterministic — the same HTML gives the same PDF, byte for byte — and contains no model making a probabilistic guess about your content.

Every response carries a conformance report naming the checks that ran and what they found. Send "verify": true and the document is also checked by veraPDF, the reference implementation, whose independent verdict is reported next to ours together with whether the two agree. A disagreement is visible rather than smoothed over.

The report covers the 87 Matterhorn failure conditions a machine can settle. The remaining 47 need human judgement and always will — that distinction is worth understanding properly before you buy any tool in this space, including this one.

You can move a pipeline over by changing the endpoint you POST your existing HTML to. If it renders in a browser today, it is a single request to see what it scores.

Point your existing HTML at it and see. If it renders in a browser today it will render here. The difference is that the response tells you what the structure tree came out as, instead of leaving you to discover it from a checker months later.

100 documents a month on the free plan. No card, and nothing expires.