Background
Why headless Chrome makes untagged PDFs
Last updated 30 August 2026
The usual way to turn HTML into a PDF is to drive a browser: Puppeteer or
Playwright over headless Chrome, or the older wkhtmltopdf. It
works, it is nearly free, and the output looks right.
Then someone runs the file through an accessibility checker and it fails, often on every document the pipeline has ever produced. This page is about why that happens, because the failure is structural rather than a matter of the wrong flag, and teams lose weeks looking for the flag.
The print path was never a document model
A browser's PDF export exists to reproduce what you would get from a printer. It walks the render tree — the boxes and glyphs after layout — and paints them onto pages. By the time that code runs, the information accessibility depends on has already been discarded.
The browser knows perfectly well that an element is a heading; it has a full
accessibility tree, which is how a screen reader works on the live page. But
that tree lives on the DOM side, and the print path does not consult it. Your
<h2> arrives at the PDF as text at a size, in a weight, at a
position. The fact that it was a heading is not encoded anywhere in the
output.
This is why no combination of CSS produces a conformant PDF. You are styling the thing that gets painted. Conformance is about the structure tree that runs alongside it, and CSS has no vocabulary for that.
It has been asked for upstream, and declined
You do not have to take our word for any of this. The history is public, and reading it is the fastest way to understand that this is a design boundary rather than a backlog item:
- Export tagged PDFs for Accessibility (August 2021) — the original feature request. Confirmed by maintainers, now closed. It notes that Chrome could carry alternative text into a PDF when a person printed one manually, but that the same thing did not happen through Puppeteer. That gap is why you will find contradictory advice on this: the browser and the automation driver do not take the same path.
- Puppeteer produces untagged PDF starting from v23.8.0 (January 2025) — a regression. Output was tagged through 23.7.1 and stopped being tagged from 23.8.0 onward, still failing at 24.1.1. Closed as not planned.
- Unable to
generate accessible PDF (July 2025) — someone running Puppeteer 24.15.0
with
--export-tagged-pdfand--force-renderer-accessibilityboth set, whose images were still announced by VoiceOver as a generic "graphic" rather than by their alternative text. Closed as not planned.
That last one is the page in miniature. The flags exist, they were set correctly, and the output was still not accessible. There is no flag that fixes this, and the two reports that asked for it to be fixed were closed as not planned.
Chrome has gained partial tagged-PDF output over the years, which is genuine progress and also why this problem is so easy to misdiagnose: a file can come back with some structure and still fail conformance comprehensively. Partial tagging produces a checker report full of individual failures rather than a clean "untagged" verdict, and that reads like a configuration problem when it is not one.
The regression is the part worth internalising. Where tagging is a side effect of the print path rather than a supported contract, it can come and go with a patch release — and nothing in your pipeline will tell you. Your documents get quietly worse and your tests stay green.
wkhtmltopdf is a separate problem
If you are on wkhtmltopdf, tagging is not the most urgent thing
on your list. The project's own status page is unusually candid: Qt 4, which it
uses, has not been supported since 2015, and the WebKit inside it has not been
updated since 2012. The same page carries the instruction
"Do not use wkhtmltopdf with any untrusted HTML", and notes
that doing so can lead to complete takeover of the server it runs on.
If your HTML is assembled from customer-supplied data — names, addresses, line items, anything a user typed — that warning is about you. It produces no structure tree either, but that is the second finding.
What actually has to happen
Conformant output requires something that holds the semantics of the source all the way through layout and writes them into the PDF deliberately:
- A structure tree built from the source markup, not reconstructed from the painted result
- Heading levels carried across intact, without skips
- Table headers tagged as headers, with scope, so a cell can be announced with the column it belongs to
- Alternative text transferred from the markup, and decorative images marked as artifacts so they leave the reading order
- Document language and title written into the catalog, with the display-title flag set
- Every font embedded, with character mappings that survive extraction
- Nothing painted outside the structure tree or the artifact marking
That is a different job from printing a page, which is why it tends to need a different tool rather than a different flag.
How Wellform does it
Wellform is not a headless browser. Layout and text shaping are done by WeasyPrint and Pango, which keep the source structure available through layout; the structure tree and document metadata are written with pikepdf. The pipeline is deterministic — the same HTML gives the same PDF, byte for byte — and contains no model making a probabilistic guess about your content.
Every response carries a conformance report naming the
checks that ran and what they found. Send "verify": true and the
document is also checked by veraPDF, the reference
implementation, whose independent verdict is reported next to ours together
with whether the two agree. A disagreement is visible rather than smoothed
over.
The report covers the 87 Matterhorn failure conditions a machine can settle. The remaining 47 need human judgement and always will — that distinction is worth understanding properly before you buy any tool in this space, including this one.
You can move a pipeline over by changing the endpoint you POST your existing HTML to. If it renders in a browser today, it is a single request to see what it scores.
Point your existing HTML at it and see. If it renders in a browser today it will render here. The difference is that the response tells you what the structure tree came out as, instead of leaving you to discover it from a checker months later.
100 documents a month on the free plan. No card, and nothing expires.