Background
Tagged is not conformant
Last updated 30 August 2026
Nearly every PDF library on the market will tell you it produces tagged PDFs. Very few will tell you whether the tagged PDF it produced actually conforms to PDF/UA-1, because those are different claims, and the second one is much harder to make.
This page explains the difference, because it is the single most common reason a team believes it has solved document accessibility and has not.
What a tag is
A tagged PDF carries a structure tree: a parallel document
alongside the visual page, describing what each piece of content is
rather than where it sits. A heading is marked /H1. A paragraph is
/P. A table is /Table containing /TR rows
and /TD cells. Assistive technology reads that tree, not the
painted glyphs.
Without it, a PDF is a set of instructions for placing marks on a page. A screen reader encountering an untagged PDF has to guess at reading order from coordinates, and it guesses badly on anything with columns, sidebars or tables.
So tagging is where accessibility starts. It is not where it finishes.
What conformance adds
A tagged PDF can be structurally complete and still fail PDF/UA-1 on a dozen separate counts. The tags being present says nothing about whether they are correct, ordered, labelled, or matched to what is actually on the page.
These are the failures we see most often in documents that were, correctly, described by the tool that made them as tagged:
- Heading levels that skip. An
/H1followed by an/H3is a broken outline. A reader navigating by heading has no way to tell whether they missed a section or the document is malformed. - Tables without header associations. A
/Tablewhose cells are all/TDhas no header row. Every cell is then an orphaned value, and a screen reader cannot announce which column a number belongs to — which, on a statement or an invoice, is the entire content of the document. - Figures with no alternative text. A
/Figuretag with an empty or absent/Altentry is a hole in the document. Worse is a decorative rule or spacer image tagged as a figure at all, when it should have been marked as an artifact and removed from the reading order entirely. - No document language, or no title. Without
/Lang, a screen reader does not know which pronunciation rules to apply and may read English in a Dutch voice. Without a title in the document catalog, and the flag telling viewers to display it, every window shows a filename. - Fonts that are not embedded. The text renders on the machine that made it and substitutes on every other one, which changes glyph widths, breaks the mapping back to characters, and can make the text unextractable.
- Content outside the structure tree. Anything painted but not tagged and not marked as an artifact simply does not exist to a reader. It is the most invisible failure of the lot, because the page looks complete.
None of those are exotic. They are what ordinary HTML turns into when it is converted by a tool whose job stopped at emitting tags.
How anyone knows
PDF/UA-1 is ISO 14289-1. The PDF Association publishes the Matterhorn Protocol, which decomposes the standard into 31 checkpoints and 136 failure conditions — a testable form of a document that is otherwise prose.
The Protocol is also honest about its own limits, and this is the number worth carrying around:
| Failure conditions | Count | Who can decide |
|---|---|---|
| Determinable by software alone | 87 | A machine |
| Require human judgement | 47 | A person |
| No test defined | 2 | Neither, yet |
| Total | 136 |
A machine can prove that a /Figure has alternative text. No
machine can tell you whether the alternative text is any good — whether "chart"
describes the chart. That judgement is yours and it does not go away.
Which means any tool claiming to certify a document as accessible is overclaiming. The most a tool can do is settle the 87 conclusively, and say clearly that it has done only that.
Where this leaves generated documents
If you produce documents at volume — statements, invoices, benefit notices, reports — the tagging question is decided once, in your pipeline, for every document you will ever send. Get it right there and the problem is closed. Get it wrong there and you own a remediation bill that scales with your customer count, at market rates of roughly $5 to $25 per page.
That asymmetry is the whole argument for fixing it at generation. Nobody remediates half a million statements.
What Wellform does about it
Wellform takes HTML and returns a PDF that conforms to PDF/UA-1, and
attaches a conformance report to every document saying
which checks ran and what they found. Set "verify": true and the
document is additionally handed to veraPDF — the open source
reference implementation the standards bodies maintain — and the report carries
veraPDF's independent verdict next to ours, along with whether the two
agree.
We report the 87. We say so in those words, in the report, on every document. What we will not do is tell you a file is accessible, because that is not a thing we can know.
If your documents are currently produced by a headless browser, the specific reason they are failing is probably covered here.
Find out what your HTML actually renders as. Post a document and read the report. It names every machine-checkable condition the file failed, by its Matterhorn number, so you are arguing with evidence rather than with a vendor.
100 documents a month on the free plan. No card, and nothing expires.