Measurement
What a town's PDFs actually look like
Last updated 6 September 2026
The DOJ rule for state and local government web content has a date on it. Entities serving 50,000 people or more have until 26 April 2027; smaller entities and special district governments until 26 April 2028. The standard named is WCAG 2.1 Level AA, and PDF/UA-1 is the recognised route for the PDF part of it.
What nobody seems to have written down is what the documents on a real municipal website are actually like today. So we measured one.
The method
We crawled a single New Jersey township's website, following links from the homepage and obeying its robots.txt at one request per second. Every PDF we reached was downloaded and opened with pikepdf, and we recorded what the file says about itself: whether it carries a structure tree, whether MarkInfo marks it as tagged, whether it declares a language and a title, and whether its XMP claims conformance with any standard.
Sixty documents were found. Fifty-two were reachable and readable, and those are the ones reported here. No file was modified, and nothing was submitted to any third-party service.
These are properties of files, not verdicts about an organisation. A missing structure tree is an observable fact. Whether a particular document meets a legal obligation is a determination for that entity's counsel, and nothing here makes it.
What the documents say about themselves
| Property | Count | Share |
|---|---|---|
| Opened successfully | 52 | 100% |
| Has a structure tree (tagged) | 7 | 13% |
| Marked as tagged in MarkInfo | 7 | 13% |
| Declares a document language | 6 | 12% |
| Declares a title | 19 | 37% |
| Declares PDF/UA conformance | 0 | 0% |
| Encrypted | 0 | 0% |
A PDF with no structure tree has no headings, no reading order and no table semantics. A screen reader is left inferring meaning from visual layout. A missing language leaves it guessing at pronunciation; a missing title leaves it announcing a filename.
Where the files come from
This is the part we did not expect, and it is more useful than the percentages. Reading the producer recorded in each file's metadata:
| Producer | Tagged | Count |
|---|---|---|
| No producer recorded | no | 25 |
| Microsoft: Print To PDF | no | 16 |
| Adobe PDF Library 26 (via Acrobat PDFMaker for Word) | yes | 5 |
| Adobe LiveCycle Designer 8.0 | yes | 1 |
| Adobe PDF Library 5.0 (InDesign 2.0) | yes | 1 |
| Adobe PDF Library 9.9 / 15.0 | no | 2 |
| macOS Quartz PDFContext | no | 1 |
| RICOH IM C6500 (a copier) | no | 1 |
Sorting the same 52 files by what is actually inside them — whether a page carries fonts, which means real text, or only an image, which means a scan:
| Document | Count |
|---|---|
| Born-digital text, untagged | 27 |
| Scanned, no text layer at all | 10 |
| Mixed: some pages text, some scanned | 8 |
| Born-digital text, tagged | 7 |
What this means, including the part that is inconvenient for us
Wellform is an API that turns HTML into a tagged PDF. Almost none of these documents has any HTML upstream of it. They are Word files printed to PDF, a scan from the copier in the municipal building, a form drawn in LiveCycle twenty years ago. A rendering API does not fix this town's documents, and we would rather say so than sell something that does not fit.
Two things follow that are worth more than a sales pitch.
The first is that somebody in that building has already solved part of it without buying anything. The five most recent tagged documents came out of Acrobat PDFMaker for Word — an export path that tags automatically, on a licence the clerk's office evidently already holds. The gap between 13% and a much higher number is, for those files, a question of which menu item gets used.
The second is that ten documents are scans with no text layer. No generator and no export setting reaches those. They need OCR and remediation, which is a different product from a different kind of vendor, and it is the part of this problem that costs real money.
Where a rendering API does apply is upstream of all of this: systems that generate documents programmatically — statements, invoices, notices, reports produced from data at volume. That is a different buyer from a township of a few thousand documents, and this measurement is part of how we worked that out.
Caveats we would want from someone else's report
- One entity, 52 documents. A sample, not a census, and one town is one data point. We are measuring more.
- The crawler follows links. Documents reachable only through a search form or a client-side index were not seen. The real denominator is larger than 60.
- Metadata can lie. A producer string is what the file claims, not proof of provenance. Structure-tree presence is a stronger fact than any string.
- Tagged is not conformant. Seven files carry a structure tree; that does not make them PDF/UA-1, and none of them claims to be. The gap between the two is its own subject.
- The entity is not named. The documents are public and the method is published, so anyone can reproduce this on any site including their own. Naming a specific small government in a vendor's marketing seemed like the wrong trade for evidence that does not need it.
You can run the same check on a single file of your own, free and without an account. It reports what the document says about itself and, where a profile exists, what veraPDF says about it.
Check a PDF The deadlines in detail