That made the form worth measuring directly. So we went looking for real ones.
How rare a real, fillable public-sector PDF form actually is
We started with a sample of 111 public-sector PDFs pulled from Hungarian
municipal websites. Not one of them contained a fillable form field. Nine of
the 111 do carry an /AcroForm entry with widget annotations — but every
widget in every one of them is a /Sig field, a slot for a digital signature,
not something a person types into.
So we went wider, to the crawl those 111 came from: 1,682 Hungarian municipal websites, 8,469 PDFs. We searched the file names for the words a Hungarian public body uses for a form — kérelem, nyomtatvány, adatlap, bejelentő, űrlap, igénylő — and got 259 distinct files, of which we downloaded 248, one request per second, identifying ourselves honestly.
Of those 259 file names, 74 had already been flagged by our crawler as scanned
images rather than text. We could download and open 70 of them: zero carry an
/AcroForm. A scanned "form" is a photograph of a form — there is no field
object anywhere in the file for a screen reader, or anything else, to reach.
You print it, fill it by hand, and post it or scan it back. That's worth
knowing up front, because it means the accessibility question doesn't even
arise for close to a third of the documents whose file names promise a
fillable form.
Of the 248 files we could actually open, four contained real /AcroForm
widgets that a person fills in: 272 fields in total. Those four are what the
rest of this article is built on.
Funnel from the 2026 crawl of Hungarian municipal websites. 8,469 PDFs were crawled. 259 file names suggested a form. 248 of those could be downloaded and opened. Only 4 contained real fillable AcroForm widgets, 272 fields in total. None of the 4 conform to PDF/UA-1.
python3 -c "import pikepdf; pdf = pikepdf.open('file.pdf'); print('/AcroForm' in pdf.Root)"
What a screen reader reads, and what it never will
Every form field has an internal name, the /T key — Acrobat needs one to
address the field in code, so it invents one if the author doesn't. That name
is not a label, it's an identifier, and several of the fields we found still
carry the ones Acrobat's form-recognition wizard assigned by default:
Check Box1, Check Box3, or a bare 0 and 1 for a pair of adjacent
fields.
The field a screen reader actually announces is a separate key, /TU — the
tooltip. The W3C's own guidance on this is blunt about it:
"The TU entry (which is the tooltip) of the field dictionary is the programmatically associated label. Therefore, add a tooltip to each field to provide a label that assistive technology can interpret." — PDF10: Providing labels for interactive form controls in PDF documents, W3C WAI
The clearest contrast in our four files is between one made carefully and one that wasn't. A pair of near-identical applications for a protected-consumer electricity and gas tariff — same template, same municipality, one field for electricity and one for gas — carry a tooltip on 15 of their 17 fields:
/T (internal name) |
/TU (what gets announced) |
|---|---|
fill_2 |
"b) address (postcode, settlement, street/road/square, house number, staircase, floor, door)" |
a családi és utóneve |
"a) family and given name" |
(The original tooltips are in Hungarian; translated here for readability.)
A local-tax data-disclosure form from a different municipality has 123 widget
fields — 74 text boxes and 49 checkboxes — and not one of them carries a /TU.
A screen reader lands on a text box and can say only that it is a text box.
Tab order is a separate setting from visual order
A page in a PDF carries its own /Tabs entry, and it isn't inferred from
where anything sits on the page. Set to /S, the reader walks the fields in
the order they appear in the structure tree — the order a sighted user would
read the page. Left unset, or set to /W, it walks them in whatever order the
annotations were added to the page: the order someone drew the boxes in the
form-design tool, not the order a person reads top to bottom.
Two of our four files set /Tabs to /S on every page carrying a field. The
canvassing-sheet request form — 115 fields across five pages — sets it to /W
throughout. veraPDF 1.30.2 (build 2026-06-03) catches exactly this on all three
of that form's pages that carry a field, against ISO 14289-1 clause 7.18.3:
"errorMessage": "A page with annotation(s) contains Tabs key with value W
instead of S"
~/.local/share/verapdf/verapdf -f ua1 --format json Ajanloiv-igenylese.pdf
The practical effect: a mouse user reads the form as laid out. A keyboard or screen reader user on the same page can land on field fourteen, then field three, then field nine, with no way to tell — short of trial and error — whether they've reached the end.
The one that did nearly everything right, and still failed
The electricity and gas application pair is the best-made form in our sample:
88% of its fields have a usable tooltip, and its tab order follows the
structure. It's also produced by Microsoft Word, and Word's export marks the
document as tagged — StructTreeRoot present, MarkInfo/Marked true, the
markers a validator checks first.
Run it through veraPDF anyway, and it fails. Every one of its 17 fields fails
clause 7.18.4: a widget annotation has to be nested inside a structure element
of type Form, and in this file none of them are. We checked this
independently of veraPDF, with pikepdf directly on the structure tree: across
all four files, not one of the 272 widgets carries a /StructParent key, and
the tree contains zero object references to any of them. The tooltips and the
tab order were clearly done with care. They sit alongside a structural gap
that Word's form export gives no setting for, and that filling in every
tooltip conscientiously cannot close.
The other two forms fail the same clause on all 123 and all 115 of their fields, too. So it isn't that this is what happens when a form is made carelessly. It's what happens to every form in our sample regardless — the best-labelled one included.
Required fields, and the error message that never comes
None of the 272 fields across the four forms use the form field's required
flag — the /Ff bit a validator or assistive technology could actually
detect. If a field is mandatory, the only signal is whatever the form printed
next to it, typically an asterisk, which is exactly what the W3C recommends
when nothing better is available:
"Required fields are implemented using the /Ff entry in the form field's dictionary. ... If errors are found, an alert dialog describes the nature of the error in text. This may be accomplished through scripting created by the author." — PDF5: Indicating required form controls in PDF forms, W3C WAI
None of our four forms carry that scripting either — no embedded JavaScript, no field-level validation action, in any of them. So there is no error message. A required field left blank doesn't get flagged; the form simply submits, or doesn't, and the applicant finds out some other way, if they find out at all.
The honest part: a PDF form is a poor medium for this
We sell PDF remediation, so this isn't the easy thing to say, but it's true:
for an application process, an HTML form is very often the better tool. A
browser enforces tab order from the page's own DOM — there's no separate
setting to get wrong. required on an <input> is native, and the browser
announces it. A label bound with <label for> isn't a discretionary tooltip an
author might skip under deadline; leave it out and most authoring tools flag it
before publication. None of what we measured above is a competence gap
specific to these authors — it's the medium giving them more ways to get it
wrong, with no equivalent of a browser's built-in checks to catch it first.
That doesn't mean every PDF form should become an HTML one tomorrow. Some
genuinely need a wet or qualified signature, a printable trail, or an existing
case-management pipeline built around the format. For those, a correctly built
PDF form — /TU on every field, /Tabs set to /S, widgets actually nested
where PDF/UA-1 requires — is a real, achievable floor. It's just a floor all
four of the forms we found, including the best of them, currently sit under.
How to check your own PDF form without a validator
Adobe Acrobat (Pro, and Reader on some plans)
Tools → Prepare Form. Every field on the page gets a small tag with its
internal name. Double-click a field and open its General tab — the
Tooltip box is the /TU value. If it's empty, that's what a screen reader
has to work with: nothing beyond "text field" or "checkbox".
Tab order, in the same tool
Right-click a page thumbnail → Page Properties → Tab Order → Use Document Structure. If that option isn't already selected, the reading order a screen reader gets is whatever order the fields were drawn in, not the order they're meant to be read in.
Without Acrobat
Our PDF accessibility checker reports the same PDF/UA-1 clauses veraPDF does, field by field, without installing anything.
If you're publishing a PDF form anyway
Check the fields that get submitted the most, not the newest upload — a fee schedule or a benefits application gets opened far more often than last month's meeting minutes. Give every field a real tooltip, not the internal name the form tool assigned it. Turn on document-structure tab order. And if the form was ever run through a compression or "shrink my PDF" step afterwards — that alone can erase every one of these settings in a single pass, the same way it erases everything else in the tag tree.
If you want the clause-by-clause list for one of your own forms, our checker walks the same ground veraPDF does, and the PDF/UA error reference explains what each failure means and whether it's fixable automatically. Missing titles and unembedded fonts are covered in the companion piece on what Word's checker misses — worth reading, since all four of our forms started life in Word. For the standard itself, our PDF/UA guide covers what conformance requires and why.