Most published guides for creating accessible PDFs are written for the document that goes smoothly: a clean export, one chart, no tables, done in twenty minutes. That document exists. It's probably not the one on your desk.
Below is that same workflow with the missing part put back: what each step costs, whether a machine can do it, and — because we measured it rather than guessed — what share of a real batch of PDF/UA failures actually needs a person.
Start with what people actually download
Before any workflow step, there's a decision that isn't on the checklist: which file first. The instinct is to start with whatever was published most recently, because that's what's visible in the CMS. That's the wrong axis. A five-year-old fee schedule that gets downloaded every week does more harm sitting broken than a report nobody has opened since the day it went up.
If your CMS or web logs can tell you what actually gets downloaded, sort by that. If they can't, page views on the page that links to the PDF are a workable stand-in. Either way, the question is "who is affected right now," not "what did we just publish" — and it's the one step in this whole process that a machine can't answer for you, because it depends on what your organisation does, not on the file.
The workflow, step by step
"Machine" means software can do the whole thing without anyone deciding what the content means. "Person" means someone has to look at the document and make a call.
| Step | Machine or person | Roughly how long | How you verify it worked |
|---|---|---|---|
| 1. Work out what you actually have | Machine | Seconds per file | Structure tree, language, title — present or not |
| 2. Fix it at the source, if it still exists | Person, tool-assisted | Minutes to hours, by length | The authoring tool's own accessibility checker |
| 3. Export without wrecking it | Machine — a setting, not a task | Seconds, once it's a habit | The Producer field on the file you publish |
| 4. The mechanical repairs | Machine | Seconds to low minutes | Re-run the validator; the clause is gone |
| 5. The judgement calls | Person | Minutes per image or table | A second person reads it back |
| 6. Validate before you publish | Machine | About 3 seconds per document | The report itself, read past the top line |
1. Work out what you actually have
Before fixing anything, find out what's there and in what state. Born-digital or scanned? Does it declare a language? And does it have a structure tree — knowing that a tag existing and a tag being right are different questions, so don't stop at yes or no. This is the one part of the job that's entirely mechanical and near-instant: a script, or a free in-browser checker, can read these signals off hundreds of files in the time it takes to make coffee.
Treat the result as a triage list, not a verdict. A structure tree being present tells you the document was tagged at some point; it says nothing about whether the tags are right, which is where step 5 comes back in.
2. Fix it at the source, if it still exists
If you still have the Word or InDesign file, this is where a person earns their keep: real heading styles rather than bold 14pt text, alt text that says what an image means rather than that it's a picture, table headers marked as headers rather than just bolded. Word's own accessibility checker flags a fair chunk of this as you work — we've written separately about which chunk, and where it quietly stops.
For a one-page form, this step is minutes. For a fifty-page annual report, it's the better part of a day, because the reading is the work: you can't write alt text for a chart without reading what the chart says.
If you don't have the source any more — a scan, an old export, a file that arrived by email from someone who's left — this step doesn't exist for you. Go straight to step 4.
3. Export without wrecking it
Two settings matter, and neither is on by default everywhere: embed the fonts — we've written up exactly where that setting hides — and don't let anything compress or "optimise" the file afterwards. A single pass of a common PDF compression tool can strip the structure straight back out of an already-tagged file; we've measured exactly what that does, and it's worse than most people expect.
It costs nothing once it's a habit, and everything the day you discover a "shrink my PDF" step someone added to save storage now runs on every attachment automatically.
4. The mechanical repairs
This is where automation is worth using, because these fixes don't get slower with volume. Writing a PDF/UA identifier into the metadata, filling in a title, embedding a font, flagging a stray mark as decoration instead of content — none of it requires knowing what the document is actually about.
As one small, real example, not a general claim: one of the 59 municipal
Word exports from our earlier sample failed on exactly two checks — no
PDF/UA identifier, no dc:title in the XMP metadata stream. A short
pikepdf script wrote both fields in 8–23 milliseconds across five runs;
re-validated, the file went from two failures to zero. That's the easy end
of the pile — embedding a missing font, or classifying hundreds of untagged
marks in a document with no structure at all, is real engineering, not two
lines of code. None of it, easy or hard, needs a person to decide what the
document means, which is what separates it from step 5.
5. The judgement calls
Nothing skips this one. Deciding what an image is trying to communicate, whether a merged cell is a genuine sub-header or a formatting habit, whether a sidebar should be read before or after the column next to it — these need someone who can read the document, not someone who can operate a tool. It's minutes per image or table, and unlike step 4 it doesn't get cheaper by doing more of it.
This is also the step most guides for making a PDF accessible skip past with a line like "add alt text to your images," as though writing four hundred descriptions were the same kind of task as ticking a checkbox.
6. Validate before you publish, not after you write
veraPDF is the open-source validator for PDF/UA and PDF/A conformance. PAC is the free PDF accessibility checker for PDF/UA and WCAG. Run either against PDF/UA-1 on the file that's actually going on the website — after every step above, not after step 2. Across our own batch of 59 files, veraPDF 1.30.2 took 3.2 seconds per document on average (188 seconds for the batch, one file at a time). Validation is not the bottleneck in this workflow; steps 2 and 5 are.
A pass means the machine-checkable part is clean. It does not mean the document is good — that's step 5's problem, and no honest validator report claims otherwise if you read past the top line.
How much of it is actually mechanical
Every guide says some fixes are automatic and some need a person. Almost none of them says how much falls into which pile, so we counted it — using the same 59 validated Word-export PDFs from our earlier piece, Hungarian municipal documents checked against PDF/UA-1 with veraPDF 1.30.2.
The rule: a failure went into the mechanical pile if fixing it means writing a fixed, computable value that doesn't depend on what the document says — a font, an identifier, a metadata field, a structural flag. It went into the needs a person pile if fixing it means deciding what something means — alt text, a table's real headers, a link's alternate description, an irregular row of cells that might be a genuine sub-table or might be a mistake.
We classified all 23 distinct PDF/UA-1 rules the sample failed on this way, then counted how many files each rule hit. Total: 277 file-level failures. 184 of those (66%) were mechanical. 93 (34%) needed a person.
A single bar split in two. 66 per cent — 184 of 277 file-level failures in the sample — are mechanical and a tool can fix them. 34 per cent — 93 failures — need a person to decide.
That 66% sits close to a different figure already on this site: roughly two-thirds of the Matterhorn Protocol's ~136 failure conditions are machine-checkable. That's not the same claim — checkable means a validator can spot the problem, fixable means a machine can also resolve it. The two numbers landed in the same neighbourhood here; worth noting, not trusting as a rule.
How we counted it
veraPDF's JSON output reports each rule as a clause and testNumber (for
example, 7.1 / 9 for a missing dc:title), with a failedChecks count
and a ruleStatus. We summed the files affected by each rule across all 59
reports and sorted every rule into one of the two piles above.
~/.local/share/verapdf/verapdf -f ua1 --format json document.pdf
The line itself is a judgement call: table-structure rules (7.5-1,
7.2-42, 7.2-43, 7.2-10) went into "needs a person" because fixing an
irregular table means deciding what its cells were meant to say, not just
adding a wrapper around them.
At the file level, the same split reads differently, and for planning a backlog it's the more useful number: 19 of the 59 files (32%) had nothing wrong that needed a person at all — every failure in them was mechanical. The other 40 had at least one thing a machine can't decide on its own. If you're estimating how much of a backlog a script alone can clear before anyone has to open a file, that's the figure to use — not the 66%.
What this doesn't solve
Automating the mechanical two-thirds doesn't make the other third disappear. A five-thousand-document backlog with the same 34% profile still has getting on for two thousand files that need a person to open them, read them, and decide what an image or a table is actually saying. No tool — ours included — changes that arithmetic. It only changes how much is left once the easy part is done.
The classification is a judgement call too, worth admitting as one. Table-row
rules like 7.2-10 went in the "needs a person" pile because an irregular
table's meaning usually isn't obvious from the markup alone; a stricter
reading could call that mechanical instead. Move it and the split shifts a
few points either way — not enough to change the headline, but enough that
66/34 is a read of one dataset, not a constant.
Where to start
Not with a plan for all of it — a backlog treated as one undifferentiated pile is where these projects stall. Sort your files by how many people actually open them. Run the mechanical pass over everything, including the files that aren't urgent, because it's nearly free to do so. Reserve the actual person-hours for the third of the failures that need them, and spend those hours on the documents at the top of the download list first.
If you want the mechanical layer handled for you, see how our pipeline does it — a free analysis, price shown upfront. If you'd rather work through your own validator's report line by line, our reference explains what each PDF/UA clause means and whether it's the kind of thing a machine can fix on its own.