An untagged PDF isn't unreadable. It's read in an order nobody chose.

"A screen reader might not be able to extract words, sentences and paragraphs in a coherent order. Instead, they may be mixed together in disconnected, confusing ways." That is not our description — it is Adobe's own official guide for screen-reader users describing what happens when a PDF has no tags. Not "cannot be read." Mixed together.

9 min read

On this page
  1. What reading an untagged document with assistive technology actually involves
  2. The measurement
  3. What each order actually produces
  4. Tables make the same problem worse
  5. What changes once the file is tagged
  6. The honest limit of this measurement
  7. What this means if you publish PDFs

That distinction is the whole subject of this article. Publishers and content teams routinely say an untagged PDF is "unreadable" for someone using a screen reader. It almost never is. The file opens, the text comes out, a voice reads it aloud — just not necessarily in the order a sighted person would follow, and not without page furniture wandering into the middle of a sentence. That is a worse problem to explain than "unreadable," which is probably why the myth persists. We measured what actually comes out, from a real file, and show it below.

What reading an untagged document with assistive technology actually involves

PDF supports "tagged" structure: an explicit, author-defined reading order that Adobe's own guide describes as marking "portions of PDF content" and organising them "in sequence," so that "assistive technology such as JAWS and Window-Eyes screen readers interpret the PDF tags in files viewed in Adobe Reader (or Acrobat)." When a file has that structure, a screen reader follows it, full stop. When it doesn't, the reader has to guess, and the same guide is unusually candid about what that guessing involves: when you open an untagged PDF, Acrobat presents a "Reading Untagged Document" dialog offering three named strategies —

  • Infer Reading Order From Document (Recommended) — "Adobe Reader determines the reading order of untagged documents using an advanced method of layout analysis." This is the default, and it is a proprietary heuristic; Adobe does not publish how it works.
  • Left-To-Right, Top-To-Bottom Reading Order — "Adobe Reader interprets the reading order of untagged documents starting from left to right on the page and moving from top to bottom."
  • Use Reading Order In Raw Print Stream — "Adobe Reader interprets the reading order of untagged documents based on the raw code in the PDF document – essentially, the order text was recorded in the print stream."

JAWS and NVDA both sit on top of this same choice when a PDF opens in Acrobat Reader. The EU Publications Office's official NVDA testing guide describes the same default in almost identical language: NVDA "will interpret the reading order of untagged documents by using advanced layout analysis" — the same inference strategy, the same lack of a guarantee.

Two of those three strategies are things we can reproduce exactly, outside any screen reader, because they are just text-extraction order. That is what we did.

The measurement

We used a real, publicly published Hungarian government PDF from a corpus of public-sector documents we've written about before — 034.pdf, a 34-page environmental-permit decision letter. pikepdf confirms it has no /StructTreeRoot at all: genuinely untagged, not just badly tagged.

pdfplumber (version 0.11.9) exposes both of Acrobat's non-default strategies directly, through one parameter:

import pdfplumber

with pdfplumber.open("034.pdf") as pdf:
    page = pdf.pages[0]
    raw_order = page.extract_text(use_text_flow=True)   # "raw print stream" order
    visual_order = page.extract_text()                    # left-to-right, top-to-bottom

use_text_flow=True skips pdfplumber's own position sort and returns characters in the order they appear in the page's content stream — the same definition Adobe gives for its raw-print-stream option. The default call sorts by vertical then horizontal position first, which is Adobe's second option in code form. Neither of these is Adobe's own renderer or NVDA's own speech engine; they are independent implementations of the same two documented ideas, which is the honest way to say what this measurement can and can't claim — more on that below.

What each order actually produces

The page has a standard government letter layout: an office footer block at the very bottom, a masthead, a two-column block of case reference fields, then the decision text.

Read in raw print-stream order, the footer comes first — before the masthead, before the case reference:

Országos Környezetvédelmi, Természetvédelmi és Hulladékgazdálkodási Főosztály
1016 Budapest, Mészáros utca 58/A.
Telefon: (06-1) 224-9100; KRID: 508260165
...
PEST VÁRMEGYEI
KORMÁNYHIVATAL
Ügyiratszám: PE/KTFO/1641-58/2026.
Ügyintéző:
Telefon: 06 (1) 224 9103
Tárgy: M100 gyorsforgalmi út M1 autópálya
csomópont – Esztergom közötti szakasz
...

("Ügyiratszám" is the case reference number, "Ügyintéző" the handling officer, "Tárgy" the subject line.) That footer block sits 770 points down a 842-point page — squarely in the bottom margin — yet it was the first thing drawn into the file, presumably because it comes from a fixed letterhead template applied before the case-specific text. A reader hearing this file top to bottom in print-stream order is told the department's phone number and email address before being told what the letter is even about.

Position-sorted order gets the footer right — it correctly lands at the very end of the page — but breaks the case-reference block instead, because two columns of fields sit at the same vertical position:

Ügyiratszám: PE/KTFO/1641-58/2026. Tárgy: M100 gyorsforgalmi út M1 autópálya
Ügyintéző: csomópont – Esztergom közötti szakasz
környezetvédelmi engedélyének
módosítása
Hiv. szám: -
Melléklet: -
Telefon: 06 (1) 224 9103

The case reference number gets glued to the start of the subject line on one row, and the (blank) handling-officer field gets glued to the subject line's continuation on the next — "Ügyintéző: csomópont – Esztergom közötti szakasz" reads as though the officer's name were "M1 motorway junction – Esztergom section," which is nonsense, because it's actually the tail end of an unrelated field one column over. Left-to-right, top-to-bottom is the more intuitive rule of the two, and it is still wrong here, for a document that is otherwise a perfectly ordinary single-column government letter.

Neither strategy is broken; both are doing exactly what Adobe documents. The letter simply doesn't carry enough information for either heuristic to recover which text belongs with which field — that information only exists in a structure tree, and this file doesn't have one.

Tables make the same problem worse

The same divergence shows up, more severely, on tabular exports. A municipal budget spreadsheet in the same corpus, printed to PDF from Excel, has a three-row stacked column header ("Cím-csoport szám," "Kiadási rovat neve," and so on). Position-sorted order reads across all three physical rows before it finishes any single column heading:

Cím- Alcím- Előir. Kiadási Feladat MÓDOSÍTOTT MÓDOSÍTOTT
csop. EREDETI
szám szám rovat jogcím előirányzat előirányzat

That's three different column headers interleaved into one incoherent line, because "read row by row" is the wrong rule for a header that spans multiple physical text rows per column. This is the mechanism behind the advice you'll see everywhere to avoid merged or stacked header cells: it isn't a stylistic preference, it's what happens to the text the moment there's no tag telling the reader which words belong to which column.

What changes once the file is tagged

Contrast that with a file that does have a structure tree: our own sample accessible PDF, walked with the same DFS logic a screen reader uses — every leaf's page, MCID and rendered text, in /K order:

1 0 H1  'Digital Accessibility'
1 1 P   'Why It Matters and How to Get Started'
1 2 P   'Digital accessibility means designing documents...'
2 0 H2  'Why Accessibility Matters'
2 2 Lbl 'Over 1 billion people worldwide live with some form of disability.'

There is no "which heuristic" question here, because there's nothing to infer. All 57 leaves across the file's 5 pages are real content — no footer, no page number, ever interrupts the flow, because that furniture is marked /Artifact and excluded from the structure tree entirely. A screen reader gets the same order every time, regardless of which of Adobe's three settings happens to be selected, because tags override all three.

The honest limit of this measurement

We did not run a live NVDA or JAWS session and listen to the output — this measures the text each reader receives, not how it's spoken, and it says nothing about pauses, punctuation announcement, or how a screen reader narrates a table cell's row and column position. And we have no way to reproduce Adobe's "Infer Reading Order" default at all: it's the recommended option, it's a proprietary algorithm, and it might handle this exact letter better than either of the two we tested. What we can say with certainty is that Adobe ships two fully-specified, documented fallback strategies for exactly this reason, and both of them audibly fail on an ordinary government letter — which is itself the point: even the vendor doesn't claim a reliable answer exists without tags.

If you want to hear the difference rather than read it, Acrobat Reader's own Read Out Loud tool (View → Read Out Loud) uses this same reading-order logic, so switching between the three options under Edit → Preferences → Reading → Reading Order and listening to the same paragraph each time reproduces what's shown above without installing anything else. NVDA users can do the same by opening the file and pressing its Say All command (NVDA+Down Arrow on the desktop layout) from the top of the page.

What this means if you publish PDFs

The fix isn't "add alt text" or "check contrast" — those matter, but they don't touch reading order. The fix is a structure tree that states, explicitly, which text is the footer, which two fields sit side by side, and which cell belongs to which column header — something no export dialog adds by default and no amount of careful visual layout substitutes for. Our PDF/UA error reference covers the specific checks a missing or broken structure tree fails, and the PDF/UA guide covers what the standard requires and why reading order is one of its clauses rather than a nice-to-have. The one thing worth doing before either: open a document you actually publish — not your most recent one — in Read Out Loud, and listen to the first paragraph. If it doesn't match what you'd say out loud yourself, you already know where to start.

Zoltán Csordás

Founder, a11yfy

Zoltán Csordás is an accessibility engineer and the founder of a11yfy. He builds the PDF/UA remediation pipeline behind a11yfy.com and spends most of his week inside tag trees, veraPDF reports and screen-reader output. In 2026 he measured 8,469 PDFs across 1,682 Hungarian municipal websites to find out how bad the problem really is. Based in Budapest.