Alt text best practice says every image needs a description. We opened 540 real ones to see what's there.

Ask anyone who has touched a content management system what "accessibility" means and alt text is usually the first thing they say — the one rule with a checkbox in nearly every publishing tool. That reputation is exactly why it's worth measuring rather than assuming: in an earlier piece on this same document sample, we found that in 59% of the files that actually contain an image, that image has no alternative text at all.

10 min read

On this page
  1. What we measured, and how
  2. What is actually inside /Alt
  3. What is actually written when there is something there
  4. Length is not the signal
  5. The gap a validator cannot close
  6. What alt text best practice actually means, case by case
  7. An honest limit on what we measured
  8. Checking your own documents

This piece is about the other 41%. We went back into the same PDFs and asked the harder question: when alt text is there, what does it actually say?

What we measured, and how

The corpus is the same 111 born-digital, public-sector PDFs, downloaded from a crawl of Hungarian municipal websites, one per site so no single office's habits dominate. The 59% figure above comes from running veraPDF against clause 7.3-1 on the 59 of those files it could validate. But presence or absence of alt text doesn't need a validator — only a structure tree — so for this piece we read all 111 files directly.

We opened every PDF with pikepdf, walked its structure tree from /StructTreeRoot down, and recorded the /Alt value on every element tagged /Figure:

uv run --quiet python3 -c "
import pikepdf

def walk(elem):
    if isinstance(elem, pikepdf.Dictionary) and elem.get('/S') == pikepdf.Name('/Figure'):
        yield elem.get('/Alt')
    kids = elem.get('/K') if isinstance(elem, pikepdf.Dictionary) else None
    if kids is None:
        return
    kids = kids if isinstance(kids, pikepdf.Array) else [kids]
    for child in kids:
        if isinstance(child, pikepdf.Dictionary):
            yield from walk(child)

pdf = pikepdf.open('document.pdf')
root = pdf.Root['/StructTreeRoot']['/K']
root = root if isinstance(root, pikepdf.Array) else [root]
for top in root:
    for alt in walk(top):
        print(repr(str(alt)) if alt is not None else None)
"

We also checked /ActualText, which can stand in for /Alt on a figure, and the alternate description on annotation-based images. Neither changed the picture below.

What is actually inside /Alt

Sixty-two of the 111 files have a structure tree at all; 24 of those tag at least one image as a /Figure. Across those 24 files we found 540 figure elements. Here is what their /Alt attribute contained:

What we found in /Alt Figures Share
No /Alt entry at all 240 44%
Present, but an empty string 235 44%
Present, but a filename, URL, or local file path 6 1%
Present, but a single character or word 3 1%
Present, under 15 characters 2 <1%
Present, a full sentence or more 53 10%
Total 540
  • 44% No /Alt entry at all
  • 44% Present, but an empty string
  • 2% A filename, a word, or under 15 characters
  • 10% A full sentence or more

Grid of 100 squares, each square one per cent of the 540 figure elements found. 44 per cent have no /Alt entry at all. 44 per cent have an /Alt entry that is an empty string. 2 per cent hold a filename, a single word, or fewer than 15 characters. 10 per cent hold a full sentence or more.

One square is one per cent of 540 figures. Of the 53 figures with a real sentence, two read as human-written — 0.4%, less than a single square in this grid.

88% of figures carry nothing usable — missing or blank. That much lines up with the 59% file-level figure from the validated sample; a busy document with dozens of images tends to leave most of them untouched even when a handful get attention. The interesting part is the last row: the 53 figures that got a real sentence.

What is actually written when there is something there

Two figures out of 540 — 0.4% — read as though someone looked at that particular image and described it. Both sit in the same 59-page regional report: short captions naming a municipal coat of arms, twenty-one and fourteen characters, accurate and on topic. We cannot confirm from the file who wrote them. What we can say is that their content breaks from every automated pattern we found everywhere else in the sample, which is the basis for counting them as human-written.

The other 51 were not written by anyone looking at the image.

Where the sentence actually came from Figures
An automatic image-recognition caption, with its own AI disclaimer left in 36
An automatic chart caption restating only the question title and reply count 15

Take one file: a one-page notice with three images, none of them missing alt text. veraPDF's clause 7.3-1 — figures must have alternative text — reports zero violations on it. Here is the first /Alt value, translated from Hungarian:

"The image shows outdoor, plant, grass, sky. AI-generated content may be incorrect."

That disclaimer is the signature of an automatic captioning feature, not a person. Word and PowerPoint can generate alt text for an inserted picture using an image-recognition model, and the feature says so, in the text itself, every time. Nobody edited it before the file was published, and nobody deleted the warning either. A longer municipal newsletter repeats the pattern nine times over — "outdoor, tree, plant, sky", "outdoor, text, tree, sign" — generic scene tags, never what the photograph is actually of.

The chart captions, from that same 59-page report, have a different but equally hollow shape:

"Forms response chart. Question title: What is your gender? Number of responses: 573 responses."

That sentence tells a screen reader user a chart exists and how many people answered it. It never says what they answered — the one thing the chart was drawn to show. A sighted reader glances at the bars and knows the split in a second; a screen reader user gets the metadata and nothing else.

Two smaller absurdities are worth naming: one figure's alt text is nothing but a full local file path — C:\Users\...\Desktop\...\[town]_crest.jpg — someone's own folder structure, left in from a copy-paste. Another holds the caption a Google Images results page shows under a thumbnail: "Image result for: '[company] logo'" — describing how the picture was found, not what it shows. Both are real values, unedited, from published government documents.

Length is not the signal

If your instinct is to check alt text by length — flag anything under, say, twenty characters — this corpus would defeat that check. The 53 populated values have a median length of 102 characters, which isn't short. It's exactly as long as the boilerplate that makes up most of it: the auto-caption and chart-metadata sentences are both comfortably wordy, while saying nothing about the specific image. A character count cannot tell prose from a template.

The gap a validator cannot close

This is the asymmetry worth sitting with. veraPDF's clause 7.3-1 asks one question: does this figure have an /Alt entry, or a replacement, that isn't empty? It has to ask exactly that, because "not empty" is the only thing a machine can check without understanding both the image and the sentence. Whether the sentence is about the image — whether it tells someone who can't see the picture what a sighted reader knows — is a judgement call, and no automated check makes judgement calls.

The file above passes 7.3-1 cleanly: every figure has a non-empty /Alt. A procurement checklist that asks "does this PDF have alt text?" would tick the box. A screen reader user who reaches that image hears "outdoor, plant, grass, sky" and a warning that it might be wrong — no closer to knowing what the picture actually is than if the field had been empty, and arguably worse, because the sentence sounds like an answer.

What alt text best practice actually means, case by case

The W3C Web Accessibility Initiative's decision tree for writing alt text starts by asking what kind of image this is, because the answer changes what belongs in the field — and every failure mode above maps onto one branch of that tree.

Decorative images

A rule, a border, a stock photograph filling space should carry no alt text at all, and in PDF terms that means being marked as an artifact rather than tagged content, not a /Figure with an empty string. The difference matters to a screen reader: an artifact is skipped silently, while an empty-string figure is still announced as an image with nothing to say. A third of our sample's figures carry that empty string, and the file alone can't tell us how many were meant that way versus simply left blank.

Charts and diagrams

These carry a finding, and the alt text has to carry it too. W3C's guidance on complex images asks for a short summary of what the image shows, with the full detail available nearby — not a restatement of the chart's own title and sample size, which is exactly what our sixteen response-chart captions did instead. "Two-thirds of respondents said they exercise regularly" earns its place. "Chart, 573 responses" does not.

Logos

A logo becomes a functional image rather than a decorative one the moment it doubles as a link: the alt text should name where the link goes, not describe the artwork. A filename or a search-engine caption does neither.

Text inside an image

A letterhead, a badge with a word baked into the graphic, has one specific rule: the alternative must contain the same text that's in the image. We wrote up a real case of that failure — a GDPR badge with the word printed on it and nothing describing it — in our piece on what's actually inside a tagged PDF.

An honest limit on what we measured

We can only see figures inside a structure tree. Forty-nine of the 111 files have none at all, so every image in them is invisible to this method — our 540 figures come from the minority of files that were tagged in the first place. The real picture across everything a municipality publishes is very likely worse than the numbers above, not better. We also couldn't tell, from the file alone, which of the 235 empty-string figures were a deliberate "this adds nothing" decision and which were simply skipped — PDF/UA permits both, and only opening each image would settle it. We didn't open all 540 by eye; that's a real gap in this measurement, not a rounding error.

Checking your own documents

Where to look, without a terminal

Microsoft Word

Right-click an image and select View Alt Text. If Word generated the description automatically, it will say so, and — per Microsoft's own instructions — asks you to review and approve it before it ships, not accept it unread. For a genuinely decorative image, tick Mark as decorative instead of leaving the field blank.

Adobe Acrobat Pro

Accessibility → Reading Order, then double-click a figure on the page to edit its alternate text with the image itself visible next to the field — the fastest way to check whether a description actually matches what's on screen.

No installation

Our PDF accessibility checker reports every figure missing alt text under the same clause veraPDF does; the PDF/UA error reference explains what the clause requires and what a validator can and can't tell you about it.

None of this needs a specialist. It needs someone to look at the image and write one sentence about what it actually shows — which is the one step that, across 540 real figures in real published documents, essentially never happened.

Zoltán Csordás

Founder, a11yfy

Zoltán Csordás is an accessibility engineer and the founder of a11yfy. He builds the PDF/UA remediation pipeline behind a11yfy.com and spends most of his week inside tag trees, veraPDF reports and screen-reader output. In 2026 he measured 8,469 PDFs across 1,682 Hungarian municipal websites to find out how bad the problem really is. Based in Budapest.