Two Experiences — descriptive transcript
The transcript is in English because the video narration is English; the video’s captions are also available in other languages.
Video: a11yfy — "Two Experiences"
Duration: 0:36 (36.0 s) · Language: English (US)
This is a descriptive transcript: it carries both the spoken audio and the information that is only shown on screen, so it works as a standalone alternative to the video (WCAG 2.2 SC 1.2.3 / 1.2.8).
Voices: two, and the difference between them is the point of the video.
- Narrator — a synthetic American English voice.
- Screen reader — an actual screen-reader synthesizer (eSpeak NG, the voice NVDA ships with), flat and robotic. It speaks the lines a real screen reader would announce for the document on screen.
Sound: whoosh transitions, a low hum during processing, a dull thud when the screen reader hits the image wall, and a rising tone under the closing logo.
Constant on screen: the A11YFY logo in the top-right corner; the social cuts carry burned-in captions along the bottom; the web cuts render clean and the player supplies the same caption text from a WebVTT track.
Layout note: the video is built around a split screen. In the 16:9 cuts the two sides sit left (sighted reader) and right (screen reader); in the 9:16 cuts the same two panels are stacked top (sighted reader) and bottom (screen reader). The content is identical.
0:00 – 0:04 · One document, two sides
Visual: A split screen. Both halves are labelled — "👁️ Sighted reader" and "🔊 Screen reader" — and both show the same white document card titled "Annual Report" with grey placeholder lines. A dark chip spans the divide: "One document. Two experiences."
Narrator: This is one document. Two very different experiences.
0:04 – 0:11 · A page, or a photo
Visual: On the sighted-reader side, the document's lines highlight in indigo one after another as they are read. On the screen-reader side, the same page is revealed for what it really is: a dense field of black-and-white pixels, with no text at all.
Narrator: To a sighted reader, it's a page. To a screen reader, it's a photo — just pixels.
0:11 – 0:15 · The wall
Visual: The background turns red. Headline: "That's all a screen reader can say." A red scan line sweeps down the pixel page, and four dark speech bubbles pop out beside it in turn, each reading "Image."
Screen reader: Image. Image. Image. Image.
(then silence)
0:15 – 0:23 · Recovering the words
Visual: The a11yfy diamond mark pulses at the top of the frame. Below it, a document card: its grey placeholder lines turn, one by one from the top down, into solid indigo lines of real text.
Narrator: a11yfy recovers the words underneath and rebuilds the structure — turning pixels back into real, readable text.
0:23 – 0:27 · The same page, read properly
Visual: A document card titled "Applicant details" with a four-column table; the header row is indigo. A blue focus rectangle moves from the heading down onto the table, the way a screen-reader cursor would. A small speaker icon sits in the corner of the card.
Screen reader: Heading. Applicant details. Table, four columns. Row one.
0:27 – 0:32 · What that means
Visual: The same card, focus rectangle resting on the table.
Narrator: Now the same page can be read, searched, and navigated.
0:32 – 0:36 · End card
Visual: The a11yfy diamond mark — an indigo diamond with a coral folded top corner — rotates into place above the A11YFY wordmark. The slogan "Accessibility, automated." rises beneath it, a coral rule sweeps out under the line, and the address A11YFY.COM fades in.
Narrator: a11yfy. Accessibility, automated.