Skip to main content
AIGC

Video Accessibility at Scale: Captions and Audio Description

July 2026 · 8 min read · Updated September 2026

Short answer. At minimum, prerecorded video needs captions for all audio content, a transcript or audio alternative, and audio description — the video-specific requirements in WCAG 2.2 at Level A and AA, the level most procurement and legislation reference. In Europe, the European Accessibility Act applies from 28 June 2025 to a defined set of products and services, with EN 301 549 operationalising it through WCAG. AI can draft captions and synthesise description, but accuracy, speaker identification, timing and non-speech information still need a human pass.

Key takeaways

  • WCAG 2.2 requires captions (Level A), a text alternative for time-based media (Level A), and audio description (Level A and AA) for prerecorded video; Level AA is the level most legislation and procurement reference.
  • Captions and subtitles are not interchangeable: captions serve viewers who cannot hear and carry speaker identification and non-speech sound, while subtitles translate dialogue for viewers who cannot understand the language.
  • The European Accessibility Act (Directive (EU) 2019/882) applies from 28 June 2025 to a defined set of products and services, with EN 301 549 as the harmonised standard that incorporates WCAG.
  • Automatic speech recognition produces a usable caption draft but still fails on proper nouns, speaker identification, and non-speech audio information without a human correction pass.
  • Audio description is far cheaper when gaps for narration are designed into the script, because most commercial video is written wall-to-wall with dialogue.

What does video actually have to provide?

WCAG is the substantive standard behind most legislation and procurement, so satisfying it at the stated level covers most of what the regimes ask for on the content side. WCAG (Web Content Accessibility Guidelines) is the W3C's standard, currently version 2.2, that most accessibility laws and procurement policies point to rather than writing their own criteria.

Requirement Level What it means in production
Captions for prerecorded audio in synchronised media A Accurate, synchronised captions carrying dialogue, speaker identity and meaningful non-speech audio
Alternative for time-based media A A transcript or equivalent text alternative covering the content of the media
Audio description for prerecorded video A and AA A narration track describing visual information the existing audio does not convey
Captions for live audio AA Real-time captioning for live streams — a different production discipline with its own suppliers
Contrast and visual presentation AA Applies to any text rendered into the picture, including burned-in titles and on-screen graphics
Audio control and background audio A and AAA Controls for anything that auto-plays; limits on background audio behind speech

Level AA is the practical target because it is what most legislation and procurement reference. Level A alone rarely satisfies a public-sector or regulated buyer.

Captions are not subtitles. Captions are for viewers who cannot hear the audio: they carry speaker identification and meaningful non-speech sound — a door slamming, a phone ringing, music that carries meaning. Subtitles are for viewers who cannot understand the language: they translate dialogue and assume the viewer hears everything else. A translated subtitle file does not satisfy a caption requirement, and that substitution is one of the most common failures in a library.

What changed in Europe?

The European Accessibility Act sets accessibility requirements for a defined set of products and services and applies from 28 June 2025, moving accessibility from a procurement preference to a market-access condition for covered categories.

The European Accessibility Act (Directive (EU) 2019/882) is an EU directive that makes accessibility a legal market-access condition, rather than a preference, for a defined set of products and services. EN 301 549 is the harmonised European standard for accessibility requirements for ICT products and services, and it incorporates WCAG for web content and documents. In practice, a producer meeting WCAG 2.2 Level AA on its media has done the substantive work the standard asks for on the content side; EN 301 549 also covers software, hardware and documentation outside a content team's scope.

Whether a specific organisation is in scope depends on the product or service and on national implementing legislation, which varies — this is a practitioner's summary, not legal advice, and scope should be confirmed with counsel. The safe operating assumption for a content team is that accessibility is a delivery requirement rather than an enhancement, and that retrofitting a library costs considerably more than producing accessibly from the start.

How do you produce captions that are actually usable?

Automatic speech recognition changed the economics of captioning, not the standard. ASR produces a good first draft on clean single-speaker audio and degrades exactly where accessibility matters most: proper nouns, technical terminology, overlapping speech, accents, and everything that is not speech at all.

The sequence below leaves ASR the mechanical work and people the work that decides whether the result is usable.

1. Transcribe with ASR, then correct against a term list. Feed in the production's pronunciation and terminology list and fix names, products and technical terms first — the highest-error and highest-salience items.

2. Add speaker identification. Who is speaking, wherever it is not obvious from the picture. This is part of what a caption is, and it is absent from every raw ASR output.

3. Add meaningful non-speech audio. Sounds that carry information — a knock, an alarm, laughter, music that signals a change. Describing every ambient sound is as unhelpful as describing none; the test is whether a hearing viewer would take meaning from it.

4. Segment for reading, not for grammar. Break at meaning boundaries so each caption is comprehensible on its own. Netflix's published English Timed Text Style Guide is a usable industry reference: up to 42 characters per line, no more than two lines on screen, and an adult reading speed of 20 characters per second.

5. Time to the audio and respect the minimums. Captions appear with the speech and hold long enough to read — a minimum event duration and a maximum, so nothing flashes and nothing lingers. Timing errors are the fastest way to make technically correct captions unusable.

6. Position around important picture content. Move captions off burned-in titles, faces and lower thirds. That is a placement decision automatic tools do not make.

7. Deliver as a sidecar file. SRT, WebVTT or TTML rather than burned in, wherever the platform supports it — editable, indexable and switchable by the viewer.

8. Review in context. Watch the video with the captions on. Errors invisible in the caption file — a caption covering the thing being discussed, a line landing after the cut — are obvious within thirty seconds of playback.

What does audio description actually require?

Audio description is a narration track conveying visual information the existing audio does not: who is on screen, what they are doing, what is written on screen, what changed. It is a WCAG requirement for prerecorded video, and it is the one most often quietly skipped, because it is the one that costs real production effort.

The difficulty is structural rather than technical. Description has to fit in the gaps between dialogue, and most commercial video is written wall-to-wall with narration. A video with no gaps cannot be described without either an extended-description version that pauses the picture, or a re-edit.

  • Design for description at the script stage. Leaving deliberate gaps costs nothing while writing and is expensive to create afterwards. This is the highest-leverage decision in the whole area.
  • Describe what matters, in priority order. Actions and on-screen text first, then setting, then detail. Description competes for time with the programme audio, and there is never enough of it.
  • Read on-screen text aloud. Titles, captions, statistics and any text rendered into the picture. This is the most commonly missed content and often the most information-dense.
  • Synthesis is appropriate here. Described audio is functional narration, which is exactly the register where synthetic voice performs well, and it is what makes description affordable across many languages — provided the content labelling and disclosure rules that apply to synthetic audio are followed.
  • Deliver as a separate track or version, per platform capability, so viewers can choose it.

How does this work across many languages?

Accessibility and localisation are the same production problem viewed from two angles, and treating them as one workflow is what makes both affordable. Both operate on a locked picture, both produce sidecar text files, both need in-market human review, and both draw on the same terminology asset.

  • One timed-text source, many derivations. The caption file is the base artefact; subtitles in other languages derive from it, and so does the transcript. Producing them independently duplicates the timing work, which is the expensive part — the same principle behind localizing one video into many languages.
  • Reuse the pronunciation lexicon across synthetic described-audio tracks in every language, exactly as for multilingual AI voice production, and draw it from a licensed voice library rather than an unlicensed one.
  • Check reading speed per language. The same sentence expands differently across languages, so caption timing that is comfortable in one is unreadable in another. That is a script-fitting decision, not a subtitling one.
  • Have an in-market reviewer watch it. Register, terminology currency and cultural readability are not visible in a caption file, and they are the difference between conformant and usable.
Reading speed = Characters in the caption ÷ Seconds the caption is on screen

Twenty characters per second is the adult ceiling in the Netflix guide for English. Recompute it per language rather than carrying the English figure across: a file that passes the character-per-line rule can still run too fast to read.

How does Lifewood handle video accessibility at scale?

Lifewood produces captions, subtitles, transcripts and described audio inside the same localisation pipeline rather than as separate services, so the timing work is done once and derived across every language. Each language variant gets region-native review, because register and terminology currency are not visible in a caption file.

Coverage runs to 100+ languages across 40+ delivery centres in 30+ countries, under a 95%+ accuracy SLA with human review on every variant — checked through the same quality-control process applied across AI-generated content, which matters most on the ASR correction pass, where errors cluster on exactly the names a viewer relying on captions most needs to be right. This sits within Lifewood's broader AIGC video production and AIGC services work.

Frequently asked questions

No. ASR produces a usable draft and fails predictably on proper nouns, technical terms, overlapping speech and accents, and it does not produce speaker identification or non-speech audio information at all — both of which are part of what a caption is. The defensible workflow is ASR plus human correction, speaker labelling, non-speech annotation and in-context review.

Captions serve viewers who cannot hear: they include speaker identification and meaningful non-speech sounds. Subtitles serve viewers who cannot understand the language: they translate dialogue and assume the rest of the audio is heard. Supplying translated subtitles where captions are required is a common and consequential substitution error, and it is the fastest diagnostic to run over an existing library.

For prerecorded video, WCAG 2.2 includes audio description at Level A and AA, and Level AA is what most legislation and procurement reference. It is the requirement most often skipped because it is the one that costs production effort — and the cost is far lower if gaps are designed into the script rather than found in a finished edit.

It applies from 28 June 2025 to a defined set of products and services, with scope depending on the category and on national implementing legislation. Rather than resolving the scope question first, most content teams find it cheaper to produce to WCAG 2.2 Level AA as standard — that covers the substantive content requirements either way and removes the need to maintain two production standards.

Yes, and it is a good fit. Described audio is functional, factual narration in a neutral register, which is where synthetic voice performs best, and it is what makes description affordable across many languages. A wholly synthetic voice raises no performer consent question, though marking and disclosure duties for AI-generated audio still apply.

Segment at meaning boundaries and keep within readable limits. Netflix's published English Timed Text Style Guide is a widely used reference: a maximum of 42 characters per line, no more than two lines on screen, minimum and maximum event durations, and an adult reading speed of 20 characters per second. Re-check the reading speed for every language, because expansion rates differ.

At the script stage, not after picture lock. Designing gaps for description, avoiding wall-to-wall narration and keeping text clear of the caption area all cost nothing while writing and are expensive afterwards. Everything downstream is cheaper on a video that was written to be described.

Sources and further reading

  1. W3C, Web Content Accessibility Guidelines (WCAG) 2.2
  2. W3C Web Accessibility Initiative — Understanding Conformance
  3. Directive (EU) 2019/882 — European Accessibility Act, EUR-Lex
  4. ETSI EN 301 549 — Accessibility requirements for ICT products and services
  5. Netflix Partner Help Center — English (USA) Timed Text Style Guide

Have an AI or visibility project in mind?

From AI evaluation and human-in-the-loop review to GEO and AEO strategy, our team can help you deploy with confidence and get found in the AI search era.

Talk to our team