Short answer. At minimum, prerecorded video needs captions for all audio content, a transcript or audio alternative, and audio description — the video-specific requirements in WCAG 2.2 at Level A and AA, the level most procurement and legislation reference. In Europe, the European Accessibility Act applies from 28 June 2025 to a defined set of products and services, with EN 301 549 operationalising it through WCAG. AI can draft captions and synthesise description, but accuracy, speaker identification, timing and non-speech information still need a human pass.
Key takeaways
- WCAG 2.2 requires captions (Level A), a text alternative for time-based media (Level A), and audio description (Level A and AA) for prerecorded video; Level AA is the level most legislation and procurement reference.
- Captions and subtitles are not interchangeable: captions serve viewers who cannot hear and carry speaker identification and non-speech sound, while subtitles translate dialogue for viewers who cannot understand the language.
- The European Accessibility Act (Directive (EU) 2019/882) applies from 28 June 2025 to a defined set of products and services, with EN 301 549 as the harmonised standard that incorporates WCAG.
- Automatic speech recognition produces a usable caption draft but still fails on proper nouns, speaker identification, and non-speech audio information without a human correction pass.
- Audio description is far cheaper when gaps for narration are designed into the script, because most commercial video is written wall-to-wall with dialogue.
What does video actually have to provide?
WCAG is the substantive standard behind most legislation and procurement, so satisfying it at the stated level covers most of what the regimes ask for on the content side. WCAG (Web Content Accessibility Guidelines) is the W3C's standard, currently version 2.2, that most accessibility laws and procurement policies point to rather than writing their own criteria.
| Requirement | Level | What it means in production |
|---|---|---|
| Captions for prerecorded audio in synchronised media | A | Accurate, synchronised captions carrying dialogue, speaker identity and meaningful non-speech audio |
| Alternative for time-based media | A | A transcript or equivalent text alternative covering the content of the media |
| Audio description for prerecorded video | A and AA | A narration track describing visual information the existing audio does not convey |
| Captions for live audio | AA | Real-time captioning for live streams — a different production discipline with its own suppliers |
| Contrast and visual presentation | AA | Applies to any text rendered into the picture, including burned-in titles and on-screen graphics |
| Audio control and background audio | A and AAA | Controls for anything that auto-plays; limits on background audio behind speech |
Level AA is the practical target because it is what most legislation and procurement reference. Level A alone rarely satisfies a public-sector or regulated buyer.
Captions are not subtitles. Captions are for viewers who cannot hear the audio: they carry speaker identification and meaningful non-speech sound — a door slamming, a phone ringing, music that carries meaning. Subtitles are for viewers who cannot understand the language: they translate dialogue and assume the viewer hears everything else. A translated subtitle file does not satisfy a caption requirement, and that substitution is one of the most common failures in a library.
What changed in Europe?
The European Accessibility Act sets accessibility requirements for a defined set of products and services and applies from 28 June 2025, moving accessibility from a procurement preference to a market-access condition for covered categories.
The European Accessibility Act (Directive (EU) 2019/882) is an EU directive that makes accessibility a legal market-access condition, rather than a preference, for a defined set of products and services. EN 301 549 is the harmonised European standard for accessibility requirements for ICT products and services, and it incorporates WCAG for web content and documents. In practice, a producer meeting WCAG 2.2 Level AA on its media has done the substantive work the standard asks for on the content side; EN 301 549 also covers software, hardware and documentation outside a content team's scope.
Whether a specific organisation is in scope depends on the product or service and on national implementing legislation, which varies — this is a practitioner's summary, not legal advice, and scope should be confirmed with counsel. The safe operating assumption for a content team is that accessibility is a delivery requirement rather than an enhancement, and that retrofitting a library costs considerably more than producing accessibly from the start.
How do you produce captions that are actually usable?
Automatic speech recognition changed the economics of captioning, not the standard. ASR produces a good first draft on clean single-speaker audio and degrades exactly where accessibility matters most: proper nouns, technical terminology, overlapping speech, accents, and everything that is not speech at all.
The sequence below leaves ASR the mechanical work and people the work that decides whether the result is usable.
1. Transcribe with ASR, then correct against a term list. Feed in the production's pronunciation and terminology list and fix names, products and technical terms first — the highest-error and highest-salience items.
2. Add speaker identification. Who is speaking, wherever it is not obvious from the picture. This is part of what a caption is, and it is absent from every raw ASR output.
3. Add meaningful non-speech audio. Sounds that carry information — a knock, an alarm, laughter, music that signals a change. Describing every ambient sound is as unhelpful as describing none; the test is whether a hearing viewer would take meaning from it.
4. Segment for reading, not for grammar. Break at meaning boundaries so each caption is comprehensible on its own. Netflix's published English Timed Text Style Guide is a usable industry reference: up to 42 characters per line, no more than two lines on screen, and an adult reading speed of 20 characters per second.
5. Time to the audio and respect the minimums. Captions appear with the speech and hold long enough to read — a minimum event duration and a maximum, so nothing flashes and nothing lingers. Timing errors are the fastest way to make technically correct captions unusable.
6. Position around important picture content. Move captions off burned-in titles, faces and lower thirds. That is a placement decision automatic tools do not make.
7. Deliver as a sidecar file. SRT, WebVTT or TTML rather than burned in, wherever the platform supports it — editable, indexable and switchable by the viewer.
8. Review in context. Watch the video with the captions on. Errors invisible in the caption file — a caption covering the thing being discussed, a line landing after the cut — are obvious within thirty seconds of playback.
What does audio description actually require?
Audio description is a narration track conveying visual information the existing audio does not: who is on screen, what they are doing, what is written on screen, what changed. It is a WCAG requirement for prerecorded video, and it is the one most often quietly skipped, because it is the one that costs real production effort.
The difficulty is structural rather than technical. Description has to fit in the gaps between dialogue, and most commercial video is written wall-to-wall with narration. A video with no gaps cannot be described without either an extended-description version that pauses the picture, or a re-edit.
- Design for description at the script stage. Leaving deliberate gaps costs nothing while writing and is expensive to create afterwards. This is the highest-leverage decision in the whole area.
- Describe what matters, in priority order. Actions and on-screen text first, then setting, then detail. Description competes for time with the programme audio, and there is never enough of it.
- Read on-screen text aloud. Titles, captions, statistics and any text rendered into the picture. This is the most commonly missed content and often the most information-dense.
- Synthesis is appropriate here. Described audio is functional narration, which is exactly the register where synthetic voice performs well, and it is what makes description affordable across many languages — provided the content labelling and disclosure rules that apply to synthetic audio are followed.
- Deliver as a separate track or version, per platform capability, so viewers can choose it.
How does this work across many languages?
Accessibility and localisation are the same production problem viewed from two angles, and treating them as one workflow is what makes both affordable. Both operate on a locked picture, both produce sidecar text files, both need in-market human review, and both draw on the same terminology asset.
- One timed-text source, many derivations. The caption file is the base artefact; subtitles in other languages derive from it, and so does the transcript. Producing them independently duplicates the timing work, which is the expensive part — the same principle behind localizing one video into many languages.
- Reuse the pronunciation lexicon across synthetic described-audio tracks in every language, exactly as for multilingual AI voice production, and draw it from a licensed voice library rather than an unlicensed one.
- Check reading speed per language. The same sentence expands differently across languages, so caption timing that is comfortable in one is unreadable in another. That is a script-fitting decision, not a subtitling one.
- Have an in-market reviewer watch it. Register, terminology currency and cultural readability are not visible in a caption file, and they are the difference between conformant and usable.
Reading speed = Characters in the caption ÷ Seconds the caption is on screen
Twenty characters per second is the adult ceiling in the Netflix guide for English. Recompute it per language rather than carrying the English figure across: a file that passes the character-per-line rule can still run too fast to read.
How does Lifewood handle video accessibility at scale?
Lifewood produces captions, subtitles, transcripts and described audio inside the same localisation pipeline rather than as separate services, so the timing work is done once and derived across every language. Each language variant gets region-native review, because register and terminology currency are not visible in a caption file.
Coverage runs to 100+ languages across 40+ delivery centres in 30+ countries, under a 95%+ accuracy SLA with human review on every variant — checked through the same quality-control process applied across AI-generated content, which matters most on the ASR correction pass, where errors cluster on exactly the names a viewer relying on captions most needs to be right. This sits within Lifewood's broader AIGC video production and AIGC services work.