Short answer. At minimum: captions for all prerecorded audio content, a transcript or audio alternative, and audio description for prerecorded video. Those are the video-specific requirements in WCAG 2.2 at Level A and AA, and Level AA is what most procurement and most legislation reference. In Europe they have teeth: the European Accessibility Act (Directive (EU) 2019/882) applies from 28 June 2025 to a defined set of products and services, with EN 301 549 as the harmonised standard that operationalises it by adopting WCAG. AI genuinely helps — speech recognition produces a caption draft, synthesis produces a described-audio track — but neither is a deliverable unaccompanied, because accuracy, speaker identification, non-speech information and timing all still need a human pass.
Accessibility work fails in predictable places, and almost none of them are technical. It fails because a translated subtitle file was filed as a caption track, because audio description was dropped, or because word-perfect captions run too fast to read.
This is a practitioner's summary of the applicable standards, not legal advice — the scope of the European Accessibility Act in particular depends on the specific product or service and on national implementing law, and should be confirmed with counsel.
What does video actually have to provide?
WCAG — the Web Content Accessibility Guidelines, published by the W3C, currently at version 2.2 — is the substantive standard. Most legislation and most procurement point at it rather than writing their own criteria, which is convenient: satisfy WCAG at the stated level and you have satisfied most of what the regimes ask for on the content side.
| Requirement | Level | What it means in production |
|---|---|---|
| Captions for prerecorded audio in synchronised media | A | Accurate, synchronised captions carrying dialogue, speaker identity and meaningful non-speech audio |
| Alternative for time-based media | A | A transcript or equivalent text alternative covering the content of the media |
| Audio description for prerecorded video | A and AA | A narration track describing visual information the existing audio does not convey |
| Captions for live audio | AA | Real-time captioning for live streams — a different production discipline with its own suppliers |
| Contrast and visual presentation | AA | Applies to any text rendered into the picture, including burned-in titles and on-screen graphics |
| Audio control and background audio | A and AAA | Controls for anything that auto-plays; limits on background audio behind speech |
Level AA is the practical target because it is what most legislation and procurement reference. Level A alone rarely satisfies a public-sector or regulated buyer.
Captions are not subtitles. Captions are for viewers who cannot hear the audio: they carry speaker identification and meaningful non-speech sound — a door slamming, a phone ringing, music that carries meaning. Subtitles are for viewers who cannot understand the language: they translate dialogue and assume the viewer hears everything else. A translated subtitle file does not satisfy a caption requirement, and that substitution is one of the most common failures in a library.
What changed in Europe?
The European Accessibility Act — Directive (EU) 2019/882 — sets accessibility requirements for a defined set of products and services and applies from 28 June 2025. Its significance for content producers is less that it invented technical requirements than that it moved accessibility from a procurement preference to a market-access condition for covered categories.
EN 301 549 is the harmonised European standard for accessibility requirements for ICT products and services, and it incorporates WCAG for web content and documents. In practice, a producer meeting WCAG 2.2 Level AA on its media has done the substantive work the standard asks for on the content side; EN 301 549 also covers software, hardware and documentation outside a content team's scope.
Whether a specific organisation is in scope depends on the product or service and on national implementing legislation, which varies. The safe operating assumption for a content team is that accessibility is a delivery requirement rather than an enhancement, and that retrofitting a library costs considerably more than producing accessibly from the start.
How do you produce captions that are actually usable?
Automatic speech recognition changed the economics of captioning, not the standard. ASR produces a good first draft on clean single-speaker audio and degrades exactly where accessibility matters most: proper nouns, technical terminology, overlapping speech, accents, and everything that is not speech at all.
The sequence below leaves ASR the mechanical work and people the work that decides whether the result is usable.
1. Transcribe with ASR, then correct against a term list. Feed in the production's pronunciation and terminology list and fix names, products and technical terms first — the highest-error and highest-salience items.
2. Add speaker identification. Who is speaking, wherever it is not obvious from the picture. This is part of what a caption is, and it is absent from every raw ASR output.
3. Add meaningful non-speech audio. Sounds that carry information — a knock, an alarm, laughter, music that signals a change. Describing every ambient sound is as unhelpful as describing none; the test is whether a hearing viewer would take meaning from it.
4. Segment for reading, not for grammar. Break at meaning boundaries so each caption is comprehensible on its own. Netflix's published English Timed Text Style Guide is a usable industry reference: up to 42 characters per line, no more than two lines on screen, and an adult reading speed of 20 characters per second.
5. Time to the audio and respect the minimums. Captions appear with the speech and hold long enough to read — a minimum event duration and a maximum, so nothing flashes and nothing lingers. Timing errors are the fastest way to make technically correct captions unusable.
6. Position around important picture content. Move captions off burned-in titles, faces and lower thirds. That is a placement decision automatic tools do not make.
7. Deliver as a sidecar file. SRT, WebVTT or TTML rather than burned in, wherever the platform supports it — editable, indexable and switchable by the viewer.
8. Review in context. Watch the video with the captions on. Errors invisible in the caption file — a caption covering the thing being discussed, a line landing after the cut — are obvious within thirty seconds of playback.
Audio description: the requirement everyone underestimates
Audio description is a narration track conveying visual information the existing audio does not: who is on screen, what they are doing, what is written on screen, what changed. It is a WCAG requirement for prerecorded video, and it is the one most often quietly skipped, because it is the one that costs real production effort.
The difficulty is structural rather than technical. Description has to fit in the gaps between dialogue, and most commercial video is written wall-to-wall with narration. A video with no gaps cannot be described without either an extended-description version that pauses the picture, or a re-edit.
- Design for description at the script stage. Leaving deliberate gaps costs nothing while writing and is expensive to create afterwards. This is the highest-leverage decision in the whole area.
- Describe what matters, in priority order. Actions and on-screen text first, then setting, then detail. Description competes for time with the programme audio, and there is never enough of it.
- Read on-screen text aloud. Titles, captions, statistics and any text rendered into the picture. This is the most commonly missed content and often the most information-dense.
- Synthesis is appropriate here. Described audio is functional narration, which is exactly the register where synthetic voice performs well — and it is what makes description affordable across many languages.
- Deliver as a separate track or version, per platform capability, so viewers can choose it.
How does this work across many languages?
Accessibility and localisation are the same production problem viewed from two angles, and treating them as one workflow is what makes both affordable. Both operate on a locked picture, both produce sidecar text files, both need in-market human review, and both draw on the same terminology asset.
- One timed-text source, many derivations. The caption file is the base artefact; subtitles in other languages derive from it, and so does the transcript. Producing them independently duplicates the timing work, which is the expensive part.
- Reuse the pronunciation lexicon across synthetic described-audio tracks in every language, exactly as for voiceover — see multilingual AI voice production.
- Check reading speed per language. The same sentence expands differently across languages, so caption timing that is comfortable in one is unreadable in another. That is a script-fitting decision, not a subtitling one.
- Have an in-market reviewer watch it. Register, terminology currency and cultural readability are not visible in a caption file, and they are the difference between conformant and usable.
Reading speed = Characters in the caption ÷ Seconds the caption is on screen
Twenty characters per second is the adult ceiling in the Netflix guide for English. Recompute it per language rather than carrying the English figure across: a file that passes the character-per-line rule can still run too fast to read.
How Lifewood approaches this
Lifewood produces captions, subtitles, transcripts and described audio inside the same localisation pipeline rather than as separate services, so the timing work is done once and derived across every language. Each language variant gets region-native review, because register and terminology currency are not visible in a caption file.
Coverage runs to 50+ languages across 40+ delivery centres in 30+ countries, under a 95%+ accuracy threshold with human review on every variant — which matters most on the ASR correction pass, where errors cluster on exactly the names a viewer relying on captions most needs to be right.
See AIGC video production, AIGC services, multilingual data collection and the QA process.
Sources and further reading
- W3C, Web Content Accessibility Guidelines (WCAG) 2.2, October 2023, and the W3C Web Accessibility Initiative overview of conformance levels.
- Directive (EU) 2019/882 — European Accessibility Act, EUR-Lex, Official Journal of the European Union, 2019; applies from 28 June 2025.
- ETSI / CEN / CENELEC, EN 301 549 — Accessibility requirements for ICT products and services (harmonised standard).
- Netflix Partner Help Center, English Timed Text Style Guide — line length, reading speed and duration limits.

