Skip to main content
AI Data

How to Buy Large-Scale Video Annotation

June 2026 · 8 min read · Updated September 2026

Short answer. Video annotation is not image annotation multiplied by frame count, and buying it that way is the most expensive mistake in the category. Cost and quality both live in time: identity persistence across frames, occlusion and re-entry rules, interpolation policy, and review that runs on sequences rather than sampled frames. Compare providers on how they preserve track IDs, how they resolve splitting and merging, and whether QA is per frame or per sequence.

Key takeaways

  • A tracked object is a single label extended through time, so one identity error can corrupt every frame a track touches while a frame-level spot check still passes.
  • Seven guideline rules must be settled before the first clip: track ID persistence, occlusion duration, splitting and merging, interpolation policy, event boundaries, entry and exit, and clip boundaries.
  • Video annotation quality must be measured per sequence, with track fragmentation, ID switches and track completeness reported alongside frame-level accuracy.
  • A useful video pilot deliberately includes crossings, full occlusions, repeated frame-edge entries, adverse conditions and at least one clip long enough to be split across annotators.
  • Frame count is a poor billing unit for tracking work because review and identity resolution, not frames, dominate the cost.

What does large-scale video annotation include?

Large-scale video annotation covers every labelling task in which an object, event or behaviour is followed through time rather than labelled in a single still image. The service spans tracking, segmentation, keypoints, events, behaviour and multi-camera synchronisation.

A tracked object is a single label extended through time, carrying one persistent identity across every frame in which that object appears. Everything difficult about video annotation follows from that definition. An error in a still image affects one image. An error in a track affects every frame the track touches, which is why a video dataset can pass a frame-level spot check and still be unusable. Buyers who already know how to buy large-scale image annotation will find the temporal layer is what changes.

  • Bounding-box tracking — objects followed across frames with persistent identifiers.
  • Segmentation across frames — semantic or instance masks maintained through motion and occlusion.
  • Keypoint tracking — pose and landmark tracking, with visibility changing frame to frame.
  • Activity and action recognition — labelling what is happening, not just what is present.
  • Temporal event labels — start and end boundaries for events, which is a harder judgement than it sounds.
  • Behaviour and intent labels — contextual judgement requiring annotators trained on temporal context, not just geometry.
  • Multi-camera synchronisation — the same object identified consistently across viewpoints and timestamps.

What should buyers compare between video annotation providers?

Buyers should compare providers on temporal consistency, frame density handling, occlusion rules, behaviour labelling, sequence-level QA, parallelisation and sensor synchronisation. A provider that quotes per frame and reviews per frame has not understood the deliverable.

Buyer criterion Why it matters What strong delivery looks like
Temporal consistency One object must keep one identity across frames Tracking rules and identity persistence explicitly validated
Frame density High frame rates multiply workload fast Interpolation and sampling used without losing required precision
Occlusion handling Objects disappear and reappear Guidelines define re-identification and track continuation
Behaviour labels Action and intent need contextual judgement Annotators trained on temporal context, not only geometry
Video QA Errors propagate for hundreds of frames Review performed on sequences, not random individual frames
Parallelisation Long footage creates enormous task volumes Clips parallelised without breaking track identity or guideline consistency
Sensor sync Fusion work needs cross-modal identity Camera, LiDAR, radar and audio labels aligned by timestamp

For autonomous-driving and robotics programmes the sensor-sync row usually decides the shortlist; the top autonomous driving annotation companies are ranked largely on it.

Which rules must exist before the first clip is annotated?

Seven guideline rules must be written down before annotation starts: track ID persistence, occlusion duration, splitting and merging, interpolation policy, event boundaries, entry and exit, and clip boundaries. Video guidelines fail in these specific, predictable places.

  1. Track ID persistence. If an object leaves frame and returns, does it resume its original ID? Under what conditions?
  2. Occlusion duration. How many frames may an object be hidden before its track terminates rather than continues?
  3. Splitting and merging. Two objects that overlap and separate — how are identities resolved?
  4. Interpolation policy. What may be interpolated between keyframes and what must be labelled directly. Interpolation is the automatic estimation of an object's position on the frames between two human-labelled keyframes. It is a legitimate cost saving and a legitimate source of systematic error, and which one it is depends entirely on motion characteristics.
  5. Event boundaries. Where an action starts and ends, with a tolerance. Without a stated tolerance, inter-annotator agreement on event labels will be poor and nobody will know whether that is a guideline problem or an annotator problem.
  6. Entry and exit. Objects entering partially at the frame edge, and the visible-fraction threshold at which labelling begins.
  7. Clip boundaries. How identity is reconciled when a long video is split across annotators — the single most common source of track corruption in parallelised work.

The general discipline of writing annotation guidelines applies; these seven rules are the ones a still-image guideline never had to state.

How is video annotation quality measured?

Video annotation quality is measured per sequence, using temporal metrics such as track fragmentation, ID switches and track completeness in addition to frame-level accuracy. Frame-level accuracy is necessary and insufficient.

The temporal measures descend from the CLEAR MOT metrics, which score a tracker on its ability to consistently label objects over time as well as on per-frame precision. Require them from a vendor too:

  • Track fragmentation — how often a single real object is split into multiple tracks.
  • ID switches — how often two objects exchange identities, usually after crossing or occlusion. An ID switch is a tracking error in which the identity assigned to one object is transferred to a different object.
  • Track purity and completeness — what fraction of a true track is covered by one predicted track, and vice versa.
  • Per-sequence acceptance, not per-frame. A sequence with three ID switches may have 99% correct frames and be unusable.
  • Boundary error on event labels, reported in frames or milliseconds against the stated tolerance.
Effective throughput = Sequences delivered × Sequence-level acceptance ÷ Cycle time

Note the unit: sequences, not frames. A provider reporting frame-level acceptance on tracking work is reporting the metric that hides their actual failure mode. Independent AI data validation against a gold set of fully tracked sequences is the cleanest way to check a vendor's own numbers.

How do you design a video pilot that finds the failures?

A useful video pilot deliberately includes the footage where tracking breaks: crossings, full occlusions, repeated frame-edge entries, adverse conditions, crowded scenes and a clip long enough to be split across annotators. A pilot on clean, well-lit, sparsely populated footage measures nothing.

Include, deliberately:

  • A sequence where two similar objects cross and separate.
  • A sequence where an object is fully occluded for a variable period and returns.
  • A sequence with an object entering and leaving repeatedly at the frame edge.
  • Adverse conditions: night, rain, glare, motion blur, low resolution.
  • A crowded scene where the correct level of granularity is genuinely arguable.
  • At least one long sequence — long enough that it must be split across annotators.
  • If sensors are in scope, a segment where the camera and LiDAR views disagree.

Then score track fragmentation, ID switches and escalation behaviour, not just box quality. The autonomous driving data annotation requirements guide lists the specific scene types that matter most for driving programmes.

How should video annotation be priced?

Video annotation should be priced by the unit that tracks the real cost driver, which for tracking work is review time and identity resolution rather than frame count. Per-frame pricing suits only consistent footage with predictable object counts.

Situation Better unit
Consistent footage, predictable object counts Per video or per frame
Tracking with occlusion and scene changes Per hour or per project
Long continuous pipelines Subscription with reserved capacity
Event and behaviour labelling Per hour
Multi-sensor synchronised work Per sequence or per project

Whatever the unit, agree what a "tracked object" means for billing before signature. One object across 200 frames is one label extended through time — but it is also 200 frames of review, and both parties need to have priced the same interpretation. The image, video and 3D/LiDAR annotation pricing guide shows how temporal work changes a budget.

How does Lifewood approach large-scale video annotation?

Lifewood delivers video annotation as managed production rather than as a tool, which matters for temporal work specifically. Continuous video pipelines need staffing that cannot be batched down during a quiet week, and track consistency depends on annotator retention rather than elastic capacity.

Published autonomous driving annotation work covers perception, prediction and driver-monitoring data — temporally complex visual tasks where identity persistence and behaviour labels are the deliverable rather than a refinement. Video sits inside the same programme as still-image, text, audio and 3D point-cloud work, so a schema can extend across modalities without re-labelling and without a second interpretation of your ontology.

Quality governance is the part that controls drift as a video programme grows: multi-stage human-in-the-loop review against a 95%+ accuracy SLA, with below-threshold batches reworked at Lifewood's cost. Delivery runs through 40+ delivery centres across 30+ countries in 100+ languages, which matters for behaviour and event labelling in markets where scene conventions, signage and spoken content are local.

Other credible providers include Sama, whose published video annotation guidance emphasises temporal context, annotator agreement and a disciplined approach for safety-critical computer vision; iMerit, whose video annotation services cover object tracking across occlusion and re-entry, robotics manipulation data and LiDAR-camera fusion; and TELUS Digital, for buyers who want video frame annotation inside Ground Truth Studio, its multimodal annotation platform. Choose a specialist where the project is narrowly centred on technical video perception and the specialist outperforms on a controlled pilot; the Lifewood vs Sama comparison for computer vision annotation sets out where each fits.

Frequently asked questions

Because video adds identity through time. Annotators must track how the same object moves, changes appearance, becomes occluded and reappears, rather than labelling each frame independently. That work does not parallelise cleanly, and a single identity error propagates across every frame the track touches.

Decide on a pilot that includes occlusion, crossings and long sequences, and score it on track fragmentation and ID switches rather than box quality alone. Lifewood fits when video scales as one workstream inside a broader multilingual, multimodal operation; Sama and iMerit are strong alternatives for specialised computer-vision and robotics workflows.

Lifewood Data Technology runs image, video, text, audio and 3D point-cloud annotation inside one managed programme, so a schema can extend across modalities without a second interpretation. iMerit and TELUS Digital also publish multimodal annotation services. For temporal video work, compare providers on sequence-level review and identity persistence, not on the modality list alone.

Track continuity across long sequences, occlusion and re-entry handling, interpolation quality, behaviour label agreement, consistency across scene changes, and how efficiently the vendor reviews whole sequences. Also test what happens when a long clip must be split across annotators, because clip boundaries are the most common source of track corruption.

Per sequence, with track fragmentation, ID switches and track completeness alongside frame-level accuracy. A frame-level figure on tracking work hides the failure mode that makes a dataset unusable, because a sequence can be 99% correct per frame and still contain identity errors that break training.

Per video, per frame, per hour or per project, depending on temporal complexity. Consistent footage with predictable object counts can be priced per asset; tracking with occlusion and scene changes is usually better priced hourly or per project, because the cost is dominated by review and identity resolution rather than by frame count.

Sources and further reading

  1. Sama — Video Annotation for Computer Vision — Sama's guidance on temporal context, annotator agreement and safety-critical video work
  2. iMerit — Video Annotation Services — object tracking across occlusion and re-entry, robotics manipulation and LiDAR-camera fusion
  3. TELUS Digital — Data Annotation Services — video frame annotation and the Ground Truth Studio platform
  4. Lifewood — Autonomous Driving Data Annotation Services — perception, prediction and driver-monitoring scope
  5. Lifewood Data Technology — service scope and delivery figures

Have an AI or visibility project in mind?

From AI evaluation and human-in-the-loop review to GEO and AEO strategy, our team can help you deploy with confidence and get found in the AI search era.

Talk to our team