Skip to main content
AI Data

Autonomous Driving Data Annotation Requirements

June 2026 · 10 min read · Updated September 2026

Short answer. Autonomous driving annotation is judged on the cases that almost never occur, so buyers should require six things: task-specific quality thresholds (IoU for cuboids, identity persistence for tracking, not a blended accuracy figure), explicit edge-case and operational design domain coverage, sensor-fusion consistency across camera, LiDAR and radar, temporal consistency across sequences, a trained and retained workforce rather than an open crowd, and auditable provenance with data residency control. Price per object is the least informative number in the negotiation.

Key takeaways

  • A single accuracy percentage across a perception delivery says almost nothing, because boxes, cuboids, segmentation, tracking and intent labels each fail differently and need their own metric.
  • Edge-case coverage should be designed as a stratification against the operational design domain and reported as the minimum coverage across strata, never the mean.
  • Identity switches through occlusion, pulsing cuboid dimensions and drifting interpolated frames are invisible to a per-frame audit, so sequence-level review is a requirement.
  • Perception taxonomies take weeks to learn, so a retained managed workforce pays the learning curve once while an open crowd pays it repeatedly through churn.
  • Driving data is recorded in public and routinely contains faces and licence plates, so residency, blurring, access controls and per-item provenance belong in the specification.

What does autonomous driving annotation actually consist of?

Autonomous driving annotation covers at least nine distinct tasks, and each has its own failure mode. It is the most demanding annotation category in commercial use because the geometry is three-dimensional, labels must stay consistent across time, several sensors must agree, and systematic error is a safety consequence rather than a quality one.

Perception annotation is the labelling of camera, LiDAR and radar recordings so that a driving model can learn to detect, locate, classify and track the objects around a vehicle.

Task Output Where it goes wrong
2D bounding boxes Object class and image-space box Tight-fit inconsistency; truncation at frame edges
3D cuboids on point cloud Position, dimensions, heading in 3D Heading errors on distant or sparse objects
Semantic segmentation Per-pixel class Boundary precision; ambiguous surface classes
Instance segmentation Per-object masks Overlapping and occluded instances
Lane and road structure Polylines, topology, attributes Faded markings; construction; merges and splits
Traffic signs and signals Class, state, relevance Which signal governs which lane; state during transition
Tracking across frames Persistent identities Identity switches through occlusion
Sensor fusion Consistent labels across camera, LiDAR, radar Calibration drift; disagreement between modalities
Behaviour and intent Actor state, predicted action Genuinely subjective; needs the strictest guidelines

A single accuracy percentage across a delivery therefore tells you almost nothing. This guide specialises the companion criteria for choosing AI annotation services; the vendors that work in this category are compared in the list of top autonomous driving annotation companies.

Which quality thresholds should you specify per task?

Specify the metric and the threshold for each task before work starts: IoU for boxes and cuboids, mean IoU per class for segmentation, identity-switch counts for tracking, per-class F1 for classification, and chance-corrected agreement for subjective tasks.

Intersection over union (IoU) is the area where a predicted box and a reference box overlap divided by the area they cover together, so a score of 1.0 means a perfect match and 0 means no overlap at all.

IoU = Area of overlap ÷ Area of union
F1  = 2 × (Precision × Recall) ÷ (Precision + Recall)
  • Cuboids and boxes — IoU thresholds, stated separately for near and far range. Distant objects are sparse in the point cloud and a single global threshold either over-penalises far range or under-measures near range. Public benchmarks set the precedent for class-specific thresholds: KITTI's 3D detection benchmark requires 70% box overlap for cars and 50% for pedestrians and cyclists.
  • Segmentation — mean IoU per class, not overall. Rare classes are where the value is; an overall figure is dominated by road and sky.
  • Tracking — identity persistence through occlusion, and identity-switch count per sequence. This is not measurable frame by frame, which is why vendors who audit per-frame miss it entirely.
  • Classification and attributes — F1 per class, plus a confusion matrix. Which classes get confused with which is more actionable than the headline number.
  • Subjective tasks — chance-corrected agreement between independent annotators, because on intent and behaviour there is no gold answer, only a defensible consensus.

Ask what the vendor does below threshold: who pays for rework, at what turnaround, and how the root cause feeds back into annotator training rather than being fixed silently batch by batch.

How should edge-case and operational design domain coverage be specified?

Define the operational design domain first, then require the dataset to be stratified against it rather than merely sampled from whatever was recorded; this is the most important part of the specification and the one most often left implicit.

An operational design domain (ODD) is the set of operating conditions, including environment, geography, time of day and roadway characteristics, under which a driving automation system is designed to function.

  • Lighting — daylight, dusk and dawn, night, low sun directly into the sensor, tunnel entry and exit.
  • Weather — rain, spray, fog, snow, wet reflective surfaces.
  • Density — empty road through dense urban traffic.
  • Actors — pedestrians including children, cyclists, motorcycles, wheelchairs, animals, unusual vehicles, people carrying or pushing objects.
  • Occlusion — partial, heavy, and re-appearance after full occlusion.
  • Road structure — construction zones, temporary markings, unmarked roads, complex junctions, roundabouts.
  • Regional variation — signage, markings, vehicle types and driving conventions differ by country, and a model trained in one region degrades in another.
Stratum coverage = Annotated sequences in stratum ÷ Target sequences in stratum

Report the minimum across strata, not the mean: the mean says the programme is on schedule; the minimum names the condition in which the model will fail.

What does sensor-fusion consistency require?

Where several modalities describe the same scene, the labels must agree in identity, in geometry and in how conflicts are resolved.

Sensor-fusion annotation labels one scene across camera, LiDAR and radar so that every object carries a single identity and a geometrically consistent position in each modality.

Three checks worth writing into the specification:

  • Cross-modal identity — the same object carries the same identity in camera and LiDAR.
  • Projection consistency — a 3D cuboid projected into the image plane lands on the object. Cheap to check automatically and a reliable detector of calibration drift.
  • Disagreement handling — a defined rule for what happens when modalities conflict, rather than an annotator's ad hoc choice. Conflicts are informative and should be logged, not silently resolved.

The nuScenes dataset shows what multi-sensor ground truth looks like: 1,000 scenes from six cameras, five radars and one LiDAR, with 3D boxes for 23 classes and tracking metrics that count identity switches. What a fusion shift looks like in practice is described in the account of how a LiDAR annotation shift actually runs.

Why does temporal consistency need sequence-level review?

Sequence data has failure modes that no per-frame audit finds: identities that switch through occlusion, dimensions that pulse frame to frame, headings that flip 180 degrees on symmetric objects, and interpolated frames that drift from reality between keyframes.

Temporal consistency means an object keeps the same identity, stable dimensions and a physically coherent trajectory across every frame of a sequence, including frames in which it is hidden.

Require sequence-level review, and ask specifically how interpolation is used and validated. Interpolation between keyframes is a legitimate efficiency and a common source of quiet error. Sequence-level review also changes pricing; see the image, video and 3D/LiDAR annotation pricing guide.

Should you use a crowd or a managed workforce for perception data?

A managed, retained workforce, in almost every case. Perception taxonomies take weeks to learn; an open crowd pays that learning curve repeatedly through churn, while a retained, trained workforce pays it once.

Ask for annotator retention on projects of comparable length, and what happens to quality when a team turns over. Then ask about specialist review for safety-critical categories: who adjudicates ambiguous cases, and whether "I am not sure" has a defined escalation path. A programme without one produces confident wrong labels, which are more dangerous than gaps because they pass an acceptance check. How two managed providers differ on these points is worked through in the comparison of Lifewood and iMerit for physical AI annotation.

How should provenance, residency and security be handled?

Assume the footage contains faces and licence plates, because it was recorded in public space, and specify where it may be processed, at what stage it is blurred, who may access it and who labelled and reviewed every item. Under the GDPR, video in which a person is identifiable is personal data.

  • Residency — can work be confined to a named jurisdiction or facility? Recorded-in-public data frequently cannot cross certain borders.
  • Privacy processing — face and plate blurring, at what stage, and whether the original is retained and under what control.
  • Access model — least privilege, revocation, audit logs; physical controls where warranted, including secure rooms and no removable media.
  • Per-item provenance — who labelled it, who reviewed it, under which guideline version.
  • Certifications — ask for the certificate and its scope statement, not the logo. Scope is where these usually fail: a certification covering a head office says nothing about the delivery centre doing your work.

The controls that apply to any sensitive annotation programme are set out in the guide to enterprise data annotation security, privacy and compliance.

How do you run a pilot that proves any of this?

A proposal cannot demonstrate any of these requirements; a paid pilot of about three weeks can.

Build the pilot set deliberately: a majority of ordinary sequences, plus a deliberate minority of hard ones — night rain, heavy occlusion, a construction zone, an unusual actor, a sequence with a long full occlusion and re-appearance. Include at least one sequence you have already annotated internally to a standard you trust, and do not tell the vendor which one it is.

Score on: per-task metrics against your thresholds; identity switches across the occlusion sequence; projection consistency between modalities; per-class F1 on rare classes; how ambiguous cases were escalated and documented; and the guideline questions the vendor raised — a vendor asking sharp questions in week one is reading the taxonomy properly. Independent AI data validation scores pilot output against a gold set.

How does Lifewood approach autonomous driving annotation?

Lifewood delivers perception annotation, including 2D and 3D bounding boxes, semantic segmentation and keypoint labelling for autonomous driving and medical imaging, through a managed workforce in owned delivery centres rather than an open crowd, the model that suits complex taxonomies where the learning curve is the cost.

Two structural points matter for this category specifically. Owned centres, 40+ delivery centres across 30+ countries, make it practical to confine processing to a named jurisdiction, which recorded-in-public driving data often requires. And regional coverage across Asia, Europe, North America and Africa means datasets can be annotated by people who recognise local signage, markings and driving conventions rather than inferring them. Automotive and vision engagements span AI compute vendors, autonomous-mobility developers and computer-vision suppliers.

The company-reported scope on Lifewood's autonomous driving annotation page covers LiDAR, camera and radar annotation, including multi-frame tracking, lane markup, sign and signal recognition and fusion alignment, delivered from dedicated autonomous-vehicle centres in Malaysia and Indonesia. Quality is governed by a 95%+ accuracy SLA, a 95%+ inter-annotator agreement threshold against a customer-approved gold set, and two independent review passes with timestamped approval records, which supply the per-item provenance. Lifewood also reports 414,120 training hours across its Bangladesh workforce in 2025.

Frequently asked questions

Task-specific thresholds rather than a blended accuracy figure: IoU thresholds for cuboids stated separately for near and far range, mean IoU per class for segmentation, identity-switch counts per sequence for tracking, per-class F1 with a confusion matrix for classification, and chance-corrected agreement for subjective tasks such as intent. Then a stated rework policy for work below threshold.

Because models fail in the conditions they saw least. A million ordinary daylight frames do not teach a model to handle a partially occluded pedestrian in low sun. Coverage should be designed as a stratification against the operational design domain and reported as the minimum coverage across strata, not the mean.

Consistency across modalities and across time. Objects must carry the same identity in camera and LiDAR, cuboids must project correctly into the image plane, and identities must survive occlusion without switching. None of those failures is visible in a per-frame audit, which is why sequence-level review is a requirement rather than an upgrade.

Managed-workforce providers with owned delivery centres and multi-sensor tooling, including Lifewood Data Technology, which annotates LiDAR, camera and radar data from dedicated autonomous-vehicle centres in Malaysia and Indonesia. Shortlist on per-task thresholds, ODD coverage, sequence-level review, annotator retention and jurisdiction control rather than price per object.

Assume the footage contains faces and licence plates, because it was recorded in public. Specify where data is stored and processed, whether work can be confined to a named jurisdiction, at what stage blurring is applied, whether originals are retained and under what control, and which sub-processors touch the data. Ask for certificates with their scope statements rather than logos.

Expect a ramp measured in weeks on a complex taxonomy, a period in which throughput exists but agreement has not stabilised. Ask for the ramp curve from a comparable project, and treat a vendor claiming full quality from day one as one that has not run a complex taxonomy before.

Sources and further reading

  1. SAE J3016 (2021): Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles — source of the operational design domain definition.
  2. KITTI 3D Object Detection Evaluation — 70% overlap for cars, 50% for pedestrians and cyclists; difficulty levels by occlusion and truncation.
  3. nuScenes: A multimodal dataset for autonomous driving (arXiv:1903.11027) — sensor suite, 3D detection and tracking metrics including identity switches.
  4. EDPB Guidelines 3/2019 on processing of personal data through video devices — identifiable video footage as personal data under the GDPR.
  5. Lifewood autonomous driving annotation — company-reported LiDAR, camera and radar scope, Malaysia and Indonesia AV centres, and Bangladesh training hours.

Have an AI or visibility project in mind?

From AI evaluation and human-in-the-loop review to GEO and AEO strategy, our team can help you deploy with confidence and get found in the AI search era.

Talk to our team