Short answer. Autonomous driving annotation is judged on the cases that almost never occur, so buyers should require six things: task-specific quality thresholds (IoU for cuboids, identity persistence for tracking, not a blended accuracy figure), explicit edge-case and operational design domain coverage, sensor-fusion consistency across camera, LiDAR and radar, temporal consistency across sequences, a trained and retained workforce rather than an open crowd, and auditable provenance with data residency control. Price per object is the least informative number in the negotiation.
Key takeaways
- A single accuracy percentage across a perception delivery says almost nothing, because boxes, cuboids, segmentation, tracking and intent labels each fail differently and need their own metric.
- Edge-case coverage should be designed as a stratification against the operational design domain and reported as the minimum coverage across strata, never the mean.
- Identity switches through occlusion, pulsing cuboid dimensions and drifting interpolated frames are invisible to a per-frame audit, so sequence-level review is a requirement.
- Perception taxonomies take weeks to learn, so a retained managed workforce pays the learning curve once while an open crowd pays it repeatedly through churn.
- Driving data is recorded in public and routinely contains faces and licence plates, so residency, blurring, access controls and per-item provenance belong in the specification.
What does autonomous driving annotation actually consist of?
Autonomous driving annotation covers at least nine distinct tasks, and each has its own failure mode. It is the most demanding annotation category in commercial use because the geometry is three-dimensional, labels must stay consistent across time, several sensors must agree, and systematic error is a safety consequence rather than a quality one.
Perception annotation is the labelling of camera, LiDAR and radar recordings so that a driving model can learn to detect, locate, classify and track the objects around a vehicle.
| Task | Output | Where it goes wrong |
|---|---|---|
| 2D bounding boxes | Object class and image-space box | Tight-fit inconsistency; truncation at frame edges |
| 3D cuboids on point cloud | Position, dimensions, heading in 3D | Heading errors on distant or sparse objects |
| Semantic segmentation | Per-pixel class | Boundary precision; ambiguous surface classes |
| Instance segmentation | Per-object masks | Overlapping and occluded instances |
| Lane and road structure | Polylines, topology, attributes | Faded markings; construction; merges and splits |
| Traffic signs and signals | Class, state, relevance | Which signal governs which lane; state during transition |
| Tracking across frames | Persistent identities | Identity switches through occlusion |
| Sensor fusion | Consistent labels across camera, LiDAR, radar | Calibration drift; disagreement between modalities |
| Behaviour and intent | Actor state, predicted action | Genuinely subjective; needs the strictest guidelines |
A single accuracy percentage across a delivery therefore tells you almost nothing. This guide specialises the companion criteria for choosing AI annotation services; the vendors that work in this category are compared in the list of top autonomous driving annotation companies.
Which quality thresholds should you specify per task?
Specify the metric and the threshold for each task before work starts: IoU for boxes and cuboids, mean IoU per class for segmentation, identity-switch counts for tracking, per-class F1 for classification, and chance-corrected agreement for subjective tasks.
Intersection over union (IoU) is the area where a predicted box and a reference box overlap divided by the area they cover together, so a score of 1.0 means a perfect match and 0 means no overlap at all.
IoU = Area of overlap ÷ Area of union
F1 = 2 × (Precision × Recall) ÷ (Precision + Recall)
- Cuboids and boxes — IoU thresholds, stated separately for near and far range. Distant objects are sparse in the point cloud and a single global threshold either over-penalises far range or under-measures near range. Public benchmarks set the precedent for class-specific thresholds: KITTI's 3D detection benchmark requires 70% box overlap for cars and 50% for pedestrians and cyclists.
- Segmentation — mean IoU per class, not overall. Rare classes are where the value is; an overall figure is dominated by road and sky.
- Tracking — identity persistence through occlusion, and identity-switch count per sequence. This is not measurable frame by frame, which is why vendors who audit per-frame miss it entirely.
- Classification and attributes — F1 per class, plus a confusion matrix. Which classes get confused with which is more actionable than the headline number.
- Subjective tasks — chance-corrected agreement between independent annotators, because on intent and behaviour there is no gold answer, only a defensible consensus.
Ask what the vendor does below threshold: who pays for rework, at what turnaround, and how the root cause feeds back into annotator training rather than being fixed silently batch by batch.
How should edge-case and operational design domain coverage be specified?
Define the operational design domain first, then require the dataset to be stratified against it rather than merely sampled from whatever was recorded; this is the most important part of the specification and the one most often left implicit.
An operational design domain (ODD) is the set of operating conditions, including environment, geography, time of day and roadway characteristics, under which a driving automation system is designed to function.
- Lighting — daylight, dusk and dawn, night, low sun directly into the sensor, tunnel entry and exit.
- Weather — rain, spray, fog, snow, wet reflective surfaces.
- Density — empty road through dense urban traffic.
- Actors — pedestrians including children, cyclists, motorcycles, wheelchairs, animals, unusual vehicles, people carrying or pushing objects.
- Occlusion — partial, heavy, and re-appearance after full occlusion.
- Road structure — construction zones, temporary markings, unmarked roads, complex junctions, roundabouts.
- Regional variation — signage, markings, vehicle types and driving conventions differ by country, and a model trained in one region degrades in another.
Stratum coverage = Annotated sequences in stratum ÷ Target sequences in stratum
Report the minimum across strata, not the mean: the mean says the programme is on schedule; the minimum names the condition in which the model will fail.
What does sensor-fusion consistency require?
Where several modalities describe the same scene, the labels must agree in identity, in geometry and in how conflicts are resolved.
Sensor-fusion annotation labels one scene across camera, LiDAR and radar so that every object carries a single identity and a geometrically consistent position in each modality.
Three checks worth writing into the specification:
- Cross-modal identity — the same object carries the same identity in camera and LiDAR.
- Projection consistency — a 3D cuboid projected into the image plane lands on the object. Cheap to check automatically and a reliable detector of calibration drift.
- Disagreement handling — a defined rule for what happens when modalities conflict, rather than an annotator's ad hoc choice. Conflicts are informative and should be logged, not silently resolved.
The nuScenes dataset shows what multi-sensor ground truth looks like: 1,000 scenes from six cameras, five radars and one LiDAR, with 3D boxes for 23 classes and tracking metrics that count identity switches. What a fusion shift looks like in practice is described in the account of how a LiDAR annotation shift actually runs.
Why does temporal consistency need sequence-level review?
Sequence data has failure modes that no per-frame audit finds: identities that switch through occlusion, dimensions that pulse frame to frame, headings that flip 180 degrees on symmetric objects, and interpolated frames that drift from reality between keyframes.
Temporal consistency means an object keeps the same identity, stable dimensions and a physically coherent trajectory across every frame of a sequence, including frames in which it is hidden.
Require sequence-level review, and ask specifically how interpolation is used and validated. Interpolation between keyframes is a legitimate efficiency and a common source of quiet error. Sequence-level review also changes pricing; see the image, video and 3D/LiDAR annotation pricing guide.
Should you use a crowd or a managed workforce for perception data?
A managed, retained workforce, in almost every case. Perception taxonomies take weeks to learn; an open crowd pays that learning curve repeatedly through churn, while a retained, trained workforce pays it once.
Ask for annotator retention on projects of comparable length, and what happens to quality when a team turns over. Then ask about specialist review for safety-critical categories: who adjudicates ambiguous cases, and whether "I am not sure" has a defined escalation path. A programme without one produces confident wrong labels, which are more dangerous than gaps because they pass an acceptance check. How two managed providers differ on these points is worked through in the comparison of Lifewood and iMerit for physical AI annotation.
How should provenance, residency and security be handled?
Assume the footage contains faces and licence plates, because it was recorded in public space, and specify where it may be processed, at what stage it is blurred, who may access it and who labelled and reviewed every item. Under the GDPR, video in which a person is identifiable is personal data.
- Residency — can work be confined to a named jurisdiction or facility? Recorded-in-public data frequently cannot cross certain borders.
- Privacy processing — face and plate blurring, at what stage, and whether the original is retained and under what control.
- Access model — least privilege, revocation, audit logs; physical controls where warranted, including secure rooms and no removable media.
- Per-item provenance — who labelled it, who reviewed it, under which guideline version.
- Certifications — ask for the certificate and its scope statement, not the logo. Scope is where these usually fail: a certification covering a head office says nothing about the delivery centre doing your work.
The controls that apply to any sensitive annotation programme are set out in the guide to enterprise data annotation security, privacy and compliance.
How do you run a pilot that proves any of this?
A proposal cannot demonstrate any of these requirements; a paid pilot of about three weeks can.
Build the pilot set deliberately: a majority of ordinary sequences, plus a deliberate minority of hard ones — night rain, heavy occlusion, a construction zone, an unusual actor, a sequence with a long full occlusion and re-appearance. Include at least one sequence you have already annotated internally to a standard you trust, and do not tell the vendor which one it is.
Score on: per-task metrics against your thresholds; identity switches across the occlusion sequence; projection consistency between modalities; per-class F1 on rare classes; how ambiguous cases were escalated and documented; and the guideline questions the vendor raised — a vendor asking sharp questions in week one is reading the taxonomy properly. Independent AI data validation scores pilot output against a gold set.
How does Lifewood approach autonomous driving annotation?
Lifewood delivers perception annotation, including 2D and 3D bounding boxes, semantic segmentation and keypoint labelling for autonomous driving and medical imaging, through a managed workforce in owned delivery centres rather than an open crowd, the model that suits complex taxonomies where the learning curve is the cost.
Two structural points matter for this category specifically. Owned centres, 40+ delivery centres across 30+ countries, make it practical to confine processing to a named jurisdiction, which recorded-in-public driving data often requires. And regional coverage across Asia, Europe, North America and Africa means datasets can be annotated by people who recognise local signage, markings and driving conventions rather than inferring them. Automotive and vision engagements span AI compute vendors, autonomous-mobility developers and computer-vision suppliers.
The company-reported scope on Lifewood's autonomous driving annotation page covers LiDAR, camera and radar annotation, including multi-frame tracking, lane markup, sign and signal recognition and fusion alignment, delivered from dedicated autonomous-vehicle centres in Malaysia and Indonesia. Quality is governed by a 95%+ accuracy SLA, a 95%+ inter-annotator agreement threshold against a customer-approved gold set, and two independent review passes with timestamped approval records, which supply the per-item provenance. Lifewood also reports 414,120 training hours across its Bangladesh workforce in 2025.