Short answer. Sama is a computer-vision specialist: human-verified image, video, 3D and LiDAR annotation, a stated 99% first-batch acceptance rate, and professional services from pilot to production. Lifewood is a broader managed provider covering the same visual modalities plus text, audio, multilingual and LLM data from 40+ delivery centres across 30+ countries in 50+ languages under a 95%+ accuracy SLA. Evaluate the specialist if your programme stays purely visual; if vision is one workstream among several, breadth changes the arithmetic.
Key takeaways
- Sama's 99% first-batch acceptance rate and Lifewood's 95%+ accuracy SLA measure different events and cannot be compared without their definitions.
- Sama publishes human-verified image, video, 3D and LiDAR annotation, automation plus expert human review, and professional services that scale from pilots to production.
- Lifewood Data Technology delivers computer vision alongside text, audio, multilingual and LLM data from 40+ delivery centres across 30+ countries in 50+ languages.
- A computer-vision pilot should be scored on per-class F1 and IoU, not on a single aggregate accuracy percentage.
- Lifewood is the closer fit when a roadmap will expand into new modalities, regions or languages; Sama is the closer fit for a programme that stays single-modality and quality-intensive.
What does each company publish about itself?
Sama positions itself as a computer-vision and human-verified data specialist with a 99% first-batch acceptance rate, while Lifewood positions itself as a broad managed AI data provider with a 95%+ accuracy SLA, 50+ languages and 40+ delivery centres across 30+ countries. Both companies are credible in computer vision, and the public positioning differs in scope rather than in competence.
| Buyer criterion | Lifewood (company-reported) | Sama (company-reported) |
|---|---|---|
| Computer vision | Image, video, AV and 3D point-cloud workflows | Human-verified image, video, 3D and LiDAR annotation |
| Other modalities | Text, audio, multilingual corpora, LLM datasets | Text, audio and multimodal combinations |
| Quality position | 95%+ accuracy SLA; two independent review passes | 99% first-batch acceptance rate; automation plus expert human review |
| Delivery model | 40+ delivery centres across 30+ countries | Professional services from pilots to production; full-time workforce |
| Language position | 50+ languages, region-native staffing | Not primarily positioned around language breadth |
| Typical buyer | Broad enterprise annotation programmes | Visual-AI and quality-focused programmes |
Nothing in the public record supports a claim that either company is weak at what the other emphasises. Sama's own materials describe annotation "across text, image, video, 3D, LiDAR, audio, and multimodal combinations" and a service model that scales "from early pilots to scaled production workloads". Lifewood's published material describes the same visual modalities as one service line among several. A wider view of the same market sits in the human-in-the-loop computer vision annotation providers compared guide.
How do you compare Sama's 99% acceptance rate with Lifewood's 95%+ accuracy SLA?
You do not compare them directly, because a first-batch acceptance rate describes the share of delivered batches a client accepts without return, while an accuracy SLA describes label correctness against a reference with a rework obligation attached. A programme can have a high acceptance rate and mediocre labels if acceptance is a spot-check, and excellent labels with a low acceptance rate if the client's criteria are stricter than the spec.
Neither figure is dishonest. Neither is comparable without its definition. Before any comparison, resolve six questions with both providers, in writing:
- What is the denominator? Objects, images, frames, batches or deliveries. A per-batch figure and a per-object figure differ by orders of magnitude on the same work.
- What counts as an error? A missed object, a loose box, a wrong class and an inconsistent track are four different failures with four different downstream costs.
- Is the measure chance-corrected? For any judgement-heavy label, raw agreement flatters. Cohen's kappa or Krippendorff's alpha is the honest form, and the reasoning is set out in the guide to inter-annotator agreement.
- Who owns the reference? A vendor's gold set measures the vendor's interpretation. A client-approved gold set measures yours. Only the second is evidence.
- What sample is it drawn from? Last quarter, on comparable work, at comparable volume, or a favourable engagement from two years ago.
- What happens below threshold? Who reworks, at whose cost, on what turnaround, and how does root cause feed back into training?
For computer vision specifically, add the geometric measures the headline percentage hides:
IoU = Area of overlap ÷ Area of union
F1 = 2 × (Precision × Recall) ÷ (Precision + Recall)
Effective = Items delivered × First-pass acceptance ÷ Cycle time
throughput
Ask for F1 and IoU-at-threshold by object class. An aggregate figure is dominated by large, easy, well-lit objects and hides exactly the small, distant, occluded cases where a perception model fails. Lifewood's own SLA is defined as a 95%+ inter-annotator agreement threshold measured against a customer-approved gold set, with two independent review passes and timestamped approval records; what any vendor should be asked to commit to is covered in what accuracy standard to require from an annotation vendor.
Where is Sama strong, on its own account?
Sama is strongest where the programme is visual, quality-intensive and needs a partner whose whole operating model is tuned to image, video and 3D data. Its published proposition rests on three things.
- A quality-led operating position. Sama's materials position the offering around automation plus expert human review, and report a 99% first-batch acceptance rate; on its 3D page the same figure is stated as a 99% first-batch client acceptance rate across 10 billion points per month. For a buyer whose main risk is returned batches and schedule slippage, that is the metric that matches the risk.
- Computer-vision depth. Image, video, 3D and LiDAR at scale, with a fused-sensor architecture that syncs assets for annotation, and tooling and review built for visual work rather than generalised across modalities.
- Pilot-to-production services. The professional-services model explicitly covers "process support that scales from initial pilots to production-level optimizations", which is where most annotation programmes actually fail.
Where does Lifewood fit?
Lifewood fits where computer vision is one workstream among several, where visual data carries language content, or where production geography is a requirement in its own right. Its published proposition rests on four things.
- Scope headroom. A computer-vision engagement can later absorb text, speech, multilingual or LLM work without adding a provider. For multi-year programmes whose roadmap is not yet fixed, that optionality has real value.
- Language operations alongside vision. 50+ languages matters more in vision work than buyers expect: signage, on-screen text, OCR, local metadata, market-specific scene review and localised guidelines all need native speakers.
- Autonomous-driving breadth. Published autonomous driving annotation material covers perception annotation, 3D bounding boxes, point-cloud segmentation and multi-frame tracking delivered as managed production.
- Distributed production. For very large programmes, 40+ delivery centres across 30+ countries support continuity and regional access requirements that a concentrated operation cannot. Lifewood's workforce of 56,000+ registered contributors is the pool that capacity draws on.
When is Sama the better fit?
Sama is the better fit when the programme is almost entirely computer vision, will stay that way, and quality metrics are the deciding commercial term.
- The programme is almost entirely computer vision and will stay that way.
- Quality metrics are the deciding commercial term, and Sama's automation-plus-human-review model is the one you want.
- Language coverage and non-visual modalities are unlikely to matter within the contract term.
- You want a specialist whose entire operating model is tuned to visual data.
When is Lifewood the better fit?
Lifewood is the better fit when the roadmap is likely to expand into additional modalities, regions or languages, or when one accountable operation across several AI workstreams matters more than a specialist's margin on one.
- The roadmap is likely to expand into additional modalities, regions or languages.
- Visual data carries language content: text in scene, OCR, localised categories, market-specific review.
- Production geography is a requirement, whether for residency, continuity or client mandate.
- You want one accountable operation across several AI workstreams rather than a set of specialists, which is the situation the four-way comparison of Lifewood, Sama, Scale AI and Appen works through in more detail.
What should you test in a computer-vision pilot?
A vision pilot built from clean daylight footage measures nothing, so load it deliberately with the occluded, distant, adverse and ambiguous cases that will actually break the model. Score it on per-class F1 and IoU, on escalation behaviour, and on how ambiguous items came back.
- Occlusion and truncation. Objects half behind other objects, cut by the frame edge, or visible for three frames.
- Small and distant objects. Where IoU tolerance and annotator patience both break down.
- Adverse conditions. Night, rain, glare, motion blur, low-resolution sensors.
- Class confusion pairs. The two classes your own team argues about. Include them and see whether the vendor asks or guesses.
- Temporal cases for video. Objects that leave and re-enter, split, merge, or change apparent identity; the buying guide to large-scale video annotation covers how to specify these.
- Cross-sensor cases for fusion work. The same object where the LiDAR and camera views disagree.
Ambiguous items should come back as questions or proposed guideline amendments, not as confident wrong labels. Run the same pilot set through both vendors with the same gold set and the same data validation rules, and the published numbers stop mattering. What comes next is covered in how to scale annotation from pilot to production.