Skip to main content
AI Data

Annotation for Robotics and Physical AI: Manipulation, Egocentric Video and Affordance

September 2026 · 10 min read · Updated September 2026

Short answer. Robotics annotation is not autonomous-driving annotation at a larger scale — it asks annotators to describe what a pair of hands is doing to an object, frame by frame, in three dimensions, from the actor's own viewpoint. That means hand-pose labels, contact-state transitions, task segmentation and affordance annotation, each with its own guideline and reviewer. It is slower and harder to get right, and it is where demand is growing fastest.

Key takeaways

  • Robotics annotation is a different discipline from driving annotation, not a bigger version of it: driving asks what is in a scene, robotics asks what a hand is doing to an object.
  • The bottleneck in 2026 physical-AI programmes is data quality and coverage, not model architecture or compute, according to multiple market analysts covering the sector.
  • Egocentric video — footage captured from the wearer's own point of view — matters because it matches what a robot-mounted camera sees, while third-person footage teaches recognition rather than performance.
  • Five label types come off the same footage — hand pose, contact state, gaze and proximity, task segmentation, and affordance — and each needs its own guideline, gold set and reviewer.
  • Affordance is the hardest of the five: datasets routinely confuse it with object function or goal-related action, and most ignore human motor capacity altogether.

Why did robotics data suddenly get so big?

Because humanoid robots stopped being demos. Once robots go into real warehouses on real production schedules, the bottleneck moves from model architecture to data — and data has to be collected by people.

The scale-up is visible from several directions at once. Figure AI reported in January 2026 that its BotQ facility had delivered more than 350 Figure 03 units and lifted production from one robot a day to one an hour. Tesla began Optimus Gen 3 production at Fremont the same month. Analysts covering the sector describe 2026 as the year the constraint on enterprise physical-AI programmes shifted from architecture and compute to data quality and distribution coverage, with models trained on carefully curated real-world demonstration data outperforming models trained on simulation alone, even at ten times the simulated scale.

Interest is climbing on the demand side too. US monthly search volume for "physical AI" grew roughly 3.5x in twelve months, driven by humanoid programmes, open-source policy releases and the arrival of factory-style data pipelines in place of passive web scraping. What buyers are actually procuring, according to marketplaces serving them, is egocentric video, teleoperation traces, manipulation demonstrations and evaluation sets, with commercial rights and consent artefacts attached — the same discipline Lifewood applies across autonomous driving annotation, where sensor licensing and consent tracking are just as central to delivery.

At Lifewood we have watched this arrive as a shift in the questions clients ask. Two years ago a robotics enquiry meant LiDAR and bounding boxes, much like the work behind top autonomous driving annotation programmes. Now it starts with wearables, hand tracking and contact states — and almost always ends with a question about how many countries and kitchens a vendor can collect in.

What does egocentric annotation actually involve?

Labelling what the robot itself will see, rather than what an outside observer would see. That single constraint reshapes everything downstream.

Egocentric video is footage captured from the wearer's own point of view, aligned with the perspective of a robot-mounted camera rather than a bystander's. Models trained on third-person footage learn to recognise actions from the outside; models trained on egocentric footage learn to perform them, because it shows the hands, the moment of contact, and the exact pixels a robot will see as it reaches. That is why egocentric capture has become the efficient way to expand manipulation datasets without buying more robots, an extension of the same multimodal data annotation discipline used to synchronise video, depth and audio in other domains.

The datasets that resulted are large, and one in particular reset expectations. Apple built EgoDex with the Vision Pro: 829 hours of 30 Hz egocentric video across 194 tabletop manipulation tasks, with SE(3) annotations for 25 joints of both hands in every frame, tracked on-device using calibrated cameras and visual-inertial SLAM. It carries language annotation, camera extrinsics and dexterous annotation together — a combination none of the earlier datasets offered.

Others went to work at industrial scale instead. Egocentric-1M, released in April 2026, captured 2,153 factory workers across real industrial sites: 1.08 billion frames at 1080p and 30fps, 16.4TB, the first egocentric dataset collected exclusively in factories rather than homes or labs. EgoVerse, from a consortium including Georgia Tech, Stanford, UC San Diego, ETH Zürich, MIT and Meta Reality Labs, contributed 1,362 hours across 1,965 tasks, 240 scenes and 2,087 demonstrators from multiple countries, with multi-robot co-training showing gains of up to 30% relative improvement across embodiments. EgoScale demonstrated a log-linear scaling law between human data volume and validation loss, with that loss correlating strongly to downstream robot performance — in effect, more annotated human demonstration reliably buys a better robot, which is why collection capacity has become the commercial constraint.

Five distinct label types come off the same footage, and each has its own failure mode, guideline and reviewer:

Label type What the annotator produces Why it matters to the policy Difficulty
Hand pose Per-frame 3D joint positions for both hands, finger by finger Fine-grained precision that wrist-only or gripper tracking cannot supply Hardware-assisted, review-heavy
Contact state The exact frame where contact begins and ends, per object A policy that cannot detect contact from its own viewpoint fails at placement Genuinely ambiguous
Gaze and proximity Where the demonstrator looked; gripper-to-target spatial relations Approach-phase errors cause grasp failures Needs synced capture
Task segmentation Atomic action boundaries and natural-language descriptions Lets language-conditioned policies map instructions to motion Guideline-sensitive
Affordance Which region of an object affords which action, and with which grasp Enables generalisation to objects never seen in training The hardest of the set

Treating these five as one task, with one guideline and one annotator, is the most common mistake in the field.

Why is affordance labelling the hard part?

Because the definition itself is contested, and most datasets get it wrong in the same three ways.

Affordance is the concept of action possibility: what a given object permits a given actor to do, based on the object's physical properties and the actor's motor capacity. Useful in principle, slippery in practice. Researchers at the University of Tokyo, proposing an annotation scheme for egocentric action video, identified three recurring problems in existing datasets: they mix up affordance with object functionality, they confuse affordance with goal-related action, and they ignore human motor capacity altogether. Their proposed fix combines goal-irrelevant motor actions with grasp types as the label, and adds the notion of mechanical action to capture what is possible between two objects — a guideline problem closely related to the work covered in writing annotation guidelines that annotators actually follow.

Read as an annotation brief, the implications land quickly. A knife's functionality is cutting. Its affordances include being gripped by the handle, pinched at the blade for a handover, and pressed down with the palm — different labels, and a guideline that does not separate them produces a dataset where "knife" means whatever each annotator assumed it meant that day. The field acknowledges the shortfall directly: affordance research still faces data scarcity, poor generalisation and difficulty deploying to the real world, with a specific lack of large-scale affordance datasets carrying precise segmentation maps.

Automated affordance extraction from egocentric video is advancing and should be used, but it inherits the same definitional problem — an automatic pipeline is only as coherent as the label schema it was built against. This is where the human layer, and disciplined inter-annotator agreement measurement against a gold set, earns its cost.

The practices that separate working programmes from stalled ones sort cleanly into two columns:

What works in practice Where programmes come unstuck
Separate guidelines per label type, not one document for all five Treating affordance as a synonym for object function
Grasp taxonomy agreed and illustrated before collection starts Ambiguous contact frames resolved silently
Contact-state adjudication by a second reviewer Demonstrator pools drawn from one country or one body type
Synced multimodal capture: video, depth, audio, gaze One annotator labelling all five types on the same clip
Diverse settings and demonstrators — kitchens, factories, countries Lab-only collection that never sees a real workspace
Consent artefacts and commercial rights captured at source PII in first-person footage discovered after delivery

Robots trained on narrow demonstrator diversity generalise about as well as that description suggests. Human demonstrators are the sensor; recruit and calibrate them like one.

What should teams get right before scaling?

The unglamorous parts, mostly, and earlier than feels necessary: privacy handling and demonstrator diversity, both of which are far cheaper to design in than to retrofit.

Contact state is the annotated moment where a hand or tool begins and ends physical contact with an object, and it is the most contested label in the pipeline because the boundary is often genuinely ambiguous frame to frame. Getting it wrong at scale is expensive to fix after delivery, which is why it belongs alongside privacy on the list of things to settle before collection starts, not during review.

Egocentric footage records whatever the wearer looked at, including faces, screens and documents nobody consented to. End-to-end PII removal, compliant storage and full audit trails have become table stakes for enterprise buyers, and ISO 27001 and SOC 2 are now baseline requirements rather than differentiators in robotics procurement — the same bar covered in enterprise annotation security and compliance. Retrofitting that after collection is painful and sometimes impossible.

A dataset collected in one lab by twenty graduate students produces a policy that works in that lab. The datasets driving real progress went the other way: thousands of demonstrators, hundreds of scenes, multiple countries. That is a logistics problem before it is an annotation problem, and it is the kind of work a delivery network across 40+ centres in 30+ countries was built for — recruiting demonstrators in genuinely different kitchens, workshops and warehouses, with native-language briefing so task instructions mean the same thing everywhere, backed by the same validation discipline applied to every dataset before delivery.

Before scaling collection, a programme should have:

  • Written the affordance schema first, separating function, goal-related action and motor affordance explicitly, with the grasp taxonomy fixed up front.
  • Split the five label types across specialised reviewers, each with its own guideline and gold set.
  • A written tie-break rule and second-reviewer adjudication for contested contact frames.
  • A deliberately diverse demonstrator pool across handedness, hand size, height, working style, culture and setting.
  • Consent and commercial-use rights captured at the moment of capture, not reconstructed afterward.
  • PII removal built into the pipeline itself, not bolted on after delivery.
  • Automated pre-labels used for pose, human judgement reserved for contact and affordance.
  • Simulation and real-world results checked against each other, since curated real demonstration data has outperformed simulation-only training even at ten times the scale.

Frequently asked questions

For pre-training, often yes — EgoDex, Ego4D, EgoVerse and Egocentric-1M are substantial. For a specific robot in a specific workspace, no. Open datasets rarely contain your objects, your tooling or your failure cases, and cross-embodiment transfer still benefits from targeted collection.

Hand-pose tracking is largely hardware-assisted, and automated pipelines for affordance extraction from egocentric video are improving. Contact-state boundaries and affordance schemas still need human judgement, because the difficulty there is definitional rather than perceptual.

It is growing quickly and forecast at a roughly 41% compound annual growth rate, but as a complement rather than a replacement. Real demonstration data has outperformed simulation-only training even when the simulated set was scaled ten times higher.

Multimodal synchronisation, 3D and temporal precision, specialist reviewers per label type, and the logistics of putting wearables on real people in real settings across many countries. It behaves more like running a film shoot than operating a labelling queue.

A policy trained on one lab's demonstrators inherits that lab's handedness, hand sizes and working habits, and generalises poorly outside it. The datasets driving real progress recruited thousands of demonstrators across hundreds of scenes and multiple countries deliberately, not incidentally.

Sources and further reading

  1. Hoque et al., "EgoDex: Learning Dexterous Manipulation from Large-Scale Egocentric Video", arXiv:2505.11709 — 338K trajectories, 194 tasks, 90M frames, language annotation, camera extrinsics and dexterous annotation
  2. "Qwen-RobotManip Technical Report", arXiv:2606.17846 — on egocentric human hand data aligning with robot-mounted camera perspective, and EgoDex capture details (Apple Vision Pro, 829 hours at 30 Hz, SE(3) for 25 joints, visual-inertial SLAM)
  3. Yu, Huang, Furuta, Yagi, Goutsu & Sato (University of Tokyo), "Precise Affordance Annotation for Egocentric Action Video Datasets", arXiv:2206.05424 — the three recurring annotation errors and the motor-action-plus-grasp-type scheme
  4. "Learning Precise Affordances from Egocentric Videos for Robotic Manipulation", arXiv:2408.10123 — on data scarcity, poor generalisation and deployment difficulty in affordance research
  5. Digital Divide Data, "Why Egocentric Datasets Are Becoming The New Standard For Training Robotics Models" — EgoScale's log-linear scaling law, EgoVerse consortium figures and cross-embodiment gains
  6. Labellerr, "10 Egocentric Datasets Reshaping Robotics and AI in 2026" — Egocentric-1M details (2,153 factory workers, 1.08B frames, 16.4TB) and EgoVerse composition
  7. Labellerr, "7 Top Egocentric Data Service Providers for Robotics 2026" — on first-person versus third-person learning, PII removal and audit-trail requirements
  8. MarketsandMarkets, Embodied AI Market — USD 4.44B (2025) to USD 23.06B (2030) at 39.0% CAGR
  9. SNS Insider, Physical AI Market — USD 5.23B (2025) to USD 87.43B (2035) at 32.53% CAGR
  10. Kaiso Research, Synthetic Data for Physical AI Market — USD 2.03B (2025) to USD 63.95B (2035) at roughly 41% CAGR
  11. MarketsandMarkets, Humanoid Robot Market — Figure AI BotQ production ramp (350+ Figure 03 units, one per hour)
  12. Truelabel, "Physical AI Data Marketplace" — on US search demand for "physical AI" and what buyers procure
  13. DataX Power, "Best Robot Training Data Services 2026" — on data quality and distribution coverage as the 2026 constraint, and real demonstration data outperforming simulation at ten times the scale
  14. Data Science Society, "7 Best Data Annotation Companies for Physical AI & Robotics in 2026" — on ISO 27001 and SOC 2 as baseline enterprise requirements

Have an AI or visibility project in mind?

From AI evaluation and human-in-the-loop review to GEO and AEO strategy, our team can help you deploy with confidence and get found in the AI search era.

Talk to our team