Skip to main content
AI Data

How a New Annotation Centre Is Opened and Trained

Short answer. Opening a centre is procurement; getting it to contracted accuracy is the work, and it takes weeks to months across four phases — recruitment, qualification testing…

Mumu D. · August 2026 · 11 min read

Download PDF

Short answer. Opening a centre is procurement; getting it to contracted accuracy is the work, and it takes weeks to months across four phases — recruitment, qualification testing, calibration rounds, then a supervised production ramp. Because annotation aptitude does not show up on a CV, recruitment runs a pool far larger than the target headcount (commonly 150–200 candidates for 60 seats) and lets the qualification test do the selecting. That test is set above the production accuracy target, since supervised performance always exceeds performance at production pace; a pass rate above 70% means the test is too easy, and below 25% means the problem is upstream in recruitment or the specification.

Opening a new delivery centre is the easy part. You sign a lease, you install the workstations, you run the network cable.

That takes weeks and it is mostly a procurement exercise.

Getting that centre to production quality is a different problem entirely, and it is the one that determines whether the expansion was worth doing. A room full of capable people who have not yet been calibrated to a client's specification will produce work that looks reasonable and fails QA. The gap between "the centre is open" and "the centre is delivering at contracted accuracy" is measured in weeks or months, and how you spend that time is the whole game.

Lifewood has done this repeatedly. The Cebu operations centre in the Philippines is one of the more established sites, running out of the I2 Building in Cebu City's Asiatown IT district. Bangladesh hubs and a dedicated Voice AI Data Center came later. Then Malaysia as an operations hub for global movement, delivery expansion across Serbia, Japan, the UK, Northern Ireland and Australia, a US site, AV expansion across Malaysia and Indonesia, and most recently Africa centres coming online. Across that expansion the crowd resource network grew from surpassing 20,000 in 2021 to over 56,000 today.

Every one of those sites went through the same sequence. This is what it looks like.


Phase 1: Recruitment, and why it is not a headcount exercise

The instinct when opening a new centre is to hire to a number. You need 60 annotators, so you recruit 60 people, train them, and start.

That approach produces attrition and rework. Annotation is a skill with a genuine aptitude component, and the people who are good at it are not always the ones who look strongest on paper. Attention to fine detail sustained over hours, comfort with ambiguity, willingness to flag uncertainty rather than guess, and the discipline to follow a specification exactly rather than apply personal judgement: these are the traits that predict a good annotator, and none of them shows up reliably in a CV.

What works better is recruiting a candidate pool substantially larger than the target headcount and letting the qualification stage do the selection. For a target of 60 production annotators, a pool of 150 to 200 candidates entering qualification is a reasonable starting point, though the ratio varies by task complexity and local labour market.

In multilingual centres the calculus changes again. For a language programme, the constraint is not general annotation aptitude but verified fluency in the target language and variety. Recruiting for Cebuano, Wolof or Fon means reaching into local networks, universities and community organisations rather than posting a generic job advert, and it means screening for the specific regional variety the programme requires rather than the standard national language.

This is a large part of why physical presence matters. A recruitment process for a language with a small professional talent pool cannot be run remotely from another continent. Someone has to know which university department teaches the relevant linguistics, which community organisations have reach into the right speaker population, and which local job platforms people actually use.


Phase 2: Qualification testing

Qualification is where candidates become annotators, and it is deliberately harder than the production work.

A qualification test is not a general aptitude assessment. It is a set of items drawn from the actual task, with known correct answers, covering the full range of difficulty the annotator will encounter. It should include:

Straightforward items that establish baseline comprehension of the task.

Edge cases near label boundaries, where the specification's decision rules matter. These are the discriminating items: a candidate who handles the easy items well and the boundary cases badly has not understood the specification, they have pattern-matched to the obvious.

Deliberately ambiguous items where the correct answer is to flag uncertainty rather than to guess. Candidates who confidently label these are demonstrating exactly the behaviour that produces silent errors in production.

Items requiring the specific knowledge the programme needs, whether that is regional language variety, domain vocabulary, or the sensor physics knowledge a LiDAR programme depends on.

The pass threshold is set against the programme's accuracy requirement with headroom, because a candidate performing at the accuracy floor during a supervised test will typically perform below it at production pace. A common approach is to set the qualification threshold several points above the production accuracy target.

Pass rates are informative in themselves. A qualification pass rate of 70% or higher usually means the test is too easy and is not discriminating. A pass rate below 25% usually means either the recruitment filter is wrong or the specification is unclear enough that even capable candidates cannot apply it. The test tells you about the guidelines as much as it tells you about the candidates.


Phase 3: Calibration rounds

This is the phase most often compressed under schedule pressure, and it is the one that determines whether the centre stabilises or spends its first six months in rework.

Calibration is a structured cycle: a batch of items is annotated by the whole new cohort, inter-annotator agreement is measured, disagreements are surfaced and discussed as a group, the guideline is clarified where the disagreement reveals ambiguity, and the cycle repeats on a fresh batch.

The critical insight, and it is one that experienced annotation managers understand and newer ones do not, is that disagreement in early calibration is usually a guideline problem, not an annotator problem. If eight of twenty annotators interpret a category boundary one way and twelve interpret it another way, the instruction for that category is ambiguous. Retraining the eight will not fix it. Clarifying the guideline will.

Practically, a calibration round runs like this:

Round 1 typically produces low agreement, often a Cohen's kappa or Krippendorff's alpha in the 0.4 to 0.6 range for a complex task. This is normal and expected. The output of round 1 is not a quality score; it is a list of ambiguities in the specification.

The disagreement review is the substance of the phase. Every item where agreement fell below threshold is examined by the group with a senior annotator or the client's specification owner present. The decision is made, documented, and the guideline is updated with a worked example.

Round 2 on a fresh batch should show meaningful improvement. If it does not, the guideline update did not address the actual source of confusion, and the review needs to go deeper.

Rounds continue until agreement stabilises above the programme threshold. Three to five rounds is typical for a moderately complex task. Highly subjective tasks may need more, and some tasks never reach high agreement because the underlying judgement is genuinely contested, which is itself a finding worth reporting to the client.

Gold items are seeded from calibration onwards, so per-annotator accuracy tracking begins before production does.

Annotators whose individual accuracy sits consistently below the cohort during calibration get targeted coaching rather than waiting for production QA to surface the problem.


Phase 4: Supervised production ramp

The centre does not go from calibration to full production volume. It ramps.

Week 1 of production runs at reduced volume with elevated QA. Sampling rates of 25 to 30% are typical for a new cohort, substantially above the 10 to 15% that established annotators with a clean track record would see. Every annotator's output is visible to the QA layer from the first day.

Shift leads run daily calibration check-ins in the early weeks: a short session reviewing the previous day's rework flags and any new edge cases that surfaced. This is where a problem that would otherwise compound across a batch gets caught in 24 hours instead of two weeks.

Volume increases as accuracy stabilises, not on a fixed schedule. A cohort that hits the accuracy threshold in week two can ramp faster than one that takes five weeks, and forcing the schedule produces exactly the rework the ramp exists to avoid.

QA sampling rates step down as individual annotators build track records. This is per annotator, not per cohort: a strong performer moves to a lower sampling rate while a struggling one stays at elevated review until their accuracy stabilises.


What the timeline actually looks like

Clients ask for a number, so here is an honest range rather than a marketing one.

For a familiar task type in an established language where the specification is mature and the centre is being staffed with experienced annotators, four to six weeks from opening to production quality is achievable. Qualification runs in week one, calibration in weeks two and three, supervised ramp from week four.

For a new task type or a complex specification, six to ten weeks is more realistic. The additional time goes almost entirely into calibration rounds, because a new specification has more ambiguity in it and each round of clarification takes a cycle to validate.

For a new language programme where no prior corpus or guideline exists, add several weeks before any of this begins. Orthographic conventions have to be decided, the label taxonomy may need adaptation to the language, and the qualification test itself has to be built in the target language rather than translated into it.

The variable that most reliably extends the timeline is guideline maturity, not annotator capability. A centre staffed with excellent annotators working from an ambiguous specification will produce inconsistent output for as long as the ambiguity persists. A centre with average annotators and a precise, well-exampled specification will stabilise faster.

This is why the calibration phase is the one worth protecting when schedules compress. Cutting a week from recruitment costs you some candidate quality. Cutting a week from calibration costs you months of rework.


What travels between centres, and what does not

Running this playbook across a network of 40-plus centres in 30-plus countries produces a body of institutional knowledge that new sites inherit, which is a genuine advantage over building each one from scratch.

What travels: the phase structure itself, the qualification test design methodology, the calibration cycle, QA sampling rate rules, the gold set construction approach, and accumulated guidance on which edge cases in a given task type reliably cause disagreement. When a new centre opens on an existing programme, it inherits a mature guideline that has already absorbed several rounds of clarification elsewhere.

What does not travel: local recruitment networks, language variety expertise, cultural knowledge relevant to the annotation task, and understanding of local field conditions. These have to be built in each place, by people who live there.

The Cebu centre's decade-plus of accumulated process knowledge transfers directly to a new site. Its knowledge of which Cebuano regional forms matter for a speech programme does not transfer to a Wolof or Fon programme in West Africa. That is precisely why the expansion is into places rather than just capacity: the transferable part is the playbook, and the nontransferable part is the reason to be there at all.


Key takeaways

  • Opening a centre is a procurement exercise; getting it to production quality is the real work, measured in weeks to months.
  • Recruit a candidate pool substantially larger than target headcount, commonly 150 to 200 candidates for 60 production seats, and let qualification do the selection.
  • Annotation aptitude traits, sustained attention to detail, comfort with ambiguity, willingness to flag uncertainty, do not show up reliably on a CV.
  • Multilingual centres recruit for verified fluency in the specific regional variety, which requires local networks rather than generic job adverts.
  • Qualification tests should mix straightforward items, boundary cases, deliberately ambiguous items where flagging uncertainty is correct, and programme-specific knowledge items.
  • Set qualification thresholds above the production accuracy target, since supervised test performance exceeds production-pace performance.
  • Pass rates above 70% suggest the test is too easy; below 25% suggests a recruitment or specification problem.
  • Calibration rounds measure inter-annotator agreement, surface disagreements, clarify the guideline and repeat.
  • Round 1 agreement of 0.4 to 0.6 is normal for complex tasks.
  • Early calibration disagreement is usually a guideline problem, not an annotator problem. Retraining will not fix an ambiguous instruction.
  • Three to five calibration rounds is typical; production ramp begins at 25 to 30% QA sampling and steps down per annotator as track records build.
  • Realistic timelines: four to six weeks for a familiar task in an established language, six to ten weeks for a new task type, longer where no prior corpus or guideline exists.
  • Guideline maturity extends timelines more than annotator capability does, which is why calibration is the phase to protect when schedules compress.
  • Process knowledge travels between centres; local recruitment networks, language variety expertise and cultural knowledge do not.

Sources and further reading

Frequently asked questions

Four to six weeks for a familiar task type in an established language with a mature specification. Six to ten weeks for a new task type or complex specification.

Because annotation aptitude does not show up reliably on a CV, and qualification testing is a better selection mechanism than interviewing. A pool of 150 to 200 candidates for 60 production seats is a common starting ratio.

Straightforward items to establish baseline comprehension, edge cases near label boundaries, deliberately ambiguous items where flagging uncertainty is the correct response, and items requiring the specific domain or language knowledge the programme needs.

Usually that the guideline is ambiguous, not that the annotators are weak. Round 1 agreement in the 0.4 to 0.6 range is normal for complex tasks. The output of round 1 is a list of specification ambiguities to resolve.

Calibration. Cutting recruitment time costs some candidate quality; cutting calibration time costs months of downstream rework because the ambiguities were never resolved.

The phase structure, qualification test methodology, calibration cycle, QA sampling rules, gold set construction approach and a mature guideline that has already absorbed clarification rounds elsewhere. Local recruitment networks and language variety expertise have to be built locally.

Have an AI or visibility project in mind?

From AI evaluation and human-in-the-loop review to GEO and AEO strategy, our team can help you deploy with confidence and get found in the AI search era.

Talk to our team