Skip to main content
AI Data

How a New Annotation Centre Is Opened and Trained

September 2026 · 9 min read · Updated September 2026

Short answer. Opening a centre is procurement; getting it to contracted accuracy is the work, and it takes weeks to months across four phases — recruitment, qualification testing, calibration rounds, then a supervised production ramp. Because annotation aptitude does not show up on a CV, recruitment runs a pool far larger than the target headcount (commonly 150–200 candidates for 60 seats) and lets the qualification test do the selecting. That test is set above the production accuracy target; a pass rate above 70% means the test is too easy, and below 25% means the problem is upstream in recruitment or the specification.

Key takeaways

  • Opening a centre is a procurement exercise; getting it to production quality is the real work, measured in weeks to months.
  • Recruit a candidate pool substantially larger than target headcount, commonly 150 to 200 candidates for 60 production seats, and let qualification testing do the selection.
  • Qualification thresholds are set above the production accuracy target, since supervised test performance exceeds production-pace performance.
  • Early calibration disagreement is usually a guideline problem, not an annotator problem; retraining will not fix an ambiguous instruction.
  • Process knowledge travels between centres, but local recruitment networks, language variety expertise and cultural knowledge have to be built in each place.

What does it take to open a new annotation centre?

Opening a delivery centre is a procurement exercise: signing a lease, installing workstations, running network cable. Getting that same centre to contracted accuracy is a separate, longer problem, and it is the one that determines whether the expansion was worth doing.

A room full of capable people who have not yet been calibrated to a client's specification produces work that looks reasonable and fails QA. The gap between "the centre is open" and "the centre is delivering at contracted accuracy" runs weeks to months, across four phases: recruitment, qualification testing, calibration rounds, and a supervised production ramp.

Lifewood has run this sequence repeatedly across its network. The Cebu operations centre in the Philippines, running out of the I2 Building in Cebu City's Asiatown IT district, is one of the more established sites. Bangladesh hubs and a dedicated Voice AI Data Center came later, followed by Malaysia as an operations hub, delivery expansion across Serbia, Japan, the UK, Northern Ireland and Australia, a US site, AV expansion across Malaysia and Indonesia, and most recently Africa centres coming online. Across that expansion the crowd resource network grew from more than 20,000 in 2021 to 56,788 registered contributors today.

Why isn't recruitment just a headcount exercise?

Recruiting to a fixed number produces attrition and rework, because annotation aptitude is a genuine skill and does not show up reliably on a CV. Qualification testing — a structured assessment against known-correct answers that decides who becomes a production annotator — does the real selecting, so recruitment needs to feed it a wide enough pool.

Sustained attention to fine detail over hours, comfort with ambiguity, willingness to flag uncertainty rather than guess, and the discipline to follow a specification exactly rather than apply personal judgement predict a good annotator, and none of these traits is visible on paper. What works better than hiring to a number is recruiting a candidate pool substantially larger than the target headcount and letting qualification select from it. For a target of 60 production annotators, a pool of 150 to 200 candidates entering qualification is a reasonable starting point, though the ratio varies by task complexity and local labour market, a topic covered in more depth in how annotators are recruited, trained and certified for specialist domains.

In multilingual centres the constraint shifts from general annotation aptitude to verified fluency in the target language and variety. Recruiting for Cebuano, Wolof or Fon means reaching into local networks, universities and community organisations rather than posting a generic job advert, and screening for the specific regional variety the programme requires rather than the standard national language. This is a large part of why physical presence matters: a recruitment process for a language with a small professional talent pool cannot be run remotely from another continent. Someone has to know which university department teaches the relevant linguistics, which community organisations reach the right speaker population, and which local job platforms people actually use.

What happens during qualification testing?

A qualification test is a set of items drawn from the actual task, with known correct answers, covering the full range of difficulty an annotator will encounter, and it is deliberately harder than the production work itself.

It should mix four kinds of items: straightforward items that establish baseline comprehension; edge cases near label boundaries, which are the discriminating items because a candidate who handles easy items well and boundary cases badly has pattern-matched rather than understood the specification; deliberately ambiguous items where the correct answer is to flag uncertainty, since confidently labelling these predicts silent errors in production; and items requiring the specific knowledge a programme needs, whether regional language variety, domain vocabulary, or sensor physics for a LiDAR programme.

The pass threshold is set with headroom above the programme's accuracy requirement, because a candidate performing at the accuracy floor during a supervised test typically performs below it at production pace. Pass rates are informative in themselves: a rate above 70% usually means the test is too easy and is not discriminating, and a rate below 25% usually means the recruitment filter is wrong or the specification is unclear even to capable candidates. The test tells you about the guideline as much as it tells you about the candidates.

How do calibration rounds work?

Calibration is a structured cycle in which a cohort annotates the same batch of items, agreement is measured, disagreements are discussed as a group, and the guideline is clarified before the cycle repeats — and it is the phase most often compressed under schedule pressure, at real cost.

The insight experienced annotation managers rely on is that disagreement in early calibration is usually a guideline problem, not an annotator problem: if eight of twenty annotators read a category boundary one way and twelve read it another, the instruction is ambiguous, and retraining the eight will not fix it. Round 1 typically produces low agreement, often inter-annotator agreement — the degree to which independent annotators assign the same label, commonly measured as Cohen's kappa or Krippendorff's alpha — in the 0.4 to 0.6 range for a complex task, and that number is a list of specification ambiguities, not a quality score. See inter-annotator agreement: Cohen's kappa, Krippendorff's alpha and what the numbers mean for how those figures are interpreted.

The disagreement review is the substance of the phase: every item below the agreement threshold is examined with a senior annotator or the client's specification owner present, the decision is documented, and the guideline is updated with a worked example. Round 2 on a fresh batch should show meaningful improvement; if it does not, the review needs to go deeper. Three to five rounds is typical for a moderately complex task, and rounds continue until agreement stabilises above the programme threshold. Gold items are seeded from calibration onwards, so per-annotator accuracy tracking — the same mechanism described in gold sets, audit sampling and consensus: three ways to QA annotated data — begins before production does, and annotators whose accuracy sits consistently below the cohort get targeted coaching early.

How does production ramp up safely?

The centre moves from calibration to full volume gradually rather than at once, because volume increases as accuracy stabilises rather than on a fixed calendar.

Week one of production runs at reduced volume with elevated QA: sampling rates of 25 to 30% are typical for a new cohort, well above the 10 to 15% an established annotator with a clean track record would see, and every annotator's output is visible to the QA layer from day one. Shift leads run daily calibration check-ins in the early weeks, a short review of the previous day's rework flags and any new edge cases, catching a compounding problem in 24 hours rather than two weeks. QA sampling then steps down per annotator, not per cohort, as individual track records build — a strong performer moves to a lower sampling rate while a struggling one stays at elevated review until accuracy stabilises.

How long does it actually take?

Clients tend to ask for a single number, so an honest range is more useful than a marketing one. For a familiar task type in an established language with a mature specification, four to six weeks from opening to production quality is achievable, with qualification in week one, calibration in weeks two and three, and supervised ramp from week four; a new task type or complex specification is more realistically six to ten weeks, with the extra time going almost entirely into calibration rounds, because a new specification carries more ambiguity and each clarification round takes a cycle to validate.

For a new language programme where no prior corpus or guideline exists, add several weeks before any of this begins, since orthographic conventions have to be decided and the qualification test itself has to be built in the target language rather than translated into it. The variable that most reliably extends the timeline is guideline maturity, not annotator capability: a centre with average annotators and a precise, well-exampled specification stabilises faster than a centre with excellent annotators working from an ambiguous one. This is why calibration is the phase worth protecting when schedules compress — cutting a week from recruitment costs some candidate quality, but cutting a week from calibration costs months of rework, a trade-off examined in how to scale AI data annotation from pilot to production.

What travels between centres, and what does not?

The phase structure, qualification test design, calibration cycle, QA sampling rules and gold set construction approach travel between centres; local recruitment networks, language variety expertise and cultural knowledge do not, and have to be rebuilt in each place.

Running this playbook across a network of 40+ delivery centres across 30+ countries produces institutional knowledge that new sites inherit rather than rebuild from scratch — how one playbook runs across a global data operation — and when a new centre opens on an existing programme, it inherits a mature guideline that has already absorbed several rounds of clarification elsewhere. The Cebu centre's decade-plus of accumulated process knowledge transfers directly to a new site, but its knowledge of which Cebuano regional forms matter for a speech programme does not transfer to a Wolof or Fon programme in West Africa. That is precisely why the expansion is into places rather than just capacity: the transferable part is the playbook, and the nontransferable part is the reason to be there at all. Lifewood's own comparison of large-scale annotation providers, top large-scale AI data annotation and labelling companies, and its overview of AI data services, set this centre-opening process in the context of the wider annotation market.

Frequently asked questions

Four to six weeks for a familiar task type in an established language with a mature specification, six to ten weeks for a new task type or complex specification, and longer for a new language programme where orthographic conventions and the qualification test itself have to be built first.

Because annotation aptitude does not show up reliably on a CV, and qualification testing is a better selection mechanism than interviewing. A pool of 150 to 200 candidates for 60 production seats is a common starting ratio, adjusted for task complexity and local labour market.

Straightforward items for baseline comprehension, edge cases near label boundaries, deliberately ambiguous items where flagging uncertainty is correct, and items requiring the specific domain or language knowledge the programme needs. The threshold is set above the production accuracy target.

Usually that the guideline is ambiguous, not that the annotators are weak. Agreement in the 0.4 to 0.6 range on round one is normal for complex tasks, and the output is a list of specification ambiguities to resolve, not a quality verdict on the cohort.

Calibration. Cutting recruitment time costs some candidate quality; cutting calibration time costs months of downstream rework because the ambiguities that cause inconsistent labelling were never resolved.

The phase structure, qualification test methodology, calibration cycle, QA sampling rules, gold set construction approach, and a mature guideline that has already absorbed clarification rounds elsewhere. Local recruitment networks and language variety expertise have to be built locally, by people who live there.

Sources and further reading

  1. Bontcheva and Sabou, "Best Practices for Managing Data Annotation Projects" (arXiv), on annotator recruitment, qualification tasks, calibration and sampling frequency adjustment
  2. TaskMonk, "The Ultimate Data Labeling Guide 2026", on calibration reviews, benchmark tasks and go/no-go quality thresholds
  3. TaskMonk, "Data Labeling Quality Guide 2026", on sampling rate ranges for new versus established annotators
  4. Annotera, "9 Best Practices for Data Annotation Quality Assurance 2026", on gold set usage for onboarding and calibration and tiered review structures
  5. Label Your Data, "Annotation QA: Best Practices for ML Model Quality", on inter-annotator agreement as a guideline diagnostic
  6. Lifewood, company timeline, on the Cebu operations centre, Bangladesh hubs, Malaysia operations hub, delivery expansion and Africa centres coming online, and crowd resource growth from 20,000 in 2021 to 56,788

Have an AI or visibility project in mind?

From AI evaluation and human-in-the-loop review to GEO and AEO strategy, our team can help you deploy with confidence and get found in the AI search era.

Talk to our team