Skip to main content
AI Data

How to Recruit Native Contributors for African Language Data

Short answer. Job boards and crowdsourcing platforms do not reach speakers of most African languages, so recruitment runs through relationships instead — research communities…

Mumu D. · July 2026 · 11 min read

Download PDF

Short answer. Job boards and crowdsourcing platforms do not reach speakers of most African languages, so recruitment runs through relationships instead — research communities, universities and existing language networks. Masakhane, founded in 2018, has produced over 400 language models and more than 20 datasets and trained over 100 African data scientists, which makes it the most productive route into these communities. The documented practice is unglamorous: asking a standing community for suggestions across three consecutive weekly meetings. Nekoto et al. (2020) showed communities contribute meaningfully to NLP without formal training.


African Language Data?

Post a job advert for a Fongbe speech contributor and you will get very few applications, most of them unsuitable. Not because the speakers do not exist, roughly two million people speak Fongbe, but because the channel is wrong. Almost nobody who speaks Fongbe as a first language is browsing an international freelance platform looking for annotation work.

This is the recruitment problem in African language data, and it is more determinative of project success than almost anything downstream. A well-designed collection protocol executed against a contributor pool you could not fill produces nothing.

The good news is that the routes that work are well documented, and one research finding in particular should change how teams think about who is eligible.


The finding that widens the pool

The default assumption in data operations is that contributors need training before they can produce usable output. For most annotation tasks that is correct.

For participatory language data collection, the evidence points the other way. Work by Nekoto and colleagues in 2020 demonstrated that communities in low-resource environments contribute significantly to NLP, even without formal training.

That single finding is what makes community recruitment viable at all. It means the eligibility criterion is native fluency and commitment rather than prior technical experience, which expands the addressable pool by orders of magnitude and shifts the burden onto guideline quality and support rather than onto candidate screening.

It does not mean training is unnecessary. It means the training is about the task, not about the language, and it can be delivered to people who have never worked in data before.


Where contributors actually come from

The channels that work are specific and mostly not commercial.

Research and practitioner communities. Masakhane is the central one for African languages: a Pan-African open-source initiative founded in 2018 that has produced more than 400 language models and over 20 datasets, covering languages including Luganda, Yoruba, isiZulu and Amharic, and has trained more than 100 African data scientists.

Practitioner literature explicitly names Masakhane alongside university mailing lists and organisational Slack or Discord channels as the recruitment routes for culturally grounded data work.

The practical mechanics are worth knowing because they are unglamorous. In one documented study, researchers recruited participants by asking the Masakhane community for suggestions during regular weekly meetings across three consecutive weeks, and via the community Slack. Not a campaign. A standing relationship, used repeatedly.

Academic networks. University linguistics and computer science departments in the relevant country, which give access to speakers who already understand structured data work. The Deep Learning Indaba and its IndabaX regional events function as the convening point for this network across the continent.

Community organisations and language associations. For languages without a strong university presence, cultural associations, radio stations and church or civic groups are frequently the only route to speakers of specific varieties.

Structured programmes. In Latin America a comparable model has used structured social service programmes engaging student volunteers in transcription and segmentation, producing eight open-access linguistic resources.

The equivalent structures exist in several African countries and are underused by commercial data programmes.

Crowdsourcing platforms work for widely spoken languages and fail for everything else. Mechanical Turk and Prolific have usable pools for Swahili and effectively none for Fongbe or Fante.


What the community-led model is now doing at scale

Worth knowing because it sets the reference point for ambition and because it changes the competitive landscape.

The Masakhane African Languages Hub is collecting high-quality multimodal data for 50 languages, targeting 500 hours of voice data per language for voice-to-voice translation. That is 25,000 hours if fully achieved, in languages where published corpora currently run to tens of hours.

The funding behind it is substantial. LINGUA Africa brings together the Gates Foundation, Google.org and the Microsoft AI for Good Lab with Masakhane, providing financial and cloud computing support for language AI initiatives.

And the model is being pushed further down to local level. At Deep Learning Indaba 2026 in Lagos, under the theme "Sovereign Intelligence," Masakhane announced Community Pilot Grants creating a direct pipeline from the Indaba to IndabaX communities to build datasets serving African people.

The framing used in that workshop is the most useful sentence I found in this research, and it should reshape how collection projects are scoped. The session was titled "Beyond the Web: Where Does Your Language Live?", opening on the data drought and the hidden wealth of offline sources.

The argument is that while web-scraped African language data is scarce, legally restrictive and culturally decontextualised, a wealth of untapped data sits in archives, radio recordings, oral histories and physical repositories across the continent. Recruitment, in that framing, is not only about finding people to speak into microphones. It is about finding the people who hold and can grant access to material that already exists.


Why "available data" is usually not usable data

Two documented examples explain why recruitment cannot be avoided by using what is already published.

Domain bias. Many baseline African language models were trained on JW300, a large parallel corpus derived from religious publications. The corpus is genuinely useful and it is heavily biased toward religious content, which was explicitly cited as a motivation for launching the AI4D African Language Program. A model trained on it handles scripture register well and conversational or technical register badly.

Orthographic inconsistency. Fongbe uses tonal diacritics such as è, é and ê. Some corpora preserve full diacritics while others strip them entirely, creating compatibility issues when datasets are combined. Non-standard orthography and limited digital presence are named as recurring barriers across African language resource surveys.

There is also a documentation gap that affects anyone trying to build on existing resources. A 2026 survey noted that Masakhane's focus on machine translation means critical metadata for resource selection is frequently absent: dataset sizes in words or sentences, licensing terms, file formats, domain coverage and preprocessing requirements.

The practical consequence is that a project which assumes it can assemble a corpus from existing sources usually discovers, several weeks in, that the sources are religious, undiacritised, unlicensed for commercial use, or all three.


Designing a recruitment process that works

Six things determine whether the pool fills.

Recruit for the variety, not the language. Twi is not one thing, and a contributor from Kumasi and one from Accra may differ in ways that matter for the dataset. Specify the target variety and recruit against it explicitly, because a general call for "Twi speakers" will over-sample whichever variety happens to have the strongest network.

Screen for fluency, not credentials. Given the Nekoto finding, requiring prior annotation experience filters out most of your viable pool for no quality benefit. Screen with a short practical task in the target language instead.

Route through people, not platforms. Every documented success runs through an existing relationship: a community, a department, an association. Budget time for building those relationships before the project starts, because they cannot be created on a two-week timeline.

Explain the purpose plainly. Contributors in participatory projects consistently respond to the argument that the work serves speakers of their language. That is not a marketing frame, it is the actual value proposition, and it recruits better than rate alone.

Offer recognition where appropriate. One documented study offered authorship to participants. That model does not fit every commercial project, but the underlying point holds: contribution to a language resource is something people want credited, and acknowledgement has real recruiting value alongside payment.

Pay properly and say so upfront. I have written elsewhere in this series about Kenya's 2026 draft AI policy proposing fair-pay benchmarking against international rates for annotation and evaluation roles. In a market where the labour conditions debate is live and regulatory attention is increasing, the terms are part of the recruitment proposition, not a back-office detail.


Retention is the harder problem

Recruitment gets attention. Retention determines whether the programme delivers.

For a language with a small qualified pool, a trained contributor is not a replaceable unit. If ten people in a project speak the target variety and produce verified output, losing three is a 30% capacity reduction that takes months to rebuild, because the recruitment channel that found them is a relationship rather than a queue.

Three things reduce attrition specifically in this context.

Predictable work. Sporadic engagement loses people to other commitments. A contributor who worked on a two-week burst nine months ago is functionally a new recruit when you return.

Progression. Contributors who become reviewers, then leads, stay. This also solves the verification staffing problem, since the second-speaker review layer requires exactly the people the collection layer has already trained.

Feedback. Telling contributors what their data was used for and how the model improved is unusually motivating in participatory work, and it costs nothing.


Where we sit in this

Declaring the interest: Lifewood brought delivery centres online in Africa as part of its recent expansion, and recruitment for low-resource language programmes is a substantial part of what those centres do.

Two observations that I think hold regardless of who runs the work.

The first is that recruitment cannot be centralised. A team in Kuala Lumpur cannot recruit Wolof speakers in Dakar, not because of competence but because the channels are local, informal and relationship-based. Knowing which university department teaches the relevant linguistics, which community radio station reaches the right speaker population, and which association convenes people who speak the specific variety you need is knowledge that only exists on the ground.

The second is that the community and commercial models are complementary rather than competing.

Masakhane and similar initiatives are building open resources at a scale and with a legitimacy no commercial operation could replicate, and the sensible commercial position is to build on that foundation rather than around it: hire from those networks, respect their licensing, contribute back where possible, and take on the work they are not structured to do, which is typically the client-specific, deadline-bound, contractually scoped collection that a volunteer community cannot commit to.


What to plan for when scoping

Budget relationship time before project time. Three to six weeks of network building before recruitment opens is normal for a language without an existing pool.

Specify the target variety, not just the language.

Screen practically, not on credentials, and expect capable contributors with no data experience.

Check what already exists before commissioning, and check it for domain bias, diacritic handling and licensing rather than assuming a published corpus is usable.

Ask what offline material exists. Archives, radio recordings and oral history collections may hold more usable audio than any collection programme could produce, and the recruitment task becomes finding the custodians.

Plan retention as deliberately as recruitment, with predictable work, progression paths and feedback loops.

Set pay against international benchmarks and state the terms in the recruitment material.


Key takeaways

  • Job boards and crowdsourcing platforms do not reach speakers of most African languages. Recruitment runs through relationships: research communities, university networks, community organisations and structured programmes.
  • Nekoto et al. (2020) demonstrated that communities in low-resource environments contribute significantly to NLP even without formal training, which makes native fluency rather than prior experience the eligibility criterion.
  • Masakhane, founded in 2018, has produced over 400 language models and more than 20 datasets and trained over 100 African data scientists, and is named in the literature as a primary recruitment channel alongside university mailing lists and organisational Slack channels.
  • Documented recruitment practice is unglamorous: asking a standing community for suggestions across three consecutive weekly meetings and via Slack.
  • The Masakhane African Languages Hub is collecting multimodal data for 50 languages, targeting 500 hours of voice data per language for voice-to-voice translation.
  • LINGUA Africa brings the Gates Foundation, Google.org and the Microsoft AI for Good Lab together with Masakhane for language AI funding and compute.
  • Community Pilot Grants announced at Deep Learning Indaba 2026 in Lagos create a pipeline from the Indaba to IndabaX communities for local dataset building.
  • The workshop framing "Beyond the Web: Where Does Your Language Live?" points at untapped archives, radio recordings, oral histories and physical repositories, making recruitment partly about finding custodians of existing material.
  • Available data is often unusable: JW300 is heavily biased toward religious content, which motivated the AI4D African Language Program.
  • Fongbe tonal diacritics are preserved in some corpora and stripped in others, creating compatibility problems when datasets are combined.
  • Existing African language resources frequently lack metadata on dataset size, licensing terms, file formats, domain coverage and preprocessing requirements.
  • Recruit for the specific variety rather than the language, screen practically rather than on credentials, route through existing relationships, explain purpose, offer recognition, and state pay terms upfront.
  • Retention matters more than in general annotation work, because a small qualified pool means losing three contributors can cut capacity by 30% with a months-long rebuild.
  • Predictable work, progression into review roles and feedback on how data was used are the three most effective retention levers.

Sources and further reading

Frequently asked questions

No. Research demonstrated that communities in low-resource environments contribute significantly to NLP without formal training. Screen for native fluency in the target variety and deliver task training, rather than filtering on credentials.

They have usable pools for a handful of widely spoken languages such as Swahili and effectively none for languages like Fongbe, Fante or Twi. Recruitment for those runs through community, academic and civic networks instead.

Check them first. Many baseline African language resources derive from JW300, which is heavily biased toward religious content, and orthographic handling varies, with some corpora preserving tonal diacritics and others stripping them.

Plan three to six weeks of network building before recruitment opens, because the channels are relationships rather than queues and cannot be created on a short timeline.

Because qualified pools are small. Losing three contributors from a ten-person variety-specific team is a 30% capacity cut that takes months to rebuild through relationship-based channels.

They are complementary. Community projects build open resources at a scale and legitimacy commercial operations cannot replicate. The sensible position is to hire from those networks, respect their licensing and contribute back, while taking on client-specific deadline-bound work volunteer communities cannot commit to.

Have an AI or visibility project in mind?

From AI evaluation and human-in-the-loop review to GEO and AEO strategy, our team can help you deploy with confidence and get found in the AI search era.

Talk to our team