Skip to main content
AI Data

From a Local Language to a Global AI Model

Short answer. It travels through seven stages, and a human being is required at nearly every one. A sentence is spoken and recorded under consent, transcribed against a chosen…

Mumu D. · August 2026 · 11 min read

Download PDF

Short answer. It travels through seven stages, and a human being is required at nearly every one. A sentence is spoken and recorded under consent, transcribed against a chosen orthography, verified by a second speaker, adjudicated where reviewers disagree, packaged with its provenance, mixed into a training corpus alongside thousands of others, and finally returned to the world as a model's ability to understand someone else speaking the same way. Most recordings that begin the journey never finish it.


Where does AI training data actually begin?

With a person saying something ordinary that has never been written down. Not with a dataset, a scrape or an API.

Start with one sentence. A woman in her sixties, in a kitchen outside Sylhet in north-eastern Bangladesh, is asked how she would tell someone the way to the nearest pharmacy. She answers in Sylheti, the language she has spoken her whole life, in about eight seconds, with a small laugh in the middle because the question is odd.

That eight-second clip is where a piece of AI training data actually begins, and it is worth pausing on why it cannot come from anywhere else.

Sylheti has roughly 11 million speakers across north-eastern Bangladesh, the Barak Valley in India and diaspora communities in the UK, the US and the Gulf. Linguists describe it as minoritised, politically unrecognised and understudied, and it is widely treated as a dialect of Bengali despite limited mutual intelligibility, which has held back efforts to document and protect it. Bengali itself is under-resourced in AI terms. Its regional varieties are further down again.

The scale of what exists is easy to state. One published parallel corpus covering Sylheti, Chittagonian and Barisali offers on the order of 1,500 words, 130 clauses and 980 sentences per dialect. A later research effort refined the Sylheti and English portion into 1,500 sentence pairs, each translated by native speakers and cross-checked. Compare that to the trillions of tokens available in English and the situation is clear. For this language, the internet is not a source. People are.


What has to happen before the record button?

A specification, a screening process, a consent conversation and a recording setup. The eight seconds are the easy part.

Before that woman is ever asked a question, several things have already been decided by people she will never meet.

A specification exists. Someone determined that the project needs spontaneous speech rather than read sentences, from speakers across a particular age range, in ordinary rooms rather than studios, on the kind of phone people actually own, covering everyday topics like directions, health and money.

She has been found and screened. Not through a job board. For a language like this, contributors are reached through local networks, community organisations and delivery teams already present in the region. She has been checked for the right variety, since Sylheti in one district is not identical to Sylheti in the next.

Consent has been taken properly. She has been told what the recording will be used for, who will hold it and how to withdraw. This is not paperwork. A voice recording that can identify a speaker is treated as biometric data under several regimes, and explicit informed consent is the only dependable basis for using it. Her compensation is agreed in advance.

The prompt has been designed. Open-ended, so she talks naturally, rather than a script that would produce read speech with no hesitation, no laugh and none of the features that make the recording useful.

Only then does anyone press record. And this is where the first losses happen. A door slams. The file clips. The connection drops mid-upload. She switches into Bengali halfway through, which is entirely natural and may or may not satisfy the specification.


What happens the moment speech becomes text?

A decision has to be made about how the language is written, and for many languages that decision is genuinely contested.

The clip now goes to a transcriber, and immediately the project confronts something that never arises in English.

Roughly 3,000 of the world's 7,000-plus languages have an established writing system, which means a great many are predominantly oral: speech, almost to the exclusion of writing, has carried the knowledge. Sylheti sits in an unusual position here. It has its own historic script, Sylheti Nagri, dating back centuries, and it is also written in Eastern Nagari and occasionally in the Latin alphabet. Three options, different communities favouring different ones, no single official standard.

So before a single word is typed, the project must decide which convention it uses, document that decision, and apply it consistently. Get this wrong and the dataset manufactures its own inconsistency: two transcribers spelling the same word differently, both correct, the model learning noise.

Then come the ordinary judgement calls, none of which software can make. Does the laugh get marked? Does the false start get transcribed or cleaned? When she switches to Bengali for three words, is that tagged as code-switching or silently normalised? Does a regional pronunciation get written as the standard form or as she said it?

Every one of those is a decision about what the model will learn. In a well-run project the answers are written down in advance and the hard cases go to a senior speaker rather than being resolved individually, ten different ways, by ten different transcribers.


Who checks the work, and against what?

A second native speaker, then a third when the first two disagree. This is the stage where most of the cost sits and most of the value is created.

The transcript now goes to review. A different speaker of the same variety listens to the audio and reads the text, and the interesting part is what they are looking for.

Automated checks have already run and passed. The file is the right length, the right format, the right sample rate, not a duplicate, not silent. None of that tells anyone whether the transcript is right.

The reviewer is checking things only a speaker can catch: whether a word was heard correctly, whether the regional form was preserved or flattened into standard Bengali, whether the code-switch was handled per the guidelines, whether the speaker is actually from the district she was recruited for. Published work in this area uses exactly this pattern. In the Sylheti benchmark mentioned earlier, translations were produced by native speakers and cross-validated, and two native Sylheti speakers scored outputs independently before their scores were averaged to reduce individual variation.

When the transcriber and reviewer disagree, the case goes to adjudication, and the decision is recorded with a reason. That last detail is what separates a project that improves from one that repeats itself. The ruling goes back into the guidelines, so the next hundred transcribers inherit the answer rather than re-deriving it.

Our eight-second clip survives this. Many do not. Between capture failures, specification mismatches, consent gaps and failed review, a meaningful share of everything recorded never reaches a dataset. That attrition is normal, and budgeting for it is the difference between a project that delivers and one that overruns.


What travels with a single clip?

Far more than audio. The recording arrives at a training pipeline wrapped in metadata, consent records and its own quality history, and without those it is close to unusable.

By the time our clip is packaged, it carries:

The audio itself, with its technical properties recorded. The verified transcript, in the documented orthography.

Speaker metadata: age band, gender, region and language variety, so the dataset can be balanced and audited later.

Recording conditions: device type, environment, background noise level. The consent record, linking the clip to a documented permission and a compensation record. Its quality history: who transcribed it, who reviewed it, whether it was adjudicated and on what grounds.

This packet is what makes the clip an asset rather than a liability. A buyer of training data increasingly has to demonstrate provenance, not merely assert quality, and a dataset without a documented consent chain carries risk regardless of how good the audio is. Reconstructing any of this afterwards is close to impossible, which is why it has to be captured while the work happens.

It is also what allows the dataset to be sliced later. When a model underperforms for older speakers in one district, someone can find out, because the metadata makes the question answerable.

Building this reliably in a place like Sylhet is a physical proposition rather than a software one. It needs people on the ground, trained, screened and equipped. Lifewood built voice AI data operations and additional hubs in Bangladesh from 2021 onward for precisely this reason: the work of collecting and verifying speech in regional languages cannot be run remotely from somewhere else.


What does a model actually do with it?

Almost nothing, individually. Its value is statistical, which is exactly why the composition of the whole set matters so much.

Our clip now joins several hundred hours of similar recordings and enters training. Taken alone it changes nothing measurable. Taken together with the others, it does three things.

It teaches acoustic patterns: how these vowels sound in this region, at this age, over a phone, with a fan running.

It contributes to the model's sense of how the language is put together, including the constructions that standard Bengali data would never supply.

And it shifts, fractionally, what the tokenizer treats as ordinary, which affects how efficiently the language is processed for the life of the model.

Two things determine whether that contribution counts. The first is volume: below a certain token or hour threshold per language, a contribution is too thin to move anything, and the language ends up listed as supported without being usable.

The second is quality, and here the research is encouraging. Studies of multilingual pretraining have found that quality-filtered corpora can match baseline performance on a small fraction of the tokens, which means a smaller, carefully verified collection can outperform a larger careless one. For a language where every hour is expensive to produce, that is the most important economic fact in the whole pipeline.

This is also the point where the human effort becomes invisible. Nobody using the finished model will see the consent form, the adjudication note or the reviewer who caught a flattened regional vowel. The work disappears into the weights.


What comes back at the end of the journey?

Somebody else, speaking the same way, being understood. That is the entire return on the journey, and it closes the loop back where it started.

Two years later, a different woman in a different district asks a health service assistant on her phone, in Sylheti, where to find a pharmacy that is open. It answers. It does not ask her to repeat herself, does not route her to Bengali, does not mishear a regional word as something else entirely.

She has no idea that a stranger's eight seconds in a kitchen contributed. That is what success looks like: entirely unremarkable, from the user's side.

It is worth being precise about what actually made that possible, because it is easy to attribute to the model. The architecture was probably open and shared internationally. The compute was rented. Neither of those was the constraint.

The constraint was that somebody went to the Sylhet region, found speakers, obtained consent, chose an orthography, recorded, transcribed, verified, adjudicated and documented, several thousand times over.

That is the journey. It begins with a person speaking and ends with a person being understood, and the whole middle section is other people doing careful work in a language the internet mostly ignores.


Key takeaways

  • AI training data for an under-resourced language begins with a person speaking, not with a scrape or a dataset purchase.
  • Sylheti has around 11 million speakers and is described by linguists as minoritised, politically unrecognised and understudied, often treated as a dialect despite limited mutual intelligibility with Bengali.
  • Published Sylheti corpora are measured in hundreds or low thousands of sentences, against trillions of tokens for English.
  • Before recording: a specification, local recruitment and screening, informed consent with compensation, and prompt design for spontaneous rather than read speech.
  • Voice recordings that identify a speaker are treated as biometric data under several regimes, where explicit informed consent is the dependable legal basis.
  • Only around 3,000 of the world's 7,000-plus languages have an established writing system, so transcription often begins with choosing an orthography.
  • Sylheti can be written in Sylheti Nagri, Eastern Nagari or Latin script, so the project must document its choice or manufacture its own inconsistency.
  • Verification is done by a second native speaker, with disagreements adjudicated by a senior speaker and the decision written back into the guidelines.
  • A delivered clip carries audio, transcript, speaker metadata, recording conditions, consent records and quality history. Provenance is now a compliance deliverable.
  • A single clip contributes statistically. Volume must clear a per-language threshold, and quality-filtered corpora have matched baselines on a fraction of the tokens.
  • The return on the journey is another speaker of the same language being understood without having to switch.
  • About the author Mumu, AI Executive, Lifewood Specialising in AI data, global multilingual data collection, AEO/GEO, AIGC, and AI quality evaluation.

Sources and further reading

Frequently asked questions

For predominantly spoken and minoritised languages, the material does not exist online in usable volume or quality. Spontaneous speech with verified transcripts and consent records has to be created.

The recording is seconds. Screening, consent, transcription, review and adjudication are what set the timeline, and a meaningful proportion of recorded material is rejected along the way.

Where a language has more than one writing convention and no single standard, transcribers will make different choices unless the project documents one, and the resulting inconsistency is learned by the model as noise.

Yes. Voice that can identify a speaker is treated as biometric data in several jurisdictions, and consent, retention and deletion rules apply. Provenance documentation is increasingly required by buyers as well as regulators.

Individually, almost not at all. Its value is statistical. What matters is whether the collection as a whole clears the volume threshold for that language and passes quality filtering.

That is set by the consent terms agreed before recording, which should state the use, the holder, the compensation and the route to withdraw.

Have an AI or visibility project in mind?

From AI evaluation and human-in-the-loop review to GEO and AEO strategy, our team can help you deploy with confidence and get found in the AI search era.

Talk to our team