Short answer. It travels through seven stages, and a human being is required at nearly every one. A sentence is spoken and recorded under consent, transcribed against a chosen orthography, verified by a second speaker, adjudicated where reviewers disagree, packaged with its provenance, mixed into a training corpus alongside thousands of others, and finally returned to the world as a model's ability to understand someone else speaking the same way. Most recordings that begin the journey never finish it.
Key takeaways
- AI training data for an under-resourced language begins with a person speaking, not with a scrape or a dataset purchase.
- Sylheti has around 11 million speakers and is described by linguists as minoritised, politically unrecognised and understudied, often treated as a dialect despite limited mutual intelligibility with Bengali.
- Published Sylheti corpora are measured in hundreds or low thousands of sentences, against trillions of tokens for English.
- Voice recordings that identify a speaker are treated as biometric data under several regimes, where explicit informed consent is the dependable legal basis.
- A single clip contributes statistically: volume must clear a per-language threshold, and quality-filtered corpora have matched baselines on a fraction of the tokens.
Where does AI training data actually begin?
With a person saying something ordinary that has never been written down. Not with a dataset, a scrape or an API.
Start with one sentence. A woman in her sixties, in a kitchen outside Sylhet in north-eastern Bangladesh, is asked how she would tell someone the way to the nearest pharmacy. She answers in Sylheti, the language she has spoken her whole life, in about eight seconds, with a small laugh in the middle because the question is odd.
That eight-second clip is where a piece of AI training data actually begins, and it is worth pausing on why it cannot come from anywhere else.
Sylheti has roughly 11 million speakers across north-eastern Bangladesh, the Barak Valley in India and diaspora communities in the UK, the US and the Gulf. Linguists describe it as minoritised, politically unrecognised and understudied, and it is widely treated as a dialect of Bengali despite limited mutual intelligibility, which has held back efforts to document and protect it. Bengali itself is under-resourced in AI terms, and its regional varieties are further down again — a pattern covered in more depth when comparing high-resource and low-resource languages in AI training.
The scale of what exists is easy to state. One published parallel corpus covering Sylheti, Chittagonian and Barisali offers on the order of 1,500 words, 130 clauses and 980 sentences per dialect. A later research effort refined the Sylheti and English portion into 1,500 sentence pairs, each translated by native speakers and cross-checked. Compare that to the trillions of tokens available in English and the situation is clear. For this language, the internet is not a source. People are.
What happens the moment speech becomes text?
A decision has to be made about how the language is written, and for many languages that decision is genuinely contested.
The clip now goes to a transcriber, and immediately the project confronts something that never arises in English. Orthography is the writing convention chosen for a spoken language — which script, spelling and punctuation rules a transcript will follow — and for many languages more than one is in active use.
Roughly 3,000 of the world's 7,000-plus languages have an established writing system, which means a great many are predominantly oral: speech, almost to the exclusion of writing, has carried the knowledge. Sylheti sits in an unusual position here. It has its own historic script, Sylheti Nagri, dating back centuries, and it is also written in Eastern Nagari and occasionally in the Latin alphabet. Three options, different communities favouring different ones, no single official standard.
So before a single word is typed, the project must decide which convention it uses, document that decision, and apply it consistently. Get this wrong and the dataset manufactures its own inconsistency: two transcribers spelling the same word differently, both correct, the model learning noise.
Then come the ordinary judgement calls, none of which software can make. Does the laugh get marked? Does the false start get transcribed or cleaned? When she switches to Bengali for three words, is that tagged as code-switching or silently normalised? Code-switching is when a speaker moves between two languages or varieties within the same utterance, and a project must decide in advance whether to tag it or normalise it away — the choices involved are covered in how to collect speech and text that mixes languages. Does a regional pronunciation get written as the standard form or as she said it?
Every one of those is a decision about what the model will learn. In a well-run project the answers are written down in advance and the hard cases go to a senior speaker rather than being resolved individually, ten different ways, by ten different transcribers.
Who checks the work, and against what?
A second native speaker, then a third when the first two disagree. This is the stage where most of the cost sits and most of the value is created.
The transcript now goes to review. A different speaker of the same variety listens to the audio and reads the text, and the interesting part is what they are looking for.
Automated checks have already run and passed. The file is the right length, the right format, the right sample rate, not a duplicate, not silent. None of that tells anyone whether the transcript is right.
The reviewer is checking things only a speaker can catch: whether a word was heard correctly, whether the regional form was preserved or flattened into standard Bengali, whether the code-switch was handled per the guidelines, whether the speaker is actually from the district she was recruited for. Published work in this area uses exactly this pattern. In the Sylheti benchmark mentioned earlier, translations were produced by native speakers and cross-validated, and two native Sylheti speakers scored outputs independently before their scores were averaged to reduce individual variation — the kind of consensus scoring discussed in gold sets, audit sampling and consensus.
When the transcriber and reviewer disagree, the case goes to adjudication, and the decision is recorded with a reason. Adjudication is the step where a senior reviewer settles a disagreement between transcriber and reviewer, and records the reasoning so the ruling can be reused. That last detail is what separates a project that improves from one that repeats itself: the ruling goes back into the guidelines, so the next hundred transcribers inherit the answer rather than re-deriving it.
Our eight-second clip survives this. Many do not. Between capture failures, specification mismatches, consent gaps and failed review, a meaningful share of everything recorded never reaches a dataset. That attrition is normal, and budgeting for it is the difference between a project that delivers and one that overruns.
What travels with a single clip?
Far more than audio. The recording arrives at a training pipeline wrapped in metadata, consent records and its own quality history, and without those it is close to unusable.
By the time our clip is packaged, it carries the audio itself, with its technical properties recorded, and the verified transcript, in the documented orthography.
It also carries speaker metadata: age band, gender, region and language variety, so the dataset can be balanced and audited later. Recording conditions are logged too: device type, environment, background noise level. The consent record links the clip to a documented permission and a compensation record, and its quality history shows who transcribed it, who reviewed it, whether it was adjudicated and on what grounds.
This packet is what makes the clip an asset rather than a liability. Provenance is the documented chain linking a piece of training data to its consent, collection conditions and review history, and a buyer of training data increasingly has to demonstrate it, not merely assert quality — a dataset without a documented consent chain carries risk regardless of how good the audio is. Reconstructing any of this afterwards is close to impossible, which is why it has to be captured while the work happens.
It is also what allows the dataset to be sliced later. When a model underperforms for older speakers in one district, someone can find out, because the metadata makes the question answerable.
Building this reliably in a place like Sylhet is a physical proposition rather than a software one. It needs people on the ground, trained, screened and equipped. Lifewood built voice AI data operations and additional hubs in Bangladesh from 2021 onward for precisely this reason: the work of collecting and verifying speech in regional languages cannot be run remotely from somewhere else, and it sits inside the wider multilingual data collection work the company runs across its delivery centres.
What does a model actually do with it?
Almost nothing, individually. Its value is statistical, which is exactly why the composition of the whole set matters so much.
Our clip now joins several hundred hours of similar recordings and enters training. Taken alone it changes nothing measurable. Taken together with the others, it does three things: it teaches acoustic patterns — how these vowels sound in this region, at this age, over a phone, with a fan running; it contributes to the model's sense of how the language is put together, including the constructions that standard Bengali data would never supply; and it shifts, fractionally, what the tokenizer treats as ordinary, which affects how efficiently the language is processed for the life of the model.
Two things determine whether that contribution counts. The first is volume: below a certain token or hour threshold per language, a contribution is too thin to move anything, and the language ends up listed as supported without being usable.
The second is quality, and here the research is encouraging. Studies of multilingual pretraining have found that quality-filtered corpora can match baseline performance on a small fraction of the tokens, which means a smaller, carefully verified collection can outperform a larger careless one. For a language where every hour is expensive to produce, that is the most important economic fact in the whole pipeline.
This is also the point where the human effort becomes invisible. Nobody using the finished model will see the consent form, the adjudication note or the reviewer who caught a flattened regional vowel. The work disappears into the weights.
What comes back at the end of the journey?
Somebody else, speaking the same way, being understood. That is the entire return on the journey, and it closes the loop back where it started.
Two years later, a different woman in a different district asks a health service assistant on her phone, in Sylheti, where to find a pharmacy that is open. It answers. It does not ask her to repeat herself, does not route her to Bengali, does not mishear a regional word as something else entirely.
She has no idea that a stranger's eight seconds in a kitchen contributed. That is what success looks like: entirely unremarkable, from the user's side.
It is worth being precise about what actually made that possible, because it is easy to attribute to the model. The architecture was probably open and shared internationally. The compute was rented. Neither of those was the constraint.
The constraint was that somebody went to the Sylhet region, found speakers, obtained consent, chose an orthography, recorded, transcribed, verified, adjudicated and documented, several thousand times over — the same operating model that underpins Lifewood's broader work on low-resource speech data and the providers ranked in the top global multilingual AI data collection companies.
That is the journey. It begins with a person speaking and ends with a person being understood, and the whole middle section is other people doing careful work in a language the internet mostly ignores.