Short answer. Buy an off-the-shelf dataset when a prebuilt corpus already matches your task, languages, modality and rights position, and time-to-start matters more than fit. Commission a custom dataset when the model's advantage depends on data your competitors do not have, or when public and licensed corpora do not represent your deployment environment — proprietary workflows, unusual equipment, local dialects, confidential documents, expert reasoning, rare operational scenarios. The comparison that decides it is not licence fee against project fee. It is cost per item that survives to training, because a cheap corpus that needs relicensing, deduplication, relabelling, format conversion and language review is not cheap. Most mature programmes end up with both, and the thing that makes the combination workable is source-level provenance.
The choice is usually framed as build versus buy and settled on headline price, which is the wrong axis. A licensed dataset and a commissioned one fail in different ways: the licensed one fails by not fitting, the commissioned one fails by being specified badly. Both failures are diagnosable before purchase, and this guide sets out how.
The two options side by side
| Off-the-shelf | Custom | |
|---|---|---|
| Time to first data | Days | Weeks to months |
| Fit to your deployment conditions | Whatever the collector needed | Whatever you specify |
| Rights position | Fixed by the licence, negotiable at the margin | Defined by you at commissioning |
| Differentiation | None — competitors can buy it too | The reason to do it |
| Ontology and schema | The vendor's | Yours, designed around the model objective |
| Cost profile | Low upfront, variable downstream rework | High upfront, low downstream rework |
| Main risk | Coverage mismatch discovered after training | A specification that describes the wrong thing |
| Refresh | Depends on the vendor's roadmap | Depends on your budget |
The differentiation row is the strategic one and the differentiation argument is often overstated. Most models are not differentiated by their pre-training corpus at all; they are differentiated by a comparatively small volume of task-specific material that nobody else has. That observation usually shifts the right answer toward licensed breadth plus custom depth rather than toward either extreme.
When is an off-the-shelf dataset the right purchase?
- Prototyping and baselines, where the point is to establish whether an approach works at all before specifying anything precisely.
- Standard tasks in well-served modalities — common speech recognition, general object detection, mainstream language pairs — where the public and commercial supply is genuinely good.
- Volume supplementation alongside a custom core, where breadth matters and precision does not.
- Evaluation baselines, where comparing against a published benchmark is part of the requirement.
The gating question is not quality but fit against your deployment conditions. A well-built dataset collected in the wrong environment is not a partial answer; it is a source of confident errors, because the model learns conditions that do not obtain where it runs.
When does a custom dataset become necessary?
Five signals, any one of which is usually sufficient:
- The data is the moat. If the model's advantage depends on material competitors cannot buy, buying it defeats the purpose.
- Public corpora do not represent deployment. Unusual equipment, proprietary workflows, specific acoustic or visual environments, industrial processes, in-house document formats.
- Language or dialect coverage is the gap. Below the well-served languages, the licensed supply thins quickly and its quality becomes hard to assess from outside.
- Domain expertise is required to produce it. Expert reasoning, specialist judgement and procedural knowledge cannot be sourced from a general corpus at any volume.
- Rights need to be unambiguous. Where the use is high-stakes or the corpus will be redistributed, commissioning is often the only way to get a clean rights position per item.
Customisation also buys something less obvious: you get to define the ontology. The schema, prompts, metadata and quality thresholds are designed around your model objective rather than around whatever the original collector needed, which is what determines whether the data can be re-used across future models rather than serving one.
How to compare cost honestly
Sticker price is the least informative number in the comparison. Compare on usable output.
Cost per usable item = (Licence or project fee + Cleaning + Relabelling + Legal review + Engineering + Storage + Refresh) ÷ Items that survive to training
The denominator is where licensed datasets are most often misjudged. Items get discarded for licence terms that exclude commercial model training, coverage that does not match the deployment environment, annotation against an incompatible taxonomy, formats requiring conversion, duplication against material you already hold, and language quality that fails review. None of that appears on the invoice.
The numerator is where custom programmes are most often misjudged, in the opposite direction. A custom project fee usually includes specification, recruitment, capture, annotation and QA — work that a licensed dataset also requires, but performs internally and books to engineering headcount rather than to the data budget. Compare like for like by costing the internal effort on both sides.
Two further items belong in the arithmetic and are routinely omitted. Refresh, because a corpus in a fast-moving domain has a shelf life and someone will pay for the next version. And re-use value, because a custom dataset with good metadata, clear rights and a documented schema can serve several models, while one built without them serves exactly one.
Due diligence before licensing a dataset
Ask for documentation, then verify it against the data rather than reading it.
- Rights and licence. Specifically: is commercial model training permitted, is redistribution of derived models permitted, are there field-of-use or territorial restrictions, and what happens on termination.
- Provenance. Source, collection method, collection dates, and consent basis where personal data is involved. A dataset that cannot describe its own origin is a liability whatever its quality.
- Coverage, by stratum. Not the total. Counts by language, region, device, scenario and class, so you can compare them against your own coverage plan.
- Annotation method. Guideline availability, annotator qualification, overlap rate, agreement figures and adjudication practice. "Human verified" is not a method.
- Known limitations. A dataset card without a limitations section has not been examined by its own producer.
- Update history. Whether it is maintained, on what cadence, and whether earlier versions remain available.
- Samples from every important segment — not the curated preview. The preview is a selection artefact; ask for a random draw from each stratum you care about.
- A pilot experiment. The only reliable judgement of a dataset is whether it improves the target model or evaluation under the conditions that matter. Run it on a slice before licensing the whole.
Red flags: a licence that is silent on model training; coverage described only in totals; annotation described as "high accuracy" with no metric or method; no named limitations; a refusal to supply random samples per stratum; provenance answered with a description of a process rather than a record.
Why the hybrid works, and what makes it work
Most mature programmes end up licensing breadth and commissioning depth. Licensed data shortens the initial timeline and covers the general case; custom collection addresses the segments that differentiate the model or that licensed data represents badly.
The enabling condition is source-level provenance carried into the merged corpus. Every record needs to retain which source it came from and under which rights, because every subsequent operation depends on it: excluding a source whose licence changed, weighting custom material more heavily, reporting composition to an auditor, honouring a deletion request, or diagnosing which portion of the corpus is responsible for a behaviour change. A merged corpus without source tags is a corpus you cannot unmerge, and that is usually discovered at the worst moment.
Second condition: keep the evaluation set custom and independent. It should be built separately from either source, screened for contamination against both, and — for multilingual work — built per language rather than translated. A benchmark drawn from a licensed corpus that a foundation model may already have seen produces confident and wrong decisions.
How Lifewood approaches this
Lifewood works from the specification rather than from a catalogue: scoping which portions of a requirement can be met by existing material and which need collecting, then building the custom layer to a written coverage plan with consent and provenance recorded per item at intake rather than reconstructed at handover.
The custom side is where the delivery footprint decides what is feasible. 50+ languages, 40+ delivery centres across 30+ countries and 56,788 registered contributors mean collection in the markets where licensed supply is thinnest, produced in-market rather than translated, with dual-layer human-in-the-loop validation held to a 95%+ accuracy threshold. Scope spans multilingual data collection, annotation across modalities, LLM training data and response evaluation. The AI-data heritage runs to 2004, with the current company established in 2018.
See global AI data, multilingual data collection, enterprise LLM training data and AI data validation.
Sources and further reading
- NIST, Artificial Intelligence Risk Management Framework (AI RMF 1.0, January 2023) — on provenance and documentation as lifecycle requirements.
- Companion guides: What Is AI Training Data, and What Makes It Good? and In-House vs Outsourced Annotation Cost.

