Short answer. Three different data products get bought under the label "LLM training data," and they do different jobs. SFT (supervised fine-tuning) teaches the model what a good response looks like, and needs written demonstrations — expensive, expert-dependent, the highest-leverage per item. RLHF (reinforcement learning from human feedback) teaches the model which of two responses is better, and needs preference comparisons — cheaper per item, needs far more of them, and lives or dies on rater agreement. Distillation transfers behaviour from a stronger model into a smaller one, and needs generated data plus human validation. Most enterprise programmes need SFT first, evaluation data second, and RLHF third — and buy them in the reverse order, which is the most common and most expensive sequencing mistake.
Teams arriving at this market usually know they need "human data" and are less clear about which kind, in what proportion, and in what order. The vocabulary does not help: vendors sell all three from the same page, and the pricing units are not comparable. This guide separates them — what each actually is, what data it consumes, how quality is measured for it, and how to sequence spend.
Key takeaways
- SFT, RLHF and distillation are three distinct data products with different costs, volumes and expertise requirements, not interchangeable line items under "training data."
- SFT needs a human to write the target response; RLHF needs a human to judge which of two responses is better; distillation needs a human to validate machine-generated data.
- Rater agreement, measured with a chance-corrected statistic such as Cohen's kappa, is the single number that predicts whether preference data will teach a model anything.
- The correct buying order is an evaluation set first, SFT second, preference data third, and distillation last — the reverse of how most programmes actually spend.
- Preference data is culturally situated and should be produced by in-market native speakers per language, not translated from an English set.
What is the difference between SFT, RLHF and distillation?
SFT teaches a model what a good answer looks like by showing it a written example; RLHF teaches it which of two answers is better by ranking; distillation teaches a smaller model to imitate a larger one, with a human validating the result.
| SFT | RLHF / preference data | Distillation | |
|---|---|---|---|
| What it teaches | What a good answer looks like | Which of two answers is better | How a stronger model behaves |
| Human produces | A written demonstration response | A ranking or comparison, with rationale | Validation and filtering of generated data |
| Cost per item | Highest | Moderate | Lowest per item, highest in compute |
| Volume needed | Lower | Higher | Highest |
| Expertise needed | High — the writer must be able to produce the target quality | Moderate to high — the rater must be able to judge it | Moderate, concentrated in review |
| Main quality risk | Inconsistent style and depth across writers | Low rater agreement; rubric ambiguity | Inheriting the teacher model's errors |
| Measured by | Rubric conformance, expert review | Chance-corrected agreement between raters | Downstream evaluation; error inheritance checks |
The economic point buried in that table: writing is harder than judging, and judging is harder than validating. Your budget should follow the difficulty, and so should your sourcing — the pools of people who can do each are different sizes and have different costs.
What does an SFT dataset actually require?
An SFT dataset is prompt-response pairs where a human writes the response the model should have produced, and its quality is capped by the writer, the rubric and the coverage of the prompt distribution.
Supervised fine-tuning (SFT) is the most direct way to move a model's behaviour, because it shows the target rather than scoring attempts at it. Enterprises use it for domain adaptation — legal, medical, financial, industrial support — and for voice and format conformance, where a model must produce output in a specific house style, structure or register. These are the same legal, medical and financial SFT datasets that carry the highest per-item cost and the highest leverage.
What decides quality, in order:
- Writer capability. The demonstration is a ceiling. A writer who cannot produce expert-quality output produces a dataset that teaches the model to be a non-expert. For specialist domains this means qualified writers, verified rather than self-declared.
- Rubric specificity. "Be helpful and accurate" produces inconsistent data. The rubric should specify structure, depth, hedging behaviour, refusal behaviour, citation practice and formatting — anything you would notice being wrong.
- Coverage of the prompt distribution. Demonstrations should span the distribution of prompts your users actually send, including the awkward ones, not the prompts that are pleasant to answer.
Before buying volume, commission twenty items across your hardest cases from two different writers and read them side by side. Variance between writers on the same prompt is the number that predicts your dataset's consistency.
How is RLHF preference data quality measured?
Preference data quality is measured by chance-corrected agreement between independent raters, and low agreement almost always signals an under-specified rubric rather than poor raters.
RLHF (reinforcement learning from human feedback) works by having a human see two or more model responses and say which is better, usually with a rationale and often with per-dimension ratings — helpfulness, accuracy, safety, tone. Enterprises use it to align behaviour where "better" is easier to recognise than to write: tone, refusal calibration, verbosity, adherence to policy, and it is also how safety behaviour is tuned in practice.
Rater agreement is the whole ballgame. If independent qualified raters disagree about which response is better, the preference signal is noise and the model learns noise. Because low agreement is the cheapest diagnostic available, it should be run on the very first batch rather than discovered after volume production. Three practical requirements follow:
- A per-dimension rubric, not a single "which is better." Aggregate preference collapses trade-offs — a response that is more accurate and less pleasant loses for reasons nobody recorded. Writing one that raters actually agree on is a skill in itself; see how to write a preference rubric raters agree on.
- Rationales captured. They are how you debug the rubric, and they are more informative than the preference labels themselves during the first weeks.
- Ties permitted and defined. Forcing a choice between two equivalent responses manufactures signal that is not there.
Preference is also culturally situated — politeness, directness and appropriate hedging differ by market. Preference data for a language should be produced by in-market native speakers, not by translating an English preference set, which produces a model that is polite in an English way in every language.
What are the risks of buying distillation data?
The main risk is that a smaller model inherits its teacher's errors, including the confident ones that read fluently enough to pass light human review.
Distillation is the process of having a stronger model generate training data that trains a smaller or cheaper one; the human role moves from producing to validating and filtering — deciding what is good enough to keep. Enterprises use it for cost and latency reduction, on-device or edge deployment, and expanding coverage of a task where a strong model already performs it well. Guard against training on AI-generated data without validation with four requirements:
- A validation rate that is stated and defended, with sampling designed to find rare errors rather than confirm common correctness.
- Error-class tracking on the teacher, so known weaknesses are screened for specifically rather than hoped against.
- Provenance labelling. Synthetic items must be identifiable in the corpus so their proportion can be controlled and their effect isolated during evaluation.
- Attention to the licence position of the teacher model's outputs for the intended use — this is a contractual question, not a technical one, and it is easier to answer before the data exists.
Why does an evaluation set matter more than any training product?
An evaluation set is the only way to tell whether SFT, RLHF or distillation spend actually worked, which makes it the highest-return purchase on this list even though it is not a training product itself.
Requirements: built independently of the training data, covering the same distribution including the hard tail, refreshed periodically to limit overfitting, and — for multilingual programmes — built per language rather than translated. Contamination screening against the training corpus is worth the effort; a benchmark the model has memorised is worse than no benchmark, because it produces confident wrong decisions.
In what order should a team buy these products?
The evaluation set comes first, SFT second, preference data third once the rubric is stable, and distillation last once behaviour is right and the goal shifts to cheaper or faster inference.
- Evaluation set first. You cannot manage what you cannot measure, and every later decision is made against this.
- SFT next, targeted narrowly at the specific behaviours that are wrong. A small, high-quality, well-covered SFT set usually beats a large diffuse one.
- Preference data third, once the rubric is stable and rater agreement has been demonstrated on a pilot batch.
- Distillation last, when behaviour is right and the goal is cheaper or faster inference.
The common inversion — buying large volumes of preference data before the rubric is stable and before an evaluation set exists — produces a dataset with low agreement, no way to prove it helped, and no diagnosis available afterwards.
What should a buyer ask a training-data vendor?
A buyer should ask which of the three products the vendor delivers in-house, how raters and writers are qualified, and what happens when a client asks for volume before the rubric is ready.
- Which of the three do you actually deliver in-house, and which do you subcontract?
- How are specialist SFT writers qualified — verified or self-declared?
- Show me rater agreement figures from a comparable preference project, per task family.
- What is your rubric development process, and who owns the rubric?
- How many rationales do you capture, and can we read them?
- For multilingual work: are preference raters in-market native speakers?
- How do you handle disagreement and edge cases — what is the escalation path?
- For distillation: what is your validation rate and how is the sample drawn?
- What is your annotator retention on projects of our length?
- What do you tell clients who ask for volume before the rubric is stable?
The last question is the most revealing. A vendor whose answer is "we would sell it to them" is selling throughput; one who describes pushing back is selling outcomes.
How does Lifewood approach RLHF, SFT and distillation?
Lifewood delivers all three in-house, through a managed workforce in owned delivery centres rather than an open crowd, so the same qualified raters stay with a rubric long enough for agreement figures to become meaningful.
Lifewood delivers RLHF, SFT and data distillation, with prompt and response evaluation across 100+ languages. That matters most for preference data: agreement figures only become meaningful when the same qualified raters stay with a rubric long enough for it to stabilise, which is precisely what high-churn sourcing prevents.
The multilingual dimension is the structural one. Preference is culturally situated, so preference data for a market should be produced in that market; 40+ delivery centres across 30+ countries and 56,000+ registered contributors make in-market rating practical in languages where the alternative is translating an English preference set. Engagements span frontier-model labs, voice-AI developers, AI compute vendors, computer-vision suppliers and autonomous-mobility programmes, with client identities withheld by agreement. Teams comparing accuracy commitments across vendors before buying can use what accuracy standard to require from an annotation vendor as a checklist, and Lifewood's own scope is detailed on the enterprise LLM training data and AI data validation pages.