Skip to main content
AI Data

RLHF, SFT and Distillation: What Enterprise Teams Buy

July 2026 · 9 min read · Updated September 2026

Short answer. Three different data products get bought under the label "LLM training data," and they do different jobs. SFT (supervised fine-tuning) teaches the model what a good response looks like, and needs written demonstrations — expensive, expert-dependent, the highest-leverage per item. RLHF (reinforcement learning from human feedback) teaches the model which of two responses is better, and needs preference comparisons — cheaper per item, needs far more of them, and lives or dies on rater agreement. Distillation transfers behaviour from a stronger model into a smaller one, and needs generated data plus human validation. Most enterprise programmes need SFT first, evaluation data second, and RLHF third — and buy them in the reverse order, which is the most common and most expensive sequencing mistake.

Teams arriving at this market usually know they need "human data" and are less clear about which kind, in what proportion, and in what order. The vocabulary does not help: vendors sell all three from the same page, and the pricing units are not comparable. This guide separates them — what each actually is, what data it consumes, how quality is measured for it, and how to sequence spend.

Key takeaways

  • SFT, RLHF and distillation are three distinct data products with different costs, volumes and expertise requirements, not interchangeable line items under "training data."
  • SFT needs a human to write the target response; RLHF needs a human to judge which of two responses is better; distillation needs a human to validate machine-generated data.
  • Rater agreement, measured with a chance-corrected statistic such as Cohen's kappa, is the single number that predicts whether preference data will teach a model anything.
  • The correct buying order is an evaluation set first, SFT second, preference data third, and distillation last — the reverse of how most programmes actually spend.
  • Preference data is culturally situated and should be produced by in-market native speakers per language, not translated from an English set.

What is the difference between SFT, RLHF and distillation?

SFT teaches a model what a good answer looks like by showing it a written example; RLHF teaches it which of two answers is better by ranking; distillation teaches a smaller model to imitate a larger one, with a human validating the result.

SFT RLHF / preference data Distillation
What it teaches What a good answer looks like Which of two answers is better How a stronger model behaves
Human produces A written demonstration response A ranking or comparison, with rationale Validation and filtering of generated data
Cost per item Highest Moderate Lowest per item, highest in compute
Volume needed Lower Higher Highest
Expertise needed High — the writer must be able to produce the target quality Moderate to high — the rater must be able to judge it Moderate, concentrated in review
Main quality risk Inconsistent style and depth across writers Low rater agreement; rubric ambiguity Inheriting the teacher model's errors
Measured by Rubric conformance, expert review Chance-corrected agreement between raters Downstream evaluation; error inheritance checks

The economic point buried in that table: writing is harder than judging, and judging is harder than validating. Your budget should follow the difficulty, and so should your sourcing — the pools of people who can do each are different sizes and have different costs.

What does an SFT dataset actually require?

An SFT dataset is prompt-response pairs where a human writes the response the model should have produced, and its quality is capped by the writer, the rubric and the coverage of the prompt distribution.

Supervised fine-tuning (SFT) is the most direct way to move a model's behaviour, because it shows the target rather than scoring attempts at it. Enterprises use it for domain adaptation — legal, medical, financial, industrial support — and for voice and format conformance, where a model must produce output in a specific house style, structure or register. These are the same legal, medical and financial SFT datasets that carry the highest per-item cost and the highest leverage.

What decides quality, in order:

  1. Writer capability. The demonstration is a ceiling. A writer who cannot produce expert-quality output produces a dataset that teaches the model to be a non-expert. For specialist domains this means qualified writers, verified rather than self-declared.
  2. Rubric specificity. "Be helpful and accurate" produces inconsistent data. The rubric should specify structure, depth, hedging behaviour, refusal behaviour, citation practice and formatting — anything you would notice being wrong.
  3. Coverage of the prompt distribution. Demonstrations should span the distribution of prompts your users actually send, including the awkward ones, not the prompts that are pleasant to answer.

Before buying volume, commission twenty items across your hardest cases from two different writers and read them side by side. Variance between writers on the same prompt is the number that predicts your dataset's consistency.

How is RLHF preference data quality measured?

Preference data quality is measured by chance-corrected agreement between independent raters, and low agreement almost always signals an under-specified rubric rather than poor raters.

RLHF (reinforcement learning from human feedback) works by having a human see two or more model responses and say which is better, usually with a rationale and often with per-dimension ratings — helpfulness, accuracy, safety, tone. Enterprises use it to align behaviour where "better" is easier to recognise than to write: tone, refusal calibration, verbosity, adherence to policy, and it is also how safety behaviour is tuned in practice.

Rater agreement is the whole ballgame. If independent qualified raters disagree about which response is better, the preference signal is noise and the model learns noise. Because low agreement is the cheapest diagnostic available, it should be run on the very first batch rather than discovered after volume production. Three practical requirements follow:

  • A per-dimension rubric, not a single "which is better." Aggregate preference collapses trade-offs — a response that is more accurate and less pleasant loses for reasons nobody recorded. Writing one that raters actually agree on is a skill in itself; see how to write a preference rubric raters agree on.
  • Rationales captured. They are how you debug the rubric, and they are more informative than the preference labels themselves during the first weeks.
  • Ties permitted and defined. Forcing a choice between two equivalent responses manufactures signal that is not there.

Preference is also culturally situated — politeness, directness and appropriate hedging differ by market. Preference data for a language should be produced by in-market native speakers, not by translating an English preference set, which produces a model that is polite in an English way in every language.

What are the risks of buying distillation data?

The main risk is that a smaller model inherits its teacher's errors, including the confident ones that read fluently enough to pass light human review.

Distillation is the process of having a stronger model generate training data that trains a smaller or cheaper one; the human role moves from producing to validating and filtering — deciding what is good enough to keep. Enterprises use it for cost and latency reduction, on-device or edge deployment, and expanding coverage of a task where a strong model already performs it well. Guard against training on AI-generated data without validation with four requirements:

  • A validation rate that is stated and defended, with sampling designed to find rare errors rather than confirm common correctness.
  • Error-class tracking on the teacher, so known weaknesses are screened for specifically rather than hoped against.
  • Provenance labelling. Synthetic items must be identifiable in the corpus so their proportion can be controlled and their effect isolated during evaluation.
  • Attention to the licence position of the teacher model's outputs for the intended use — this is a contractual question, not a technical one, and it is easier to answer before the data exists.

Why does an evaluation set matter more than any training product?

An evaluation set is the only way to tell whether SFT, RLHF or distillation spend actually worked, which makes it the highest-return purchase on this list even though it is not a training product itself.

Requirements: built independently of the training data, covering the same distribution including the hard tail, refreshed periodically to limit overfitting, and — for multilingual programmes — built per language rather than translated. Contamination screening against the training corpus is worth the effort; a benchmark the model has memorised is worse than no benchmark, because it produces confident wrong decisions.

In what order should a team buy these products?

The evaluation set comes first, SFT second, preference data third once the rubric is stable, and distillation last once behaviour is right and the goal shifts to cheaper or faster inference.

  1. Evaluation set first. You cannot manage what you cannot measure, and every later decision is made against this.
  2. SFT next, targeted narrowly at the specific behaviours that are wrong. A small, high-quality, well-covered SFT set usually beats a large diffuse one.
  3. Preference data third, once the rubric is stable and rater agreement has been demonstrated on a pilot batch.
  4. Distillation last, when behaviour is right and the goal is cheaper or faster inference.

The common inversion — buying large volumes of preference data before the rubric is stable and before an evaluation set exists — produces a dataset with low agreement, no way to prove it helped, and no diagnosis available afterwards.

What should a buyer ask a training-data vendor?

A buyer should ask which of the three products the vendor delivers in-house, how raters and writers are qualified, and what happens when a client asks for volume before the rubric is ready.

  1. Which of the three do you actually deliver in-house, and which do you subcontract?
  2. How are specialist SFT writers qualified — verified or self-declared?
  3. Show me rater agreement figures from a comparable preference project, per task family.
  4. What is your rubric development process, and who owns the rubric?
  5. How many rationales do you capture, and can we read them?
  6. For multilingual work: are preference raters in-market native speakers?
  7. How do you handle disagreement and edge cases — what is the escalation path?
  8. For distillation: what is your validation rate and how is the sample drawn?
  9. What is your annotator retention on projects of our length?
  10. What do you tell clients who ask for volume before the rubric is stable?

The last question is the most revealing. A vendor whose answer is "we would sell it to them" is selling throughput; one who describes pushing back is selling outcomes.

How does Lifewood approach RLHF, SFT and distillation?

Lifewood delivers all three in-house, through a managed workforce in owned delivery centres rather than an open crowd, so the same qualified raters stay with a rubric long enough for agreement figures to become meaningful.

Lifewood delivers RLHF, SFT and data distillation, with prompt and response evaluation across 100+ languages. That matters most for preference data: agreement figures only become meaningful when the same qualified raters stay with a rubric long enough for it to stabilise, which is precisely what high-churn sourcing prevents.

The multilingual dimension is the structural one. Preference is culturally situated, so preference data for a market should be produced in that market; 40+ delivery centres across 30+ countries and 56,000+ registered contributors make in-market rating practical in languages where the alternative is translating an English preference set. Engagements span frontier-model labs, voice-AI developers, AI compute vendors, computer-vision suppliers and autonomous-mobility programmes, with client identities withheld by agreement. Teams comparing accuracy commitments across vendors before buying can use what accuracy standard to require from an annotation vendor as a checklist, and Lifewood's own scope is detailed on the enterprise LLM training data and AI data validation pages.

Frequently asked questions

SFT gives the model written demonstrations of the response it should produce; RLHF gives it comparisons showing which of several responses is better. SFT is more expensive per item and needs fewer items, because a demonstration carries more information than a preference. RLHF is cheaper per item, needs far more of them, and depends entirely on raters agreeing with each other.

An evaluation set, then SFT targeted at the specific behaviours that are wrong, then preference data once the rubric is stable, then distillation if cost or latency is the remaining problem. Buying preference data first — the common pattern — produces low-agreement data and no way to prove whether it helped.

By chance-corrected agreement between independent raters, per task family, plus rubric conformance. Raw agreement is misleading on skewed comparisons. Low agreement usually indicates an under-specified rubric rather than poor raters, and it should be measured on the first pilot batch rather than after volume production.

It should not be. Preference judgements encode culturally situated expectations about politeness, directness, hedging and appropriate detail. A translated English preference set trains a model to be polite in an English way in every language — fluent and subtly wrong everywhere.

The student inherits the teacher's errors, including confident ones that read fluently and pass light review. Manage it with a defended validation rate, sampling designed to surface rare errors, explicit screening for the teacher's known weak classes, and provenance labelling so synthetic items can be isolated during evaluation. Licence terms for the teacher model's outputs are a separate and equally important question.

There is no universal figure — it depends on the base model, the gap being closed, and the narrowness of the task. The reliable planning heuristic is relative: SFT needs the fewest items and the highest quality per item; preference data needs substantially more items at lower cost each; distillation needs the most and the lightest human touch per item. Start small on each, measure against the evaluation set, and scale the one that moves it.

Sources and further reading

  1. Cohen's Kappa Explained — definition and formula for the chance-corrected agreement measure used to score preference and judgement tasks.
  2. Lifewood: Enterprise LLM Training Data — scope of Lifewood's RLHF, SFT and distillation delivery.

Have an AI or visibility project in mind?

From AI evaluation and human-in-the-loop review to GEO and AEO strategy, our team can help you deploy with confidence and get found in the AI search era.

Talk to our team