Short answer. With credentialed professionals in the loop at every step: experts author and verify the instruction–response demonstrations, work from rubrics with explicit failure modes, resolve disagreements by consensus, and every record is de-identified and documented before training. The stakes justify the cost — general models hallucinate on 69–88% of legal queries and 64% of unmitigated medical summaries — and the research is clear that a small, carefully curated expert set beats orders of magnitude more noisy data.
Key takeaways
- General LLMs hallucinate on 69–88% of legal queries (75%+ on court holdings), purpose-built legal tools still err 17–43% of the time, and unmitigated medical summaries hallucinate at 64.1% — against a under-2% clinical deployment target.
- Domain fine-tuning demonstrably closes the gap: major clinical hallucinations fell 58.8–84.8% in one 2026 study, statute-grounding metrics rose in legal work, and Gartner projects domain-specific models going from 1% to over half of enterprise GenAI by 2027.
- Quality beats quantity: small, carefully curated expert sets match models trained on orders of magnitude more data, and canonical domain evaluation sets run in the thousands of items, not millions.
- Expert SFT records are demonstrations a professional would sign: grounded in the authoritative source, correctly cited, terminologically exact, with refusals where a professional would refuse.
- Privacy is a measured pipeline stage — automated de-identification recalls only 95–98% of identifiers, so a two-pass, human-reviewed process is standard, and data preparation consumes roughly 80% of clinical fine-tuning work.
Why does generic SFT fail in expert domains?
Because the error baselines are catastrophic, the tolerances are near-zero, and the fix demonstrably runs through expert-built training data.
The baselines deserve to be stated plainly. Stanford RegLab and HAI researchers found general LLMs hallucinating on 69–88% of specific legal queries, with at least a 75% error rate on questions about court holdings — and a follow-up Stanford study of purpose-built legal research tools still measured 17% hallucination for Lexis+ AI, 33% for Westlaw's AI-assisted research and 43% for GPT-4, counting both fabricated cases and subtler mischaracterisations of real ones. Medicine looks no better unassisted: one clinical documentation analysis recorded hallucinations in 64.1% of medical case summaries without mitigation, against a clinical-deployment target practitioners set at under 2%. Finance adds a regulator: the SEC has already charged advisers over false AI claims, and a wrong number in a filing summary is not a UX problem.
Supervised fine-tuning (SFT) is the process of further training a model on curated instruction–response pairs so it learns to answer the way a qualified reviewer would. Fine-tuning on domain data is the documented countermeasure to the baselines above. A 2026 clinical study of a fine-tuned on-device model recorded major hallucinations falling 58.8% on a public benchmark and 84.8% on an internal evaluation set, with factual-correctness scores rising sharply; in law, work on LegalHalBench shows supervised fine-tuning plus preference optimisation materially improving non-hallucinated statute rates and legal-claim truthfulness. The market has drawn the conclusion already — Gartner projects more than half of enterprise GenAI models will be domain-specific by 2027, against roughly 1% in 2024, a shift covered in more detail in this guide to what enterprise teams buy in RLHF, SFT and distillation.
What decides whether that fine-tune helps or hurts is the training data itself, and here the research finding that reframed the field applies with full force: annotation quality matters far more than quantity, with a small set of carefully curated examples matching or exceeding models trained on orders of magnitude more data. In domains where the annotator must know the law, the medicine or the accounting standard to label correctly, "carefully curated" has a precise meaning: built by experts, verified by experts.
Baseline hallucination rates before domain fine-tuning:
| Source | Hallucination rate |
|---|---|
| General LLMs on specific legal queries (Stanford RegLab/HAI) | 69–88% |
| Unmitigated medical case summaries | 64.1% |
| GPT-4 on legal research tasks | 43% |
| Purpose-built legal AI tools (Lexis+ AI, Westlaw) | 17–33% |
| Clinical deployment target after mitigation | under 2% |
What does expert SFT data actually look like?
Instruction–response demonstrations a professional would sign: grounded, cited, terminologically exact, with each domain adding its own non-negotiables.
The common core. An SFT record is an instruction paired with the response the model should learn to give, and in expert domains that response must be one a qualified professional would stand behind: claims grounded in the authoritative source (the statute, the guideline, the filing), citations that resolve, refusals and hedges where a professional would refuse or hedge, and terminology used with the precision the field demands. The reference dataset scales are instructive: canonical evaluation sets run in the thousands, not millions — MedQA at 12,873 items, LegalBench at 8,452, FinanceBench at 10,150 — consistent with the quality-over-quantity finding that shapes modern SFT.
Legal: jurisdiction and citation discipline. Legal demonstrations must pin every claim to the right authority in the right jurisdiction — the exact failure surface the Stanford studies mapped — and handle hierarchy: what a court held versus said, what binds versus persuades. Terminology work here is expert work by definition: the TermGPT project's regulatory corpus was annotated by a panel of experts in financial law, with cross-validation and consensus-based resolution protocols, because a term like "material" carries meanings a lay annotator cannot adjudicate.
Medical: safety semantics at physician grain. Clinical demonstrations encode distinctions where errors harm people — contraindication versus caution, symptom versus diagnosis — and the scale of expert involvement in the field's reference artefacts shows what credible looks like: HealthBench, the canonical 2026 medical evaluation, rests on 48,562 rubric criteria written by 262 physicians across 26 specialties and 60 countries. Training data for clinical use is held to the same authorship standard, with safety-critical responses reviewed by clinicians rather than generalists.
Financial: numbers, standards and jurisdictional regimes. Financial demonstrations live and die on quantitative fidelity — figures that reconcile, calculations that follow the named standard, terminology aligned to the applicable regulatory regime — and fine-tuned financial models consistently outperform general ones on tasks like event and causality extraction precisely because the domain's language is this constrained. Across all three verticals one more dimension is chronically under-served: language. Law, medicine and finance exist in every jurisdiction and language, non-English legal benchmarks show even frontier models failing, and an SFT corpus built only in English trains a model for a fraction of the world it will be asked about — the same gap covered in this piece on multilingual LLM training data quality.
How is it produced — experts, process, and privacy?
Credentialed people, dual review, explicit rubrics, and a de-identification pipeline that treats privacy as part of the dataset rather than paperwork around it.
Recruit for verifiable expertise, then equip it. Published expert-evaluation projects show the operational template: credential requirements checked (health or medical degrees in one representative study), detailed written guidelines, tutorials before work begins, informed consent, and structured batches with multiple experts per item. Naive aggregation is the documented weakness — most pipelines majority-vote annotations and discard who labelled what, while recent research argues annotator identity and expertise should weight the outcome. In practice that means routing hard items to the most qualified reviewers and resolving disagreements by consensus discussion, not arithmetic — the same recruitment and certification discipline described in how annotators are recruited, trained and certified for specialist domains.
Independent second review is the quality mechanism. The pattern that recurs across every credible dataset in this space — expert panels with cross-validation, consensus resolution, physician-written rubrics — is the same dual-layer discipline Lifewood applies across its AI data work: one qualified pass produces the demonstration, an independent pass verifies it against the source and the rubric, and the decision is recorded. Lifewood's Bangladesh workforce alone logged 414,120 training hours in 2025, the kind of sustained calibration this category depends on. It is also where the operational reality of this category lives — sourcing credentialed annotators is a supply-chain problem, and Lifewood runs it through 40+ delivery centres across 30+ countries with coverage in 50+ languages, which is precisely what the multilingual gap in legal and medical corpora demands; its healthcare-document work runs at the company's strictest review thresholds for exactly the reasons this article describes, governed by the kind of enterprise annotation security and compliance controls that regulated data requires.
De-identification is the process of stripping direct and indirect identifiers from source records before they are used for training. Privacy is a pipeline stage with a measured failure rate: clinical source data must be collected under IRB approval or a quality-improvement exemption, then de-identified to HIPAA's standards — removing 18 categories of identifiers — and the sobering operational fact is that automated de-identification tools achieve only 95–98% recall, which is why practitioner guidance prescribes a two-pass approach (rule-based plus transformer NER) with human review to catch residual PHI before anything trains. Legal and financial data carry parallel duties: privilege, confidentiality and MNPI screening. And the budget reality frames all of it: in clinical fine-tuning projects, data preparation accounts for roughly 80% of the work — against compute costs of $500–$5,000 for LoRA-scale runs on 1,000–5,000 examples — so the expert data is not a line item in the project; it substantially is the project.
The expert SFT pipeline, in order:
- Recruit and equip — credential-verified experts with written rubrics, tutorials, consent and calibration batches.
- Author to rubric — grounded, cited demonstrations with explicit failure modes, refusals and hedges included.
- Dual review and consensus — independent expert verification, disagreements resolved by discussion, expertise weighted, decisions recorded.
- De-identify and document — two-pass PHI/PII removal with human review; provenance, reviewer identity and consent travel with every record.
How do you verify the fine-tune worked?
Expert-graded evaluation against explicit numeric targets, run on your own golden set, feeding a loop that turns production review into the next training round.
Set numeric safety targets and grade them with experts. Clinical practice shows the shape: have clinicians flag hallucinations across 500+ model outputs, target under 2% for deployment, and track harmful-recommendation rate as its own metric, with automatic rollback if a new model version exceeds the established baseline. Public domain benchmarks (MedQA, LegalBench, FinanceBench and their successors) are useful smoke tests, but they inherit every caveat of public benchmarks — saturation, contamination, and distance from your workload — so the deciding evaluation runs on a golden set built from your own cases, graded by the same calibre of experts who built the training data. Some of these same evaluation sets also double as red-teaming material; see this walkthrough of building safety and jailbreak datasets for LLM red-teaming for the adjacent discipline.
Close the loop. The audit trail regulated deployments already require — every inference logged with input, output, model version and the professional's action — doubles as the best data-collection instrument the programme will ever have: each expert correction in production is a candidate SFT example for the next round. Mature programmes treat evaluation and data production as one continuous cycle, which is also the direction the measurement field is heading — away from one-off, exam-style benchmarks and toward continuous, supervised evaluation inside real workflows, the way junior doctors and lawyers have always been assessed. Buyers comparing vendors for this stage often start from a broader list of LLM training data companies before narrowing to firms with genuine domain-expert bench strength, and pair that evaluation work with independent AI data validation and enterprise LLM training data services.
A caution on the numbers. The hallucination rates cited are study-specific — different models, prompts, and time windows produce different figures, and model generations move quickly; the cost figures are practitioner estimates for typical project shapes; and benchmark statistics describe those datasets as published. Treat all of them as directional, verify against the original studies, and treat nothing here as legal or medical advice.