Skip to main content
AI Data

Domain-Expert Data: Building Legal, Medical and Financial SFT Datasets

September 2026 · 10 min read · Updated September 2026

Short answer. With credentialed professionals in the loop at every step: experts author and verify the instruction–response demonstrations, work from rubrics with explicit failure modes, resolve disagreements by consensus, and every record is de-identified and documented before training. The stakes justify the cost — general models hallucinate on 69–88% of legal queries and 64% of unmitigated medical summaries — and the research is clear that a small, carefully curated expert set beats orders of magnitude more noisy data.

Key takeaways

  • General LLMs hallucinate on 69–88% of legal queries (75%+ on court holdings), purpose-built legal tools still err 17–43% of the time, and unmitigated medical summaries hallucinate at 64.1% — against a under-2% clinical deployment target.
  • Domain fine-tuning demonstrably closes the gap: major clinical hallucinations fell 58.8–84.8% in one 2026 study, statute-grounding metrics rose in legal work, and Gartner projects domain-specific models going from 1% to over half of enterprise GenAI by 2027.
  • Quality beats quantity: small, carefully curated expert sets match models trained on orders of magnitude more data, and canonical domain evaluation sets run in the thousands of items, not millions.
  • Expert SFT records are demonstrations a professional would sign: grounded in the authoritative source, correctly cited, terminologically exact, with refusals where a professional would refuse.
  • Privacy is a measured pipeline stage — automated de-identification recalls only 95–98% of identifiers, so a two-pass, human-reviewed process is standard, and data preparation consumes roughly 80% of clinical fine-tuning work.

Why does generic SFT fail in expert domains?

Because the error baselines are catastrophic, the tolerances are near-zero, and the fix demonstrably runs through expert-built training data.

The baselines deserve to be stated plainly. Stanford RegLab and HAI researchers found general LLMs hallucinating on 69–88% of specific legal queries, with at least a 75% error rate on questions about court holdings — and a follow-up Stanford study of purpose-built legal research tools still measured 17% hallucination for Lexis+ AI, 33% for Westlaw's AI-assisted research and 43% for GPT-4, counting both fabricated cases and subtler mischaracterisations of real ones. Medicine looks no better unassisted: one clinical documentation analysis recorded hallucinations in 64.1% of medical case summaries without mitigation, against a clinical-deployment target practitioners set at under 2%. Finance adds a regulator: the SEC has already charged advisers over false AI claims, and a wrong number in a filing summary is not a UX problem.

Supervised fine-tuning (SFT) is the process of further training a model on curated instruction–response pairs so it learns to answer the way a qualified reviewer would. Fine-tuning on domain data is the documented countermeasure to the baselines above. A 2026 clinical study of a fine-tuned on-device model recorded major hallucinations falling 58.8% on a public benchmark and 84.8% on an internal evaluation set, with factual-correctness scores rising sharply; in law, work on LegalHalBench shows supervised fine-tuning plus preference optimisation materially improving non-hallucinated statute rates and legal-claim truthfulness. The market has drawn the conclusion already — Gartner projects more than half of enterprise GenAI models will be domain-specific by 2027, against roughly 1% in 2024, a shift covered in more detail in this guide to what enterprise teams buy in RLHF, SFT and distillation.

What decides whether that fine-tune helps or hurts is the training data itself, and here the research finding that reframed the field applies with full force: annotation quality matters far more than quantity, with a small set of carefully curated examples matching or exceeding models trained on orders of magnitude more data. In domains where the annotator must know the law, the medicine or the accounting standard to label correctly, "carefully curated" has a precise meaning: built by experts, verified by experts.

Baseline hallucination rates before domain fine-tuning:

Source Hallucination rate
General LLMs on specific legal queries (Stanford RegLab/HAI) 69–88%
Unmitigated medical case summaries 64.1%
GPT-4 on legal research tasks 43%
Purpose-built legal AI tools (Lexis+ AI, Westlaw) 17–33%
Clinical deployment target after mitigation under 2%

What does expert SFT data actually look like?

Instruction–response demonstrations a professional would sign: grounded, cited, terminologically exact, with each domain adding its own non-negotiables.

The common core. An SFT record is an instruction paired with the response the model should learn to give, and in expert domains that response must be one a qualified professional would stand behind: claims grounded in the authoritative source (the statute, the guideline, the filing), citations that resolve, refusals and hedges where a professional would refuse or hedge, and terminology used with the precision the field demands. The reference dataset scales are instructive: canonical evaluation sets run in the thousands, not millions — MedQA at 12,873 items, LegalBench at 8,452, FinanceBench at 10,150 — consistent with the quality-over-quantity finding that shapes modern SFT.

Legal: jurisdiction and citation discipline. Legal demonstrations must pin every claim to the right authority in the right jurisdiction — the exact failure surface the Stanford studies mapped — and handle hierarchy: what a court held versus said, what binds versus persuades. Terminology work here is expert work by definition: the TermGPT project's regulatory corpus was annotated by a panel of experts in financial law, with cross-validation and consensus-based resolution protocols, because a term like "material" carries meanings a lay annotator cannot adjudicate.

Medical: safety semantics at physician grain. Clinical demonstrations encode distinctions where errors harm people — contraindication versus caution, symptom versus diagnosis — and the scale of expert involvement in the field's reference artefacts shows what credible looks like: HealthBench, the canonical 2026 medical evaluation, rests on 48,562 rubric criteria written by 262 physicians across 26 specialties and 60 countries. Training data for clinical use is held to the same authorship standard, with safety-critical responses reviewed by clinicians rather than generalists.

Financial: numbers, standards and jurisdictional regimes. Financial demonstrations live and die on quantitative fidelity — figures that reconcile, calculations that follow the named standard, terminology aligned to the applicable regulatory regime — and fine-tuned financial models consistently outperform general ones on tasks like event and causality extraction precisely because the domain's language is this constrained. Across all three verticals one more dimension is chronically under-served: language. Law, medicine and finance exist in every jurisdiction and language, non-English legal benchmarks show even frontier models failing, and an SFT corpus built only in English trains a model for a fraction of the world it will be asked about — the same gap covered in this piece on multilingual LLM training data quality.

How is it produced — experts, process, and privacy?

Credentialed people, dual review, explicit rubrics, and a de-identification pipeline that treats privacy as part of the dataset rather than paperwork around it.

Recruit for verifiable expertise, then equip it. Published expert-evaluation projects show the operational template: credential requirements checked (health or medical degrees in one representative study), detailed written guidelines, tutorials before work begins, informed consent, and structured batches with multiple experts per item. Naive aggregation is the documented weakness — most pipelines majority-vote annotations and discard who labelled what, while recent research argues annotator identity and expertise should weight the outcome. In practice that means routing hard items to the most qualified reviewers and resolving disagreements by consensus discussion, not arithmetic — the same recruitment and certification discipline described in how annotators are recruited, trained and certified for specialist domains.

Independent second review is the quality mechanism. The pattern that recurs across every credible dataset in this space — expert panels with cross-validation, consensus resolution, physician-written rubrics — is the same dual-layer discipline Lifewood applies across its AI data work: one qualified pass produces the demonstration, an independent pass verifies it against the source and the rubric, and the decision is recorded. Lifewood's Bangladesh workforce alone logged 414,120 training hours in 2025, the kind of sustained calibration this category depends on. It is also where the operational reality of this category lives — sourcing credentialed annotators is a supply-chain problem, and Lifewood runs it through 40+ delivery centres across 30+ countries with coverage in 50+ languages, which is precisely what the multilingual gap in legal and medical corpora demands; its healthcare-document work runs at the company's strictest review thresholds for exactly the reasons this article describes, governed by the kind of enterprise annotation security and compliance controls that regulated data requires.

De-identification is the process of stripping direct and indirect identifiers from source records before they are used for training. Privacy is a pipeline stage with a measured failure rate: clinical source data must be collected under IRB approval or a quality-improvement exemption, then de-identified to HIPAA's standards — removing 18 categories of identifiers — and the sobering operational fact is that automated de-identification tools achieve only 95–98% recall, which is why practitioner guidance prescribes a two-pass approach (rule-based plus transformer NER) with human review to catch residual PHI before anything trains. Legal and financial data carry parallel duties: privilege, confidentiality and MNPI screening. And the budget reality frames all of it: in clinical fine-tuning projects, data preparation accounts for roughly 80% of the work — against compute costs of $500–$5,000 for LoRA-scale runs on 1,000–5,000 examples — so the expert data is not a line item in the project; it substantially is the project.

The expert SFT pipeline, in order:

  1. Recruit and equip — credential-verified experts with written rubrics, tutorials, consent and calibration batches.
  2. Author to rubric — grounded, cited demonstrations with explicit failure modes, refusals and hedges included.
  3. Dual review and consensus — independent expert verification, disagreements resolved by discussion, expertise weighted, decisions recorded.
  4. De-identify and document — two-pass PHI/PII removal with human review; provenance, reviewer identity and consent travel with every record.

How do you verify the fine-tune worked?

Expert-graded evaluation against explicit numeric targets, run on your own golden set, feeding a loop that turns production review into the next training round.

Set numeric safety targets and grade them with experts. Clinical practice shows the shape: have clinicians flag hallucinations across 500+ model outputs, target under 2% for deployment, and track harmful-recommendation rate as its own metric, with automatic rollback if a new model version exceeds the established baseline. Public domain benchmarks (MedQA, LegalBench, FinanceBench and their successors) are useful smoke tests, but they inherit every caveat of public benchmarks — saturation, contamination, and distance from your workload — so the deciding evaluation runs on a golden set built from your own cases, graded by the same calibre of experts who built the training data. Some of these same evaluation sets also double as red-teaming material; see this walkthrough of building safety and jailbreak datasets for LLM red-teaming for the adjacent discipline.

Close the loop. The audit trail regulated deployments already require — every inference logged with input, output, model version and the professional's action — doubles as the best data-collection instrument the programme will ever have: each expert correction in production is a candidate SFT example for the next round. Mature programmes treat evaluation and data production as one continuous cycle, which is also the direction the measurement field is heading — away from one-off, exam-style benchmarks and toward continuous, supervised evaluation inside real workflows, the way junior doctors and lawyers have always been assessed. Buyers comparing vendors for this stage often start from a broader list of LLM training data companies before narrowing to firms with genuine domain-expert bench strength, and pair that evaluation work with independent AI data validation and enterprise LLM training data services.

A caution on the numbers. The hallucination rates cited are study-specific — different models, prompts, and time windows produce different figures, and model generations move quickly; the cost figures are practitioner estimates for typical project shapes; and benchmark statistics describe those datasets as published. Treat all of them as directional, verify against the original studies, and treat nothing here as legal or medical advice.

Frequently asked questions

Fewer than most teams assume: LoRA-scale clinical projects typically run on 1,000–5,000 well-curated examples, full fine-tunes on 10,000+, and the research consistently shows small expert-curated sets beating far larger noisy ones. Spend the budget on curation depth, not row count.

Not for the judgments that matter. Deciding whether a citation supports a holding, a dosage note is safe, or a term matches the regulatory definition requires the domain itself — guidelines help experts be consistent; they cannot substitute for expertise. Reserve generalists for structural tasks under expert review.

It can extend them — expert-seeded generation with expert verification is common — but unverified synthetic data in these domains launders model errors into training truth. The working pattern is silver synthetic drafts promoted to gold only after expert review.

Data preparation — roughly 80% of the work in clinical fine-tuning — and within it, de-identification and expert review time. Compute is the cheap part; credentialed attention is the constraint to plan around.

Treat each language-jurisdiction pair as its own data problem: laws, clinical guidelines and regulatory terms differ by country, not just by translation. That requires in-country domain experts, which is why non-English benchmarks embarrass even frontier models — this work has rarely been done at quality.

Sources and further reading

  1. Stanford Law School / RegLab & HAI, "Hallucinating Law: Legal Mistakes with Large Language Models are Pervasive", on the 69–88% legal hallucination rates
  2. AI Law Librarians, "What the Science Says About Hallucinations in Legal Research", reporting the Stanford follow-up rates for Lexis+ AI, Westlaw and GPT-4
  3. SQ Magazine, "LLM Hallucination Rate: 40+ Stats", compiling the 64.1% unmitigated medical-summary hallucination figure and related benchmarks
  4. Kili Technology, "Domain-Specific LLM Benchmarks: 2026 Vertical AI Map", on Gartner's domain-specific projection, HealthBench's physician-authored rubrics and LegalBench-RAG
  5. "REALM: Reliable Expertise-Aware Language Model Fine-Tuning from Noisy Annotations" (arXiv), on annotator-expertise weighting and the quality-over-quantity SFT literature
  6. "TermGPT: Multi-Level Contrastive Fine-Tuning for Terminology Adaptation in Legal and Financial Domain" (arXiv), on expert-panel annotation with cross-validation and consensus resolution
  7. "Why Supervised Fine-Tuning Fails to Learn" (arXiv), for the MedQA, LegalBench and FinanceBench dataset statistics
  8. Nirmitee, "Fine-Tuning AI for Patient Care", on clinical fine-tuning costs, the 80% data-preparation share, HIPAA de-identification recall rates, the under-2% deployment target and audit-trail practice
  9. "MedHalu: Hallucinations in Responses to Healthcare Queries" (arXiv), for the expert-evaluation logistics: credential requirements, guidelines, batching and consent
  10. "An On-Device AI Model for Medical Transcription and Note Generation" (medRxiv), on post-fine-tuning hallucination and omission reductions
  11. "Fine-tuning LLMs for Improving Factuality in Legal Question Answering" (LegalHalBench, arXiv), on SFT plus preference optimisation for statute grounding

Have an AI or visibility project in mind?

From AI evaluation and human-in-the-loop review to GEO and AEO strategy, our team can help you deploy with confidence and get found in the AI search era.

Talk to our team