Short answer. Horizontal LLM training data builds general capability — broad coverage across many domains, languages and task types, sourced at scale, judged on breadth and consistency. Vertical LLM training data builds domain competence — narrow, deep, expert-produced material for one field, judged on correctness by someone qualified to know. Enterprises almost always need both, and the mistake that costs the most is buying horizontal volume when the model's failure is vertical: a model that sounds fluent and gets your industry's specifics wrong will not be fixed by more general data, at any volume.
Two products, one label. "LLM training data" covers a corpus assembled from broad general material and a corpus written by qualified specialists in a single field, and the two differ in how they are sourced, who produces them, what they cost per item, and how their quality is measured.
This guide separates them and gives a diagnostic for deciding which your model actually needs.
The distinction in one table
| Horizontal | Vertical | |
|---|---|---|
| Goal | General capability, breadth, robustness | Domain competence and correctness |
| Coverage | Many domains, languages, task types | One field, in depth |
| Producer | Generalist annotators and writers, at scale | Qualified domain specialists |
| Volume | High | Low relative to horizontal |
| Cost per item | Lower | Substantially higher |
| Quality measured by | Consistency, coverage against a stratification plan, chance-corrected agreement | Expert review; factual correctness; currency of practice |
| Main risk | Skew toward whatever was easiest to source | Too narrow; a model expert in one sub-area and weak beside it |
| Fixes | Fluency, format, general reasoning behaviour, language coverage | Terminology, procedure, edge cases, regulatory specifics |
The producer row is the one that drives everything else. Generalist production scales; specialist production does not, and its cost is set by the market rate for the expertise, not by annotation rates.
When you need horizontal data
- The model is generally weak — poor instruction following, inconsistent formatting, weak reasoning across all topics rather than in one.
- Language coverage is the gap. A model that works in English and degrades in your other markets needs breadth, in those languages, produced natively.
- You are training or substantially adapting a base model rather than fine-tuning a strong one.
- Robustness is the complaint — the model handles the expected phrasing and falls over on the awkward variants.
Quality here is about coverage design and consistency, not depth. The recurring failure is a corpus that is large and skewed: over-representing whatever was easy to source, which is typically news, encyclopaedia and forum text, and under-representing the long tail of domains and the languages with the least available material. Stratify deliberately and report the minimum coverage per stratum, not the mean.
When you need vertical data
- Fluent and wrong. The model produces confident, well-formed output containing errors a practitioner would catch instantly. This is the diagnostic signature of a vertical gap, and it is not improved by general data.
- Terminology drift — using a term correctly in the general sense and incorrectly in your field's sense.
- Procedural gaps — knowing what a process is called and not the order or the exceptions.
- Regulatory and currency problems — stating a rule that was true, or that is true in a different jurisdiction.
- Edge cases that only a practitioner recognises as significant.
Quality here is correctness judged by someone qualified, which changes the sourcing problem entirely. Three requirements:
- Verified expertise. Qualification checked, not self-declared. The demonstration is a ceiling: a writer who cannot produce expert-quality output produces data that teaches the model to be a non-expert.
- Currency. Fields move. A corpus reflecting practice from five years ago will train a model to be confidently out of date, and there should be a review cadence for time-sensitive material.
- Jurisdiction and market specificity. Regulated fields differ by market. "Finance" is not one domain; it is one domain per regulatory regime.
How to diagnose which you need
Run a structured error review before buying anything. Sample real failures, and classify each:
| Symptom | Likely gap |
|---|---|
| Wrong format, ignored instruction, rambling | Horizontal — behaviour |
| Correct in English, poor in another language | Horizontal — language coverage |
| Reasoning breaks on multi-step problems generally | Horizontal — capability |
| Fluent output, domain-specific factual errors | Vertical |
| Right general term, wrong field-specific meaning | Vertical |
| Correct for one jurisdiction, wrong for yours | Vertical — market specificity |
| Fails only on rare inputs, fine otherwise | Either — check coverage before buying depth |
Vertical share of errors = Domain-specific errors ÷ Total errors reviewed
If that ratio is high, general volume will not move it. If it is low, expensive specialist writing is the wrong purchase. The review costs a few days and routinely redirects a budget by an order of magnitude.
How the two combine
They are sequential more often than simultaneous.
- Establish behaviour horizontally. Format, instruction following and language coverage first, because vertical data cannot fix a model that will not follow instructions.
- Add vertical depth narrowly, targeted at the specific error classes the review found.
- Build evaluation sets for both, separately. A general benchmark will not detect a domain regression, and a domain benchmark will not detect a general one.
- Re-test for interference. Heavy vertical fine-tuning can degrade general capability; a model that became excellent at your field and worse at everything else is a known and avoidable outcome, and only a general evaluation set will reveal it.
A useful budgeting heuristic: vertical data buys correctness where you are judged, horizontal data buys the competence that makes the model usable at all. Programmes that skip the horizontal layer produce a model that is expert and unusable; programmes that skip the vertical layer produce one that is pleasant and wrong.
What to ask a supplier
- Which do you actually produce in-house — general-scale production, specialist production, or both?
- How is domain expertise verified for vertical work?
- What is your stratification plan for horizontal coverage, and what is the minimum coverage per stratum?
- How do you keep vertical content current, and at what cadence?
- For multilingual work, is vertical content produced in-market or translated?
- How is quality measured differently for the two — and can I see both figures?
- Who owns the corpus, and can we export it in full?
Red flags: one blended price for both, which usually means specialist work is being produced by generalists; expertise described as self-declared; a horizontal corpus with no stratification plan; vertical content in non-English markets produced by translating English source material, which imports the wrong jurisdiction along with the language.
How Lifewood approaches this
Lifewood delivers both as distinct services rather than one blended offering, because they are different production problems: horizontal LLM data is a scale-and-coverage operation, vertical LLM data is a specialist-sourcing operation, and pricing them identically means one of them is being produced by the wrong people.
The multilingual dimension cuts across both and is where the delivery model matters most: 50+ languages, 40+ delivery centres across 30+ countries and 56,788 contributors mean both breadth and depth can be produced in-market rather than translated — which for vertical content is not a quality preference but a correctness requirement, since a translated domain corpus imports the source market's regulatory assumptions. Broader LLM scope includes RLHF, SFT, data distillation and response evaluation. The AI-data heritage runs to 2004, with the current company established in 2018.
See type B horizontal LLM data, type C vertical LLM data, enterprise LLM training data and type A data servicing.
Sources and further reading
- Companion guides: RLHF, SFT and Distillation and Multilingual LLM Training Data.
- Lifewood LLM data scope is published at lifewood.com/enterprise-llm-training-data.

