Short answer. Choose a generative model for production by building a small evaluation set from your own work and scoring it blind, because public leaderboards measure general capability on tasks that are not yours. Fifty to a hundred items, weighted towards hard cases and scored by two reviewers against a written rubric, separate candidates more reliably than any ranking. Then decide on the criteria that actually constrain production: output rights, data handling, latency, cost at real volume, language coverage and version stability.
Key takeaways
- Public leaderboards are a useful shortlist filter but a poor final decision tool, because they measure general capability on standard tasks, can be inflated by benchmark contamination, average away subtask variance and say nothing about rights, data handling or cost.
- An evaluation set of fifty to a hundred items drawn from real briefs, over-sampling difficult cases, scored blind by two independent reviewers against a rubric written before any output was seen, is the most reliable way to separate serious candidates.
- Rankings reorder between languages: the MMLU-ProX benchmark reports gaps of up to 24.3 percent between high- and low-resource languages on identical questions across 36 models, so a model chosen on English performance can be the wrong choice for other markets.
- Output rights and data handling eliminate a leading candidate more often than output quality does, and neither appears in any quality comparison, so check them before spending a week scoring.
- Keep the evaluation set frozen and re-run it on every model version change, keep the pipeline model-agnostic, and record the decision, its date and the decisive criterion so the choice can be explained later.
What can public benchmarks tell you, and what can't they?
Public benchmarks reliably tell you which models are serious candidates and which are not, but they cannot tell you which serious candidate is best for your specific production task. Leaderboards answer a real question, how these systems compare on a standard set of tasks, and it is usually not the question a production team has.
This is a procurement question, not a testing question. How to evaluate a system before you deploy it, including thresholds, gates and what to measure once it is live, is covered in the guide to designing enterprise evaluation benchmarks for AI systems. What follows is narrower: how to pick which generative model runs in your production pipeline, and how to be able to explain the choice afterwards.
A public benchmark is a fixed, openly published set of test items and a scoring method that ranks models on general capability rather than on any one organisation's production task. The gap between a benchmark score and production performance has several specific causes, and each suggests a different correction.
| Limitation | What it does to your decision |
|---|---|
| Task mismatch | Benchmarks measure reasoning, knowledge and general instruction-following; production tasks are narrow: rewrite this in our tone, describe this product accurately, generate a shot matching this reference |
| Contamination | Benchmark items circulate publicly and may appear in training data, inflating scores in ways that do not transfer |
| Aggregation | A single score averages across subtasks, so a model can lead overall while being clearly worse at the one thing you need |
| Thin language coverage | Most headline benchmarks are English-only, and rankings genuinely reorder by language |
| No operational criteria | No leaderboard scores output rights, data retention, regional availability or version stability |
Benchmark contamination is the leakage of published test items into a model's training data, which inflates its score without improving its real capability. It is one reason a leaderboard position should be treated as an upper bound rather than a prediction.
The language point is the one most often assumed rather than tested. MMLU-ProX, published at EMNLP 2025 and also available as arXiv preprint 2503.10497, poses identical items across 29 languages and reports gaps of up to 24.3 percent between high- and low-resource languages across 36 evaluated models. A model that wins in English can lose badly in your third market, which is why a serious selection needs multilingual evaluation sets rather than a translated English set.
The constructive use of public benchmarks is as a shortlist filter: they reliably separate serious candidates from unserious ones. Holistic frameworks such as Stanford CRFM's HELM are more useful at this stage than single-number rankings, because HELM reports seven metrics (accuracy, calibration, robustness, fairness, bias, toxicity and efficiency) across 16 core scenarios rather than collapsing to one figure. Choosing among the serious candidates is your own work.
How do you build an evaluation set from your own work?
Build an evaluation set by collecting fifty to a hundred real inputs from actual briefs, weighting them towards hard cases, writing a scoring rubric before seeing any output, running every candidate on the identical set, and scoring blind with two independent reviewers. This takes about a week of one person's time and is repeatable at every model change, which is what makes it worth building rather than improvising.
An evaluation set is a frozen collection of representative production inputs, paired with a written rubric, that every candidate model is run against under identical conditions. The steps below are the ones that make the result trustworthy.
1. Collect items from real work, weighted towards the hard cases. Fifty to a hundred inputs drawn from actual briefs. Deliberately over-sample the difficult ones: unusual products, sensitive topics, your worst-formatted source material, your smallest markets. A set of typical items will show every serious candidate performing well and tell you nothing.
2. Write the rubric before seeing any output. The dimensions that matter for the task (accuracy, tone, instruction adherence, format compliance, brand fit) with severity levels and worked examples. A rubric written after seeing outputs describes the first model you happened to look at.
3. Run every candidate on the identical set. Same inputs, same parameters where comparable, same number of attempts. Where a model needs different prompting to perform well, that is a real finding about integration cost. Record it rather than quietly tuning one candidate harder than the others.
4. Score blind. Strip model identity, randomise order, and have two reviewers score independently. Blind scoring is a review method in which reviewers rate outputs without knowing which model produced them, so that the score measures the output rather than expectations about the brand.
5. Check reviewer agreement before trusting the result. If two reviewers disagree substantially, the rubric is ambiguous and the comparison is not yet meaningful. Fix the rubric and re-score. Chance-corrected agreement in the sense of Landis and Koch (Biometrics, 1977) is the standard way to report this; the guide to inter-annotator agreement statistics explains what the kappa and alpha numbers mean in practice.
6. Evaluate in every language you operate in. Not a translated English set: items authored in each target language. Rankings reorder between languages, and a selection made on English performance can be the wrong choice for most of your markets.
7. Test the failure modes, not only the successes. How does each candidate behave on out-of-scope requests, ambiguous briefs, and inputs designed to elicit a confident wrong answer? A model that fails loudly is far easier to operate than one that fails plausibly.
8. Freeze the set and re-run on every version change. This is what converts "the model feels different this month" into a measured difference. Keep the set out of any training or fine-tuning use so it stays a clean reference.
Separation = Mean rubric score of leading candidate − Mean rubric score of runner-up
If that difference is smaller than the disagreement between your two reviewers, the candidates are not separated on quality and the decision belongs to the operational criteria.
Which criteria actually decide the selection?
Once a shortlist sits within a reasonable band on output quality, the selection is decided by operational criteria: output rights, data handling, regional availability, latency and throughput, cost at real volume, version stability, language coverage and provenance support. Shortlists sit closer on quality than vendor comparisons imply, so these criteria settle most real selections.
| Criterion | The question to ask | Why it decides cases |
|---|---|---|
| Output rights | What may we do with the outputs, and does the provider claim anything? | A quality advantage is worthless if the licence does not permit your use |
| Data handling | Are inputs retained, used for training, or logged? Where, and for how long? | Client confidentiality obligations routinely eliminate otherwise strong candidates |
| Regional availability and residency | Can we call it from, and store data in, the regions we operate in? | A model unavailable in a market is not a candidate for that market |
| Latency and throughput | Per-item latency at our concurrency, and the real rate limits | Batch production is throughput-bound; interactive tools are latency-bound, and these select differently |
| Cost at real volume | Cost per finished deliverable including retries and rejected takes | Per-token or per-image pricing understates the cost of takes you discard |
| Version stability | How are versions pinned, deprecated and announced? | A model that changes under a production pipeline without notice is an operational risk regardless of quality |
| Language coverage | Measured performance in your actual markets | The criterion most often assumed rather than tested |
| Provenance support | Does it mark outputs, in a way that survives your pipeline? | EU AI Act Article 50 transparency duties make this a compliance input rather than a nice-to-have |
Version stability is a provider's documented practice of pinning, deprecating and announcing model versions, so that a production pipeline does not change behaviour without notice. It is the criterion most often discovered the hard way, usually when outputs drift and nobody can say what changed.
Output rights and data handling are the two that most often eliminate a leading candidate outright, and they are the two least visible in a quality comparison. Check them before you spend a week scoring; the explainer on what happens to your data at a generative AI vendor sets out the retention and training questions to put to each provider.
Provenance has moved from a preference to a compliance input. Article 50 of the EU AI Act requires providers of generative AI systems to ensure that synthetic audio, image, video and text outputs are marked in a machine-readable format and detectable as artificially generated, with those obligations applying from 2 August 2026. Whether a model's marking survives your editing and delivery pipeline is a practical question, and the comparison of C2PA, SynthID and what survives covers what each approach does when a file is transcoded or cropped.
How should you document a model selection decision?
Document a model selection by versioning the evaluation set and rubric, recording the scores per candidate and per language with reviewer agreement, listing how each candidate was assessed against the non-quality criteria, and dating the decision with the criterion that settled it. Model selection is increasingly something an organisation is expected to be able to explain rather than merely to have done.
The NIST AI Risk Management Framework (AI RMF 1.0, released January 2023 and intended for voluntary use) is widely referenced as a baseline, and its core functions, Govern, Map, Measure and Manage, treat documentation, measurement and traceability as ordinary practice. In practice, a client or auditor asking why a particular system was chosen is asking for exactly the artefacts a good evaluation produces anyway:
- The evaluation set and rubric, versioned.
- The scores, per candidate and per language, with reviewer agreement recorded.
- The non-quality criteria and how each candidate was assessed against them.
- The decision and its date, including which criterion was decisive.
- The re-evaluation record at each model version change.
Keeping these costs almost nothing at the time and is unreconstructable later. That is the recurring theme of running AI in production: the work is usually fine, and the evidence that the work was done is what goes missing.
Which two patterns make a model choice easier to live with?
Two patterns make a model selection easier to live with: do not standardise on a single model, and keep the pipeline separate from the model so that swapping one is a configuration change rather than a rebuild. Both reduce the cost of being wrong, which matters in a field where the best candidate changes every few months.
Do not standardise on one model. Different tasks in the same pipeline are genuinely won by different systems, and the integration cost of supporting two or three behind a common interface is modest against the quality cost of forcing one everywhere. It also removes a single point of failure when a provider changes terms, pricing or availability.
Separate the model from the pipeline. Prompts, references, rubrics, review workflow, provenance and delivery should be model-agnostic, so that swapping a model is a configuration change rather than a rebuild. Given how quickly this field moves, a pipeline welded to one provider's specifics is a pipeline that will be rewritten within the year. Treating prompts as enterprise assets, versioned and owned independently of any one provider, is the practical form this takes.
How does Lifewood choose generative models for production?
Lifewood operates third-party generative models rather than training its own, so model evaluation is part of ordinary operations rather than a one-off procurement event. The evaluation set is drawn from real briefs, frozen, re-run on version changes and run per language, with blind, dual-reviewer scoring.
The language dimension is where rankings actually move, and because Lifewood delivers in 100+ languages across 40+ delivery centres in 30+ countries, an English-only selection would be wrong for most of the work. Scoring is blind and dual-reviewer, feeding the same 95%+ accuracy SLA applied across delivery programmes, with two independent review passes and timestamped approval records.
This evaluation discipline sits inside the managed AIGC services and AI video production offer, where model choice is one of the things a client is paying not to have to make alone. Buyers weighing providers on this basis can see how the field lines up in the comparison of AIGC video production providers.