Skip to main content
AIGC

How to Choose a Generative Model for Production

July 2026 · 11 min read · Updated September 2026

Short answer. Choose a generative model for production by building a small evaluation set from your own work and scoring it blind, because public leaderboards measure general capability on tasks that are not yours. Fifty to a hundred items, weighted towards hard cases and scored by two reviewers against a written rubric, separate candidates more reliably than any ranking. Then decide on the criteria that actually constrain production: output rights, data handling, latency, cost at real volume, language coverage and version stability.

Key takeaways

  • Public leaderboards are a useful shortlist filter but a poor final decision tool, because they measure general capability on standard tasks, can be inflated by benchmark contamination, average away subtask variance and say nothing about rights, data handling or cost.
  • An evaluation set of fifty to a hundred items drawn from real briefs, over-sampling difficult cases, scored blind by two independent reviewers against a rubric written before any output was seen, is the most reliable way to separate serious candidates.
  • Rankings reorder between languages: the MMLU-ProX benchmark reports gaps of up to 24.3 percent between high- and low-resource languages on identical questions across 36 models, so a model chosen on English performance can be the wrong choice for other markets.
  • Output rights and data handling eliminate a leading candidate more often than output quality does, and neither appears in any quality comparison, so check them before spending a week scoring.
  • Keep the evaluation set frozen and re-run it on every model version change, keep the pipeline model-agnostic, and record the decision, its date and the decisive criterion so the choice can be explained later.

What can public benchmarks tell you, and what can't they?

Public benchmarks reliably tell you which models are serious candidates and which are not, but they cannot tell you which serious candidate is best for your specific production task. Leaderboards answer a real question, how these systems compare on a standard set of tasks, and it is usually not the question a production team has.

This is a procurement question, not a testing question. How to evaluate a system before you deploy it, including thresholds, gates and what to measure once it is live, is covered in the guide to designing enterprise evaluation benchmarks for AI systems. What follows is narrower: how to pick which generative model runs in your production pipeline, and how to be able to explain the choice afterwards.

A public benchmark is a fixed, openly published set of test items and a scoring method that ranks models on general capability rather than on any one organisation's production task. The gap between a benchmark score and production performance has several specific causes, and each suggests a different correction.

Limitation What it does to your decision
Task mismatch Benchmarks measure reasoning, knowledge and general instruction-following; production tasks are narrow: rewrite this in our tone, describe this product accurately, generate a shot matching this reference
Contamination Benchmark items circulate publicly and may appear in training data, inflating scores in ways that do not transfer
Aggregation A single score averages across subtasks, so a model can lead overall while being clearly worse at the one thing you need
Thin language coverage Most headline benchmarks are English-only, and rankings genuinely reorder by language
No operational criteria No leaderboard scores output rights, data retention, regional availability or version stability

Benchmark contamination is the leakage of published test items into a model's training data, which inflates its score without improving its real capability. It is one reason a leaderboard position should be treated as an upper bound rather than a prediction.

The language point is the one most often assumed rather than tested. MMLU-ProX, published at EMNLP 2025 and also available as arXiv preprint 2503.10497, poses identical items across 29 languages and reports gaps of up to 24.3 percent between high- and low-resource languages across 36 evaluated models. A model that wins in English can lose badly in your third market, which is why a serious selection needs multilingual evaluation sets rather than a translated English set.

The constructive use of public benchmarks is as a shortlist filter: they reliably separate serious candidates from unserious ones. Holistic frameworks such as Stanford CRFM's HELM are more useful at this stage than single-number rankings, because HELM reports seven metrics (accuracy, calibration, robustness, fairness, bias, toxicity and efficiency) across 16 core scenarios rather than collapsing to one figure. Choosing among the serious candidates is your own work.

How do you build an evaluation set from your own work?

Build an evaluation set by collecting fifty to a hundred real inputs from actual briefs, weighting them towards hard cases, writing a scoring rubric before seeing any output, running every candidate on the identical set, and scoring blind with two independent reviewers. This takes about a week of one person's time and is repeatable at every model change, which is what makes it worth building rather than improvising.

An evaluation set is a frozen collection of representative production inputs, paired with a written rubric, that every candidate model is run against under identical conditions. The steps below are the ones that make the result trustworthy.

1. Collect items from real work, weighted towards the hard cases. Fifty to a hundred inputs drawn from actual briefs. Deliberately over-sample the difficult ones: unusual products, sensitive topics, your worst-formatted source material, your smallest markets. A set of typical items will show every serious candidate performing well and tell you nothing.

2. Write the rubric before seeing any output. The dimensions that matter for the task (accuracy, tone, instruction adherence, format compliance, brand fit) with severity levels and worked examples. A rubric written after seeing outputs describes the first model you happened to look at.

3. Run every candidate on the identical set. Same inputs, same parameters where comparable, same number of attempts. Where a model needs different prompting to perform well, that is a real finding about integration cost. Record it rather than quietly tuning one candidate harder than the others.

4. Score blind. Strip model identity, randomise order, and have two reviewers score independently. Blind scoring is a review method in which reviewers rate outputs without knowing which model produced them, so that the score measures the output rather than expectations about the brand.

5. Check reviewer agreement before trusting the result. If two reviewers disagree substantially, the rubric is ambiguous and the comparison is not yet meaningful. Fix the rubric and re-score. Chance-corrected agreement in the sense of Landis and Koch (Biometrics, 1977) is the standard way to report this; the guide to inter-annotator agreement statistics explains what the kappa and alpha numbers mean in practice.

6. Evaluate in every language you operate in. Not a translated English set: items authored in each target language. Rankings reorder between languages, and a selection made on English performance can be the wrong choice for most of your markets.

7. Test the failure modes, not only the successes. How does each candidate behave on out-of-scope requests, ambiguous briefs, and inputs designed to elicit a confident wrong answer? A model that fails loudly is far easier to operate than one that fails plausibly.

8. Freeze the set and re-run on every version change. This is what converts "the model feels different this month" into a measured difference. Keep the set out of any training or fine-tuning use so it stays a clean reference.

Separation = Mean rubric score of leading candidate − Mean rubric score of runner-up

If that difference is smaller than the disagreement between your two reviewers, the candidates are not separated on quality and the decision belongs to the operational criteria.

Which criteria actually decide the selection?

Once a shortlist sits within a reasonable band on output quality, the selection is decided by operational criteria: output rights, data handling, regional availability, latency and throughput, cost at real volume, version stability, language coverage and provenance support. Shortlists sit closer on quality than vendor comparisons imply, so these criteria settle most real selections.

Criterion The question to ask Why it decides cases
Output rights What may we do with the outputs, and does the provider claim anything? A quality advantage is worthless if the licence does not permit your use
Data handling Are inputs retained, used for training, or logged? Where, and for how long? Client confidentiality obligations routinely eliminate otherwise strong candidates
Regional availability and residency Can we call it from, and store data in, the regions we operate in? A model unavailable in a market is not a candidate for that market
Latency and throughput Per-item latency at our concurrency, and the real rate limits Batch production is throughput-bound; interactive tools are latency-bound, and these select differently
Cost at real volume Cost per finished deliverable including retries and rejected takes Per-token or per-image pricing understates the cost of takes you discard
Version stability How are versions pinned, deprecated and announced? A model that changes under a production pipeline without notice is an operational risk regardless of quality
Language coverage Measured performance in your actual markets The criterion most often assumed rather than tested
Provenance support Does it mark outputs, in a way that survives your pipeline? EU AI Act Article 50 transparency duties make this a compliance input rather than a nice-to-have

Version stability is a provider's documented practice of pinning, deprecating and announcing model versions, so that a production pipeline does not change behaviour without notice. It is the criterion most often discovered the hard way, usually when outputs drift and nobody can say what changed.

Output rights and data handling are the two that most often eliminate a leading candidate outright, and they are the two least visible in a quality comparison. Check them before you spend a week scoring; the explainer on what happens to your data at a generative AI vendor sets out the retention and training questions to put to each provider.

Provenance has moved from a preference to a compliance input. Article 50 of the EU AI Act requires providers of generative AI systems to ensure that synthetic audio, image, video and text outputs are marked in a machine-readable format and detectable as artificially generated, with those obligations applying from 2 August 2026. Whether a model's marking survives your editing and delivery pipeline is a practical question, and the comparison of C2PA, SynthID and what survives covers what each approach does when a file is transcoded or cropped.

How should you document a model selection decision?

Document a model selection by versioning the evaluation set and rubric, recording the scores per candidate and per language with reviewer agreement, listing how each candidate was assessed against the non-quality criteria, and dating the decision with the criterion that settled it. Model selection is increasingly something an organisation is expected to be able to explain rather than merely to have done.

The NIST AI Risk Management Framework (AI RMF 1.0, released January 2023 and intended for voluntary use) is widely referenced as a baseline, and its core functions, Govern, Map, Measure and Manage, treat documentation, measurement and traceability as ordinary practice. In practice, a client or auditor asking why a particular system was chosen is asking for exactly the artefacts a good evaluation produces anyway:

  • The evaluation set and rubric, versioned.
  • The scores, per candidate and per language, with reviewer agreement recorded.
  • The non-quality criteria and how each candidate was assessed against them.
  • The decision and its date, including which criterion was decisive.
  • The re-evaluation record at each model version change.

Keeping these costs almost nothing at the time and is unreconstructable later. That is the recurring theme of running AI in production: the work is usually fine, and the evidence that the work was done is what goes missing.

Which two patterns make a model choice easier to live with?

Two patterns make a model selection easier to live with: do not standardise on a single model, and keep the pipeline separate from the model so that swapping one is a configuration change rather than a rebuild. Both reduce the cost of being wrong, which matters in a field where the best candidate changes every few months.

Do not standardise on one model. Different tasks in the same pipeline are genuinely won by different systems, and the integration cost of supporting two or three behind a common interface is modest against the quality cost of forcing one everywhere. It also removes a single point of failure when a provider changes terms, pricing or availability.

Separate the model from the pipeline. Prompts, references, rubrics, review workflow, provenance and delivery should be model-agnostic, so that swapping a model is a configuration change rather than a rebuild. Given how quickly this field moves, a pipeline welded to one provider's specifics is a pipeline that will be rewritten within the year. Treating prompts as enterprise assets, versioned and owned independently of any one provider, is the practical form this takes.

How does Lifewood choose generative models for production?

Lifewood operates third-party generative models rather than training its own, so model evaluation is part of ordinary operations rather than a one-off procurement event. The evaluation set is drawn from real briefs, frozen, re-run on version changes and run per language, with blind, dual-reviewer scoring.

The language dimension is where rankings actually move, and because Lifewood delivers in 100+ languages across 40+ delivery centres in 30+ countries, an English-only selection would be wrong for most of the work. Scoring is blind and dual-reviewer, feeding the same 95%+ accuracy SLA applied across delivery programmes, with two independent review passes and timestamped approval records.

This evaluation discipline sits inside the managed AIGC services and AI video production offer, where model choice is one of the things a client is paying not to have to make alone. Buyers weighing providers on this basis can see how the field lines up in the comparison of AIGC video production providers.

Frequently asked questions

As a shortlist filter, yes: leaderboards reliably separate serious candidates from unserious ones. As a decision, no. Benchmarks measure general capability on tasks that are not yours, they can be affected by contamination, they aggregate away subtask variance, and they score nothing about rights, data handling, availability or cost at your volume.

Because unblinded scoring measures expectations about brands rather than output. Strip model identity, randomise the order, and use two independent reviewers against a rubric written before anyone saw a result. If the two reviewers disagree substantially, the rubric is ambiguous and the comparison is not yet meaningful.

Yes, with items authored in each language rather than translated from English. Rankings reorder between languages and the gap can be large: MMLU-ProX reports disparities of up to 24.3 percent between high- and low-resource languages on identical questions across 36 models. A selection made on English performance can be actively wrong for most of your markets.

At every model version change, and on a periodic cadence regardless; quarterly is a reasonable default. Providers update models and behaviour moves in both directions. A frozen evaluation set turns "it feels different" into a measurement, which is the only sound basis for a pipeline change.

Output rights and data handling, in most real selections. If the licence does not permit your use, or the provider retains and trains on your inputs in a way your client contracts prohibit, quality is irrelevant. After those: regional availability, cost at genuine volume including discarded takes, and version stability.

Lifewood Data Technology produces AI video and content as a managed service with human review built in. Third-party generative models are evaluated blind by two reviewers on a frozen set drawn from real briefs, and delivery runs against a 95%+ accuracy SLA with two independent review passes, across 50+ languages and 40+ delivery centres in 30+ countries.

Sources and further reading

  1. MMLU-ProX: A Multilingual Benchmark for Advanced Large Language Model Evaluation (EMNLP 2025)
  2. MMLU-ProX, arXiv preprint 2503.10497
  3. Holistic Evaluation of Language Models (HELM), Stanford CRFM, arXiv 2211.09110
  4. NIST AI Risk Management Framework (AI RMF 1.0)
  5. Landis and Koch, The Measurement of Observer Agreement for Categorical Data, Biometrics 33(1), 1977
  6. EU AI Act, Article 50: Transparency Obligations for Providers and Deployers of Certain AI Systems

Have an AI or visibility project in mind?

From AI evaluation and human-in-the-loop review to GEO and AEO strategy, our team can help you deploy with confidence and get found in the AI search era.

Talk to our team