Short answer. A brand book is written for people who will interpret it. A generative pipeline cannot interpret; it needs the brand expressed as artefacts it can be conditioned on — an approved fact pack, a termbase with do-not-translate entries, positive and negative examples, a claims matrix per market, and a reviewer rubric that scores against those artefacts rather than against taste. Do not expect a model to recall your brand correctly from general knowledge: supply the approved material every time. The work is one-time per brand and per language, and it is the difference between a pipeline that scales and one that produces a fresh argument for every asset.
This guide covers the verbal and factual side of brand control — terminology, claims, tone and localisation. The visual side, where the mechanism is reference conditioning and compositing rather than glossaries, is a different problem covered separately.
Why a brand book does not survive contact with a generative pipeline
Traditional guidelines describe logos, colours, typography, tone and messaging in prose, and rely on a designer or writer to apply judgement. That works because the reader has context. A model has no context beyond what is in the prompt, and it will confidently produce a plausible version of your brand assembled from everything it has seen about companies like yours.
The failure modes are specific and repetitive:
- Terminology drift. The product feature is called four things across a campaign, none of them the approved name.
- Invented specifics. A number, a certification, a customer count that nobody supplied and nobody can source.
- Regressed claims. A superlative that legal removed two years ago reappears, because it is the kind of sentence companies write.
- Tone flattening. Fluent, generic marketing register that could belong to any competitor.
- Market-inappropriate framing. A claim that is permitted in the source market and a regulatory problem in another.
None of these are model quality problems. All of them are supply problems: the model was not given the thing it needed, so it produced the average of what it had seen.
The five artefacts a pipeline actually needs
| Artefact | What it contains | What it prevents |
|---|---|---|
| Approved fact pack | Every figure, date, capability and claim the brand may state, each with a source | Invented specifics; stale numbers |
| Termbase | Approved terms per language, with do-not-translate entries and forbidden synonyms | Terminology drift; mistranslated product names |
| Example set | Strong examples of approved output and negative examples with a note on why each fails | Tone flattening; the model averaging toward generic |
| Claims matrix | Which claims are permitted in which market, and which require qualification | Regulatory exposure created by a translator's reasonable choice |
| Reviewer rubric | Scored checks tied to the four artefacts above, with a defined pass threshold | Review that is an opinion rather than a measurement |
The fact pack is the highest-return item and the one most often missing. A model given a fact pack containing the eight numbers a company is allowed to state will use those eight numbers. A model given nothing will produce numbers, because marketing copy contains numbers.
The negative examples matter more than they look. "Do not sound generic" is uninterpretable; three examples of output that was rejected, each with one sentence explaining the defect, is a specification.
Designing a termbase that survives fifty markets
A termbase is not a glossary of translations. It is a control surface, and the fields it needs are more than a word pair:
- Approved term, per language, with the source-language head term.
- Do-not-translate flag for product names, feature names and legally controlled phrases.
- Forbidden alternatives — the plausible synonyms that must not be used, which is what an automated pre-check screens for.
- Context or product scope, because some terms translate differently depending on which product they appear in.
- Pronunciation, phonetically, per language, for anything that will be spoken. Brand-name pronunciation is a brand decision rather than a linguistic one, and it has to be decided explicitly per market.
- Owner and last-reviewed date, because a termbase nobody owns becomes wrong quietly.
Track exceptions rather than suppressing them. A term that a market's reviewers keep overriding is a term whose approved translation is wrong, and the exception log is the only place that shows up.
Localisation is a set of decisions, not a step
Fluent output is not the same as correct output for a market. Before any generation, decide explicitly what is fixed and what may adapt:
| Element | Typically fixed | Typically adapts |
|---|---|---|
| Product and feature names | Yes | No — unless a market-specific name exists |
| Factual claims and figures | Yes | No |
| Regulatory disclosures | No | Yes, per jurisdiction |
| Examples and scenarios | No | Yes, where a local reference is clearer |
| Formality and register | No | Yes — this is where translated copy most often reads wrong |
| Humour, wordplay, taglines | No | Yes, by transcreation from intent rather than words |
| Units, dates, currency, formats | No | Yes, as data-driven fields rather than typed strings |
The row that causes the most rework is register. A translation can be accurate at word level and wrong at the level of how a company should address a buyer in that market, and that judgement requires someone living in it. A fluent speaker abroad catches grammar and misses register — and register is most of what a brand is buying.
What a reviewer should actually check
A rubric turns "the German is off" into a list of arguable defects. Score each asset against the artefacts, not against preference:
- Factual accuracy against the fact pack — every figure traceable, no additions.
- Terminology against the termbase, including do-not-translate compliance.
- Claims against the market's row in the claims matrix.
- Register and naturalness — would a company in this market address a buyer this way.
- Cultural fit — examples, references, imagery, gestures, anything that reads as imported.
- Technical quality — formatting, units, dates, names, numbers, and for spoken content, pronunciation and pacing.
Define the pass threshold before work starts, in defects per thousand words at each severity, and make it contractual. A threshold agreed after the first delivery is a negotiation. Sample rather than reviewing everything, but sample randomly across the whole delivery rather than the first section of each file, and escalate a failed sample to full review of that batch rather than accepting a corrected sample as evidence for the rest.
Automate the boundary, not the judgement
The checks that scale are the mechanical ones, run before a human sees the asset: forbidden terms present, approved term absent, a figure that does not appear in the fact pack, a claim not permitted in this market's row, a missing required disclosure, an unapproved logo file.
Everything left after those checks is judgement, and judgement is what the human reviewer's time should be spent on. Teams that skip the automated boundary end up using expensive in-market reviewers to catch banned words, and then wonder why review is the bottleneck.
How Lifewood approaches this
Lifewood builds these artefacts as the first phase of a multilingual AIGC programme rather than as documentation produced alongside it, because the fact pack, termbase and claims matrix are what make the later volume reviewable at all. Review is run in-market — 50+ languages, 40+ delivery centres across 30+ countries, 56,788 registered contributors — on a dual-layer human-in-the-loop process held to a 95%+ accuracy threshold, with the rubric scored against the client's own artefacts rather than against a house style.
The honest limit: this is set-up cost with no output attached to it, and it is the phase clients most often want to compress. Compressing it does not remove the work; it moves it into per-asset argument, where it costs more and produces an inconsistent result.
See AIGC services.
Sources and further reading
- NIST, AI Risk Management Framework — the governance vocabulary underlying the artefact-and-rubric structure.
- Google Search Central, guidance on generative AI content — on quality expectations for published generated material.
- Companion guides: Keeping AI Imagery On-Brand at Scale (the visual equivalent of this system) and What Human-in-the-Loop Review Actually Does.

