A leading China-based AI technology company
Benchmarking in-vehicle infotainment AI across 313 sub-functions in five real-world environments
Published 8 September 2026
At a glance
| Industry | Automotive AI / in-vehicle systems | in-vehicle infotainment AI evaluation |
|---|---|---|
| Client | China-based AI technology company | unnamed under the confidentiality terms of the engagement |
| Region | China | testing conducted in-market |
| Service | Independent benchmark testing and evaluation | not data production — third-party assessment |
| Test environments | 5 | distinct real-world conditions rather than a single lab setting |
| Major functions | 9 | speech recognition · interaction · synthesis · user experience and others |
| Sub-functions | 313 | the granular level at which results were recorded |
| Method | Repeated real-world testing | observed outcomes recorded, not expected outcomes |
| Process | 5 stages | equipment prep · scenario selection · testing · result recording · report generation |
| Reporting | Analytics and visualised reports | delivered as full coverage of every scenario and function |
| Status | Delivered | client applied results to system optimisation |
This is a Lifewood Data Technology case study in Benchmark testing — an engagement delivered for a leading China-based AI technology company across China. Benchmarking in-vehicle infotainment AI across 313 sub-functions in five real-world environments
Vendor-reported performance is measured under conditions the vendor chooses
A leading China-based AI technology company needed to benchmark in-vehicle infotainment AI systems against each other. The problem with relying on vendor-reported figures is not that they are dishonest — it is that they are measured under conditions the vendor selects, and a moving vehicle cabin is not one of those conditions. Road noise, competing passenger speech, HVAC output, variable microphone distance and accented commands all degrade recognition in ways a quiet-room benchmark never surfaces. The second difficulty is resolution. Nine major functions describe an infotainment system at a level too coarse to act on: knowing that speech recognition scores well tells an engineering team nothing about which of its constituent behaviours is failing. Useful benchmarking has to go down to the level where a fix can actually be made, which is where the 313 sub-function breakdown comes from.
A five-stage evaluation covering every function in five real-world environments
Lifewood ran the benchmark through a dedicated process engineering team for planning, analysis and evaluation, with evaluation specialists executing the tests. The programme followed five stages. 1. Equipment preparation. Test rigs and recording setup standardised so results from different environments remain comparable. 2. Scenario selection. Five distinct real-world environments defined, rather than a single controlled setting, so the results describe behaviour in conditions the system will actually meet. 3. Testing. Evaluation experts ran multiple real-world tests across all 9 major functions and 313 sub-functions, spanning speech recognition, interaction, synthesis and user experience. 4. Result recording. Observed outcomes were recorded rather than expected ones — the discipline that separates a benchmark from a specification review. 5. Report generation. Results delivered through an analytics and visualisation platform as reports covering every scenario and function. Independence is the point of the exercise. Because Lifewood produces training data but did not build the systems under test, it has no stake in which system scores best, and the benchmark carries a credibility a first-party evaluation cannot.
Full-coverage test reports the client used to optimise system performance
Lifewood delivered comprehensive test reports covering all 5 environments, all 9 major functions and all 313 sub-functions, with no scenario left unreported. The client applied the results to optimise and improve its in-vehicle infotainment system performance. Complete coverage is the meaningful outcome on a benchmark engagement. A partial benchmark tells an engineering team where a system performs well and leaves the gaps ambiguous — which is the failure mode that makes many evaluations unusable for prioritisation.
Related Lifewood services
Verified outcomes
| Metric | Value | Baseline | How measured |
|---|---|---|---|
| Test environments covered | 5 | real-world conditions, not a single lab setup | Scenario selection stage |
| Major functions evaluated | 9 | speech, interaction, synthesis, user experience and others | Client evaluation specification |
| Sub-functions evaluated | 313 | the granular level a fix can be made at | Result recording stage |
| Scenario coverage | Complete | no scenario or function left unreported | Delivered report set |
| Outcome recording | Observed, not expected | the distinction between a benchmark and a spec review | Evaluation specialist records |
| Client application | System optimisation | results applied to improve IVI performance | Client-reported |
Method and verification. Figures on this page are Lifewood-reported and describe delivered scope. Environment, function and sub-function counts are the evaluation specification as executed, and coverage was complete against that specification. Recorded outcomes are observed test results rather than expected or vendor-stated values. Lifewood did not build any of the systems under test and had no commercial interest in the relative rankings, which is the basis on which the benchmark is offered as independent. The number of systems benchmarked and the engagement dates are not published because they are not yet confirmed. The client is not named on this page under the confidentiality terms of the engagement, and per-system scores are the client's property and are not published.
Questions about this programme
On Lifewood benchmark engagements the value is that performance is measured under real-world conditions rather than conditions the vendor selects — a quiet-room speech benchmark does not predict behaviour in a moving cabin.
Lifewood recorded this in-vehicle infotainment benchmark at sub-function level because nine major function scores are too coarse to act on; a fix has to be made at the level the failure occurs.
A Lifewood AI benchmark programme runs equipment preparation, scenario selection, testing, result recording and report generation, with observed outcomes recorded rather than expected ones.
Lifewood did not build any of the systems evaluated on this benchmark engagement and held no interest in the relative rankings, which is the basis for offering the evaluation as independent.
Run a similar program with Lifewood?
Tell us your scope and accuracy bar. We will scope a comparable engagement within one call.
Talk to our team