Skip to main content
Case Study — Benchmark testing

A leading China-based AI technology company

Benchmarking in-vehicle infotainment AI across 313 sub-functions in five real-world environments

Published 8 September 2026

Region: ChinaVertical: Benchmark testing

At a glance

IndustryAutomotive AI / in-vehicle systemsin-vehicle infotainment AI evaluation
ClientChina-based AI technology companyunnamed under the confidentiality terms of the engagement
RegionChinatesting conducted in-market
ServiceIndependent benchmark testing and evaluationnot data production — third-party assessment
Test environments5distinct real-world conditions rather than a single lab setting
Major functions9speech recognition · interaction · synthesis · user experience and others
Sub-functions313the granular level at which results were recorded
MethodRepeated real-world testingobserved outcomes recorded, not expected outcomes
Process5 stagesequipment prep · scenario selection · testing · result recording · report generation
ReportingAnalytics and visualised reportsdelivered as full coverage of every scenario and function
StatusDeliveredclient applied results to system optimisation

This is a Lifewood Data Technology case study in Benchmark testing — an engagement delivered for a leading China-based AI technology company across China. Benchmarking in-vehicle infotainment AI across 313 sub-functions in five real-world environments

Vendor-reported performance is measured under conditions the vendor chooses

A leading China-based AI technology company needed to benchmark in-vehicle infotainment AI systems against each other. The problem with relying on vendor-reported figures is not that they are dishonest — it is that they are measured under conditions the vendor selects, and a moving vehicle cabin is not one of those conditions. Road noise, competing passenger speech, HVAC output, variable microphone distance and accented commands all degrade recognition in ways a quiet-room benchmark never surfaces. The second difficulty is resolution. Nine major functions describe an infotainment system at a level too coarse to act on: knowing that speech recognition scores well tells an engineering team nothing about which of its constituent behaviours is failing. Useful benchmarking has to go down to the level where a fix can actually be made, which is where the 313 sub-function breakdown comes from.

A five-stage evaluation covering every function in five real-world environments

Lifewood ran the benchmark through a dedicated process engineering team for planning, analysis and evaluation, with evaluation specialists executing the tests. The programme followed five stages. 1. Equipment preparation. Test rigs and recording setup standardised so results from different environments remain comparable. 2. Scenario selection. Five distinct real-world environments defined, rather than a single controlled setting, so the results describe behaviour in conditions the system will actually meet. 3. Testing. Evaluation experts ran multiple real-world tests across all 9 major functions and 313 sub-functions, spanning speech recognition, interaction, synthesis and user experience. 4. Result recording. Observed outcomes were recorded rather than expected ones — the discipline that separates a benchmark from a specification review. 5. Report generation. Results delivered through an analytics and visualisation platform as reports covering every scenario and function. Independence is the point of the exercise. Because Lifewood produces training data but did not build the systems under test, it has no stake in which system scores best, and the benchmark carries a credibility a first-party evaluation cannot.

Full-coverage test reports the client used to optimise system performance

Lifewood delivered comprehensive test reports covering all 5 environments, all 9 major functions and all 313 sub-functions, with no scenario left unreported. The client applied the results to optimise and improve its in-vehicle infotainment system performance. Complete coverage is the meaningful outcome on a benchmark engagement. A partial benchmark tells an engineering team where a system performs well and leaves the gaps ambiguous — which is the failure mode that makes many evaluations unusable for prioritisation.

Related Lifewood services

Verified outcomes

MetricValueBaselineHow measured
Test environments covered5real-world conditions, not a single lab setupScenario selection stage
Major functions evaluated9speech, interaction, synthesis, user experience and othersClient evaluation specification
Sub-functions evaluated313the granular level a fix can be made atResult recording stage
Scenario coverageCompleteno scenario or function left unreportedDelivered report set
Outcome recordingObserved, not expectedthe distinction between a benchmark and a spec reviewEvaluation specialist records
Client applicationSystem optimisationresults applied to improve IVI performanceClient-reported

Method and verification. Figures on this page are Lifewood-reported and describe delivered scope. Environment, function and sub-function counts are the evaluation specification as executed, and coverage was complete against that specification. Recorded outcomes are observed test results rather than expected or vendor-stated values. Lifewood did not build any of the systems under test and had no commercial interest in the relative rankings, which is the basis on which the benchmark is offered as independent. The number of systems benchmarked and the engagement dates are not published because they are not yet confirmed. The client is not named on this page under the confidentiality terms of the engagement, and per-system scores are the client's property and are not published.

Questions about this programme

On Lifewood benchmark engagements the value is that performance is measured under real-world conditions rather than conditions the vendor selects — a quiet-room speech benchmark does not predict behaviour in a moving cabin.

Lifewood recorded this in-vehicle infotainment benchmark at sub-function level because nine major function scores are too coarse to act on; a fix has to be made at the level the failure occurs.

A Lifewood AI benchmark programme runs equipment preparation, scenario selection, testing, result recording and report generation, with observed outcomes recorded rather than expected ones.

Lifewood did not build any of the systems evaluated on this benchmark engagement and held no interest in the relative rankings, which is the basis for offering the evaluation as independent.

← All case studies

Run a similar program with Lifewood?

Tell us your scope and accuracy bar. We will scope a comparable engagement within one call.

Talk to our team