AI's measurable returns in physical product development come not from generating novel geometry but from evaluating what already exists — inspecting parts, reducing test programmes, and compressing simulation cycles.
Akash SinghSEPTEMBER 20268 MIN READ
The dominant narrative around AI in physical product development centres on generation: algorithms that propose novel shapes, topologies, and material combinations a human designer would not conceive. General Motors' generative-design seat bracket — 40 per cent lighter, eight parts consolidated into one — remains the canonical example, six years after its debut.¹ It is a real result. It is also a prototype. The part has not entered series production, and the programme has not been replicated at scale across GM's vehicle range.
The quieter story is evaluation. Inspection, simulation triage, test-plan optimisation, and digital-twin-based operational feedback are where AI has moved from demonstration to deployment — and where the financial returns are documented rather than projected.
99.2 per cent detection accuracy redefines the economics of inspection
AI-powered visual inspection is the most commercially mature application of machine learning in physical manufacturing. The numbers are no longer disputed in narrow domains. AI vision systems sustain 99.2 per cent detection accuracy across production shifts; human inspectors peak at 87 per cent and decline to 70 per cent after four hours, with 34 per cent variability between individuals.² The throughput gap is equally stark: 2.4 seconds per part versus 38 seconds, a 15-fold difference.²
The market reflects this. Cognex reported $994 million in revenue in 2025, adding 9,000 new manufacturing accounts. Keyence recorded $7.16 billion, commanding 28–32 per cent of the semiconductor inspection segment.² NVIDIA's Metropolis platform now partners with TSMC for nanometre-scale defect detection.² These are not pilot programmes.
A modelled deployment at a facility processing 1,200 parts per day shows a three-year return of 374 per cent on an initial $240,000 investment, with payback at eight months. The savings divide roughly into $342,000 in annual labour displacement and $144,000 in escaped-defect cost reduction, against a $78,000 annual platform fee.² These figures are vendor-modelled, not independently audited, and should be read as directional.
Instrumental, which counts Apple, Tesla, and Logitech among its customers, found that 4.6 per cent of units passing three sequential human operators carried real defects that its system intercepted.³ Sandia National Laboratories established the baseline in 2012: a single human inspection pass catches roughly 80 per cent of defects.⁴ That figure has not been seriously challenged by subsequent studies, which suggests the constraint is biological, not procedural.
17 per cent fewer physical tests is modest — and that is why it matters
AI-driven test-plan optimisation is a younger category, and its reported gains are smaller. Nissan Technical Centre Europe, working with Monolith AI, achieved a 17 per cent reduction in physical bolt-joint validation testing on the Nissan LEAF programme.⁵ Monolith's own beta users report 30–60 per cent reductions in validation test counts, though these figures depend on initial test-plan inefficiency and are vendor-reported.⁶
JOTA Sport, the endurance racing team, used the same platform to cut aerodynamic and vehicle-setup test effort by up to 80 per cent.⁷ Motorsport is a useful leading indicator: test budgets are constrained, feedback loops are short, and failure consequences are immediate. The transfer to series-production automotive programmes is underway — Nissan extended its Monolith partnership through 2027 — but remains early.⁵
The mechanism matters more than the headline number. These systems do not simulate physics. They analyse historical test data to identify which proposed tests carry the most information value, then recommend a test sequence that reaches the same confidence interval with fewer physical runs. Nissan's model drew on nearly 90 years of accumulated test records.⁷ The approach works precisely because it does not attempt to replace physics-based simulation; it triages which simulations and physical tests are worth running.
McKinsey's 2025 R&D Leaders Forum reported a 20 per cent reduction in rework achievable through deep-learning surrogates in product development, with overall development time halved in some medtech applications.⁸ These are forum-reported figures, not peer-reviewed findings, and the phrase "achievable" signals aspiration more than realisation.
Surrogate models run in seconds — inside the bounds of their training data
AI surrogate models — neural networks trained on finite-element or computational-fluid-dynamics results to approximate the same outputs in seconds rather than hours — are the application closest to replacing traditional engineering computation. Comsol's senior VP of product management, Bjorn Sjodin, states the capability plainly: within pre-defined parametric ranges, a surrogate "could be arbitrarily good. Actually, it could be just as good as the finite element model."⁹
The operative phrase is "within pre-defined parametric ranges." A surrogate trained on five or six driving parameters — CAD dimensions, material properties, boundary conditions — with defined minimum and maximum values will perform well inside that envelope. Outside it, reliability degrades without warning. As Sjodin also notes: "Sometimes it hallucinates, you know, gives the wrong answer."⁹
The practical constraint is that a surrogate model cannot exceed the accuracy of the physics-based simulation used to generate its training data.¹⁰ Every surrogate is a compression of a higher-fidelity model, not an independent source of truth. Domain expertise becomes "much more important, not less" when deploying these systems, because the engineer must recognise when a prediction has left the training envelope.¹⁰
The speed gain is real: a simulation that takes an hour can return in seconds, and factory-floor decisions that need answers in under five minutes become feasible.⁹ Ansys SimAI, Siemens Simcenter, Altair PhysicsAI, and Neural Concept all offer commercial platforms. Bombardier has signed a multimillion-dollar AI contract for simulation acceleration.⁶ But a 2026 review of every major tool on the market concluded: "None universally replaces validated FEA, CFD or physical testing."⁶
BMW projects a 30 per cent cut in planning costs across 30 factories
Digital twins occupy the capital-intensive end of the evaluation spectrum. BMW Group operates digital replicas across more than 30 production sites globally, built on NVIDIA's Omniverse platform. The company projects a 30 per cent reduction in production planning costs and has compressed collision-check cycles from approximately four weeks to three days.¹¹ Between now and 2027, BMW intends to integrate more than 40 new or updated vehicles into global production using this infrastructure.¹¹
IKEA digitalised 37 retail stores across East Asia in nine months through Akila's platform, connecting 7,000 data points from 10 different HVAC manufacturers and monitoring 6,000 units. The result was a 30 per cent reduction in HVAC energy consumption for central supply systems, with annual savings estimated in the millions of dollars.¹²
A Hexagon survey of 660 executives found that 92 per cent of companies tracking digital-twin returns saw gains exceeding 10 per cent, and 50 per cent reported returns above 20 per cent. Eighty per cent of respondents said AI had heightened their interest in the technology.¹³ These are self-reported figures from a vendor-sponsored survey, and the selection bias — executives who track returns are predisposed to report them — is evident.
The digital-twin market is valued at $36.19 billion in 2026 and projected to reach $240.3 billion by 2035, implying a 30.54 per cent compound annual growth rate.¹⁴ Alternative projections run as high as $626.07 billion.¹⁴ The spread between estimates suggests the category's boundaries remain contested.
Regulators have not agreed that AI evaluation replaces physical evidence
The falsification condition for this argument is regulatory acceptance. If AI-generated simulation and inspection data were accepted as equivalent to physical test evidence by the FDA, EASA, or ISO certification bodies, the evaluation layer would subsume the physical layer entirely. That has not happened.
The FDA issued its first warning letter for AI compliance failures in 2026.¹⁵ The core finding, as a Forbes analysis summarised it: "A model may perform exceptionally well and still lack the evidence package necessary to support regulatory review."¹⁶ The gap is not accuracy but traceability — 21 CFR Part 11 requirements for documentation chains, predetermined change-control plans for model updates, and continuous post-market monitoring.¹⁶
McKinsey reports that 71 per cent of organisations now use generative AI regularly in at least one business function. One per cent of executives describe their rollouts as mature.¹⁷ The distance between adoption and maturity is where the actual work lies.
Physical testing persists not because AI cannot match its accuracy in controlled conditions, but because the regulatory and liability architecture of manufactured goods demands evidence that a neural network's training envelope cannot yet guarantee. The evaluation layer accelerates the path to that evidence. It does not replace the evidence itself.
What changes for one person
A test engineer at an automotive OEM now spends less time running redundant physical tests and more time deciding which tests carry information the model has not already captured — a different job, requiring the same domain knowledge and a new statistical fluency.
Sources
- [1]Modern Machine Shop, "What This Seat Bracket Says About the Future of Automotive Manufacturing," 2018. https://www.mmsonline.com/articles/what-this-seat-bracket-says-about-the-future-of-automotive-manufacturing
- [2]IIoT World, "AI Vision in Manufacturing Quality Inspection," 2026. https://www.iiot-world.com/smart-manufacturing/ai-vision-quality-manufacturing-2026/
- [3]Instrumental, "Manual Inspection vs. AI Inspection with Instrumental," undated. https://instrumental.com/build-better-handbook/machine-vision-vs-manual-inspection-vs-instrumental
- [4]Sandia National Laboratories, human inspection reliability baseline, 2012. Referenced via Instrumental.
- [5]Monolith AI, "Nissan Harnesses Power of AI to Speed Up Physical Vehicle Tests," November 2025. https://www.monolithai.com/press-release/nissan-harnesses-power-of-ai-to-speed-up-physical-vehicle-tests-monolith
- [6]CoLab Software, "Best AI Simulation Tools for Mechanical Engineers (2026)," 2026. https://www.colabsoftware.com/guides/ai-powered-simulation-tools-smarter-faster-design-validation
- [7]Monolith AI, "Engineering AI at Scale: Use Cases That Stood Up in 2025," 2025. https://www.monolithai.com/blog/engineering-ai-at-scale-use-cases-2025
- [8]McKinsey & Company, "Breakthroughs in AI-Augmented R&D: Recap from the 2025 R&D Leaders Forum," 2025. https://www.mckinsey.com/capabilities/operations/our-insights/operations-blog/breakthroughs-in-ai-augmented-r-and-d-recap-from-the-2025-r-and-d-leaders-forum
- [9]Engineering.com, "Simulation Trends for 2025: Get Ready for AI and Surrogate Models," 2025. https://www.engineering.com/simulation-trends-for-2025-get-ready-for-ai-and-surrogate-models/
- [10]Digital Engineering 24/7, "Read This Before You Develop a Surrogate Simulation Model," 2025. https://www.digitalengineering247.com/article/read-this-before-you-develop-a-surrogate-simulation-model
- [11]BMW Group, "BMW Group Scales Virtual Factory," 2025. https://www.press.bmwgroup.com/global/article/detail/T0450699EN/bmw-group-scales-virtual-factory
- [12]World Economic Forum, "Pairing AI and Digital Twin Technology to Cut Emissions," March 2024. https://www.weforum.org/stories/2024/03/how-digital-twin-technology-can-work-with-ai-to-boost-buildings-emissions-reductions/
- [13]Hexagon, "The Digital Twin Industry Report," 2024. Via Visual Capitalist: https://www.visualcapitalist.com/dp/charted-the-return-on-investment-of-digital-twins/
- [14]MindInventory, "Digital Twin Statistics 2026," 2026. https://www.mindinventory.com/blog/digital-twin-statistics/
- [15]MasterControl, "FDA Issues First Warning Letter for AI Compliance," 2026. https://www.mastercontrol.com/gxp-lifeline/first-fda-warning-letters-for-ai-compliance/
- [16]Forbes Tech Council, "Your AI Model Passed Validation But Your FDA Audit Didn't — Here's Why," July 2026. https://www.forbes.com/councils/forbestechcouncil/2026/07/14/your-ai-model-passed-validation-but-your-fda-audit-didnt-heres-why/
- [17]IBTimes, "McKinsey Says AI Is Compressing 9-Month Product Cycles Into 2 Weeks," 2025. https://www.ibtimes.com/mckinsey-says-ai-compressing-9-month-product-cycles-2-weeks-most-companies-arent-ready-3802992


