AI quality inspection: evidence beyond accuracy scores
A factory does not buy an inspection system to achieve a benchmark score. It buys a way to find unacceptable parts, avoid unnecessary rejection and keep production moving. New industrial anomaly-detection research makes the gap between those objectives increasingly clear.

As of May 20, 2025. Three research releases this spring examine difficult lighting, fine surface geometry and internal defects. Together they suggest a better buying question: what evidence shows that the proposed inspection will work on this line, with this defect definition and this production mix?
What the new benchmarks test
The March 2025 MVTec AD 2 preprint introduces eight inspection scenarios with more than 8,000 high-resolution images. Its challenges include transparent and overlapping objects, difficult illumination and very small anomalies. It also examines changes in lighting between training and testing. These choices matter because the appearance of a normal object can change without its quality changing.
The paper reports that strong existing methods still struggle on the new benchmark. A score on an older, easier dataset therefore provides limited evidence about a tougher application. The authors also distinguish metrics that summarize performance across thresholds from results at a selected operating threshold. Neither should be casually renamed production accuracy.
This is benchmark research, not an independent acceptance test of a purchased installation. Its practical contribution is to expose failure conditions that a local trial should investigate. A manufacturer can use those conditions to challenge a vendor demonstration without assuming that every defect in its own process resembles the benchmark.
The image must contain the relevant evidence
The April 19 Real-IAD D3 preprint combines ordinary color images, photometric-stereo information and three-dimensional point clouds across 20 product categories. Its researchers deliberately introduced defects during dataset preparation. This creates controlled comparisons, but the resulting defect mix should not be treated as the natural defect frequency of a production line.
The study also examines reduced point-cloud resolution. Fine flaws can become harder to detect when detail is removed. That result connects algorithm performance to the acquisition system: a model cannot reliably recover a defect that the selected imaging arrangement fails to represent. More modalities may help particular tasks, but this study does not establish that every factory needs a three-dimensional scanner.
A different May preprint, CXR-AD, studies X-ray inspection of internal component defects. It describes low contrast, complex internal backgrounds and defects at different scales. Its relevance to a surface-camera buyer is a boundary, rather than a competing product recommendation. A visible-light surface image cannot be assumed to reveal every internal quality problem.
These studies support separating two decisions. First establish how the relevant defect can be observed. Then compare methods that interpret that evidence. Beginning with a fashionable model and trying to make the defect fit its input can produce an impressive demonstration of the wrong inspection task.
Why a high accuracy figure can mislead
Consider a deliberately simplified example: 10,000 parts contain 100 defective items. A system that passes every part classifies 9,900 correctly and therefore reports 99 percent accuracy. It also misses every defect. These are illustrative numbers, not the results of any cited study or commercial system.
For a production discussion, the missing information includes how many defects escaped, how many conforming parts were rejected and which defect types caused those errors. A plant might tolerate additional review of uncertain cosmetic marks while requiring a different response to a functional defect. Combining all of them into one percentage conceals that distinction.
The denominator matters too. A false-rejection percentage calculated among conforming parts is different from the share of rejected parts that later prove acceptable. Both can be useful, but they answer different questions. A proposal should define the numerator, denominator and sample composition before its figures are compared with another proposal.
A useful acceptance report would therefore retain counts as well as rates. It would identify the test population, the number of each defect and the decisions made. If a rare but important defect appears only a handful of times, the resulting evidence is limited even when every example is detected.
Choose the operating decision before the final test
Many anomaly-detection systems produce a score rather than a complete production decision. Someone still chooses the point at which a score triggers rejection, review or another measurement. Moving that threshold changes the balance between escaped defects and false alarms. A factory needs to understand that balance at the setting it intends to use.
As an editorial testing recommendation, the threshold should be settled using development data before the final acceptance sample is assessed. Repeatedly adjusting it after viewing the final results can make the apparent performance depend on knowledge of the test itself. Keeping a separate final sample gives the buyer a clearer account of what has actually been demonstrated.
Amazon's earlier VisA dataset provides another useful reference point. Its public registry describes 12 object classes and separately counts normal and anomalous images. That documentation illustrates why dataset composition should accompany a score. It does not demonstrate the outgoing quality of a particular factory, and its sample proportions need not match local production.
The commercial implication is straightforward: an attractive public benchmark result can justify a local trial. It cannot substitute for defining the plant's own acceptance decision and evaluating it against independently checked examples.
Test ordinary variation alongside defects
A local trial should include acceptable variation that is likely to occur in routine work. Different approved suppliers, batches, surface finishes and product variants are sensible candidates for investigation. Whether each one matters will depend on the actual process. The aim is to test the system against the intended operating envelope, not to invent an unlimited list of hypothetical failures.
For example, a hypothetical machined-component line might compare images before and after a scheduled lighting adjustment, while retaining the same approved parts. If rejection rises, the team can investigate whether the change affects the image or the parts. This is a proposed diagnostic comparison, not a claim that a particular system will fail.
NIST's 2023 AI Risk Management Framework provides relevant background by calling for evaluation under conditions resembling deployment and continued assessment during operation. It is a voluntary framework, not a certification of an inspection product. Applied here, its value is the connection between initial testing, operating conditions and later monitoring.
A practical monitoring arrangement might track review volume and confirmed errors by product family. That creates a way to notice when the operating result changes. The process also needs a named owner who can investigate a change and decide whether further validation is needed before the system resumes its intended role.
Measure the production result
A camera decision is only one step in an inspection station. The plant still has to identify the part, route it correctly, retain the relevant result and handle an uncertain decision. A trial should therefore observe the complete sequence at the intended production pace. A fast model running on stored images does not establish the throughput of that sequence.
The commercial comparison should include the work created by false alarms. As an illustrative budgeting exercise, a buyer could estimate the review time per rejected item and apply it to the trial's measured rejection counts. It should keep that estimate separate from confirmed reductions in scrap or customer complaints, which require their own evidence.
The spring research offers useful advances in how inspection problems are tested. For manufacturers, the most valuable next step is a narrower and more demanding evaluation: the right sensing method, an explicit defect definition, a fixed decision rule and evidence from the intended line. Those details connect research progress to an inspection system that can earn its place in production.
Source: MVTec and Technical University of Munich, The MVTec AD 2 Dataset · Cover: AI-generated illustration
