When is a factory digital twin trustworthy enough to guide a decision?
A digital twin can present a convincing picture of a machine while giving an unreliable answer about its next operating condition. Manufacturers need evidence that the model is suitable for the decision being made, alongside evidence that its data connection works.
As of February 28, 2025. NIST's Security and Trust Considerations for Digital Twin Technology, finalized on February 14, puts trust alongside the technology's potential uses. Its discussion includes accuracy, synchronization and the relationship between a representation and a changing physical object. The report is technical guidance, not a certification that a particular commercial twin will deliver reliable predictions. For a factory, that distinction matters when a model moves from explaining yesterday's production to recommending tomorrow's settings. A useful visualization and a dependable operational forecast are different achievements. Before relying on a recommendation, the team should know what was tested, under which conditions and with what remaining uncertainty.
Begin with the decision the model will support
The phrase digital twin can cover a broad set of capabilities. A maintenance team may want to compare observed vibration with expected behaviour. A production planner may want to estimate how a different product mix affects a bottleneck. A process engineer may want to investigate a proposed operating change. These uses need different evidence.
The Digital Twin Consortium's 2024 capabilities guide provides a way to connect use cases with required capabilities. It helps structure a requirements discussion; it does not establish that an implementation is accurate. That distinction is valuable when suppliers present a long list of functions without explaining which functions the customer actually needs.
An illustrative pumping-system project makes the issue concrete. A model might reproduce steady operation at a familiar flow rate. The maintenance team then asks whether it can predict behaviour during an unusual startup sequence. That second question introduces conditions which the original comparison may never have tested.
The project brief should specify the output, the operating range and the action that follows. It should also state who can accept or reject a recommendation. This makes validation a practical production requirement, rather than an abstract request for a model to be accurate in every possible situation.
A shared framework does not settle model credibility
ISO 23247-1:2021 describes general principles and requirements for a manufacturing digital twin framework. A framework helps establish a common structure and vocabulary. The public scope does not offer a performance guarantee for a specific factory model, and a reference to the standard should not be read as such a guarantee.
An earlier NIST analysis of the ISO 23247 series discusses the need for verification, validation and uncertainty quantification throughout the twin's life. It identifies credibility assessment as an area needing further development beyond the initial framework. This is a useful distinction between organizing the system and showing that its conclusions are dependable.
For a manufacturer, a procurement review can ask two separate questions. Can the system represent and exchange the required information? Can its predictions support the intended operational decision? Success on the first question is necessary for many applications, but it leaves the second question open.
The acceptance test should therefore contain examples of the decision itself. If the model is intended to help schedule maintenance, showing an attractive equipment screen is insufficient. The review should examine whether its outputs distinguish the situations that would lead the maintenance team to act, investigate or continue monitoring.
Test the model against conditions it has not memorized
NASA's 2019 handbook for models and simulations discusses empirical validation, its domain and the risk of overfitting. It recommends preserving data for validation rather than using everything to tune the model. These are useful background concepts for industrial modelling, although NASA's handbook is not a general factory compliance requirement.
The practical concern is straightforward. If engineers repeatedly adjust a model until it matches one historical production run, the match may say more about that run than about future operations. Testing against separate runs provides a more meaningful challenge. The separation should reflect the real way the model will be used.
For example, an illustrative quality model could be tested on later production batches, including the normal variation in material supply and operating conditions. Randomly dividing highly similar readings from the same batch may create an easier test than deployment will present. The team should explain why its test represents the intended use.
A validation report also needs a comparison point. If an existing rule or simple calculation performs just as well for the decision, a more elaborate model may not justify its additional operating burden. Complexity should earn its place through useful performance, broader coverage or a capability that simpler methods cannot provide.
The result should identify where confidence ends. A model tested within one temperature range should not silently retain the same status outside it. Recording that boundary gives operators a reason to seek further evidence when conditions change.
Decide how uncertainty affects action
A forecast can be helpful without being exact. The relevant question is whether its uncertainty changes the decision. A modest error may be tolerable for a broad planning estimate while being unacceptable for a tightly constrained process adjustment. There is no universal accuracy percentage that resolves every use case.
Consider an illustrative maintenance forecast which suggests a component should be inspected during the next planned shutdown. If a reasonable uncertainty range moves the inspection date into the current production campaign, the decision deserves additional review. Displaying only one predicted date would conceal information that matters to the maintenance planner.
Teams can define action bands around the decision. One region may support routine use, another may require human investigation, and a third may fall outside the model's approved purpose. The actual thresholds should reflect process knowledge, consequence and evidence, rather than a convenient colour scheme on a dashboard.
Errors should also be examined by operating condition. A single average can hide poor performance on a less common product or during a transition. Those cases may occupy little of the dataset while accounting for much of the practical concern. A useful review makes their performance visible.
Keep the model aligned with the changing asset
The physical system will not remain exactly as it was during development. Maintenance can replace a component, a control change can alter behaviour, and a sensor can be recalibrated. The twin's usefulness depends on a process for deciding which changes require investigation or renewed validation.
A practical ownership record should connect the model version with the equipment configuration, data sources and approved purpose. Someone should be responsible for assessing changes. Without that ownership, a model can remain available long after the assumptions supporting its use have become uncertain.
Data freshness should be judged against the decision. A slow update may be acceptable for a long-term capacity study but unsuitable for a short operational response. The system should make missing, delayed or substituted data visible, so that a plausible output is not mistaken for a fully supported one.
The manufacturer also needs a fallback. If the model is unavailable or outside its validated range, operators should know which existing procedure applies. That operational arrangement helps prevent convenience from turning an advisory tool into an undocumented dependency.
A credible digital twin therefore comes with more than a demonstration. It has a defined use, evidence from relevant tests, visible limits and a maintained connection to the asset it represents. For factories considering their next investment, those features provide a stronger basis for trust than visual realism or the number of data points displayed on screen.
