Evaluating AI for Real-World Performance
An AI model can perform well in a test environment, yet struggle when it is used in the real world. Different data, changing conditions, or an unexpected use case can expose weaknesses.
So how do you know whether an AI model will actually work when and where you need it?
A useful evaluation can help you understand where a model performs reliably, what causes its performance to decline, and how those limitations will impact your application. Here are practical ways to put that approach into practice.
1. Start With What You Need to Learn
Before selecting a benchmark or metric, define the decision the evaluation needs to support by asking:
- What task does the model need to perform?
- What types of data will it encounter?
- Which aspects of performance matter for that task?
- Which failures would have the greatest impact?
Identifying these requirements gives the evaluation a clear purpose. It also makes it easier to determine whether a benchmark is valid and represents the conditions that matter based on how the technology is going to be used.
This is the approach we take through our work on DARPA’s Artificial Intelligence Quantified (AIQ) program. As part of that effort, Kitware leads test and evaluation through the Mathematical Assurance and Generative AI Network Evaluation Toolkit (MAGNET). For example, we evaluated an AIQ research approach designed to predict how models would perform on benchmarks based on their responses to a set of queries. The evaluation included approximately 100 different models and benchmarks spanning different domains. This allowed us to test whether the approach could predict performance across a broader range of models and evaluation settings.
Fitting AIQ into Industry T&E Processes

The same principle applies when evaluating your own AI system: Start with the question the evaluation needs to answer, then select the data, conditions, and methods that will give you meaningful evidence.
2. Find Where Performance Begins to Change
A single aggregate score gives you limited insight into a model’s behavior. In examining how performance changes as the evaluation gets harder, you can identify the conditions where accuracy or other important metrics begin to decline.
This type of stress testing is the idea behind Kitware’s Natural Robustness Toolkit (NRTK). Rather than requiring teams to collect entirely new datasets, NRTK systematically changes existing data to simulate conditions a computer vision system might encounter in the real world, such as haze, changing light, noise, lower resolution, or lens distortion. Teams can control the severity of these changes and measure how the model responds as conditions become more challenging.
This helps reveal where the model’s performance starts to break down. Those boundaries can help teams determine where additional testing, model development, or data collection may be needed.
3. Document What the Result Actually Means
That means documenting the model and dataset used, testing conditions, metrics, relevant assumptions, and known limitations.
Keep this information together in an evaluation record that clearly states what was tested, how it was tested, and what conclusions the evidence supports. In MAGNET, Kitware uses Evaluation Cards as a structured format for documenting evaluation results and the evidence behind them. This gives future reviewers enough context to determine whether the result still applies as the model, data, or application changes.
This is particularly important when an evaluation result may later support a deployment decision. Someone reviewing the result should be capable of understanding both what it demonstrates and what it does not.
4. Build Evaluation Into the Development Process
Evaluation should not become a single checkpoint immediately before deployment.
Models change. Data changes. Applications expand. The operating environment may introduce conditions that were not represented during the original evaluation. Plan to reevaluate when the model or data changes, when the system is applied to a new use case, or when operating conditions shift in ways that could affect performance.
The process is iterative:
Define what matters → evaluate → identify limitations → refine the model or evaluation → test again.
Build this cycle into the development process so that decisions about new versions, data, use cases, or operating conditions are based on evidence that still applies.
Make Your AI Evaluation More Meaningful
The goal of AI evaluation is not simply to produce another score. It is to generate evidence that helps teams decide where an AI system can be used reliably and where more work may be needed.
Before using an evaluation result to support a development or deployment decision, make sure your team can answer three questions:
- What was the model expected to do?
- When does its performance begin to change?
- What evidence supports the conclusions?
A score becomes much more useful when your team understands what it represents, where its limitations are, and how it applies to the intended use.
These recommendations are grounded in the same approaches Kitware applies in our own AI research and development. Through DARPA’s AIQ program, including our work supporting the development of MAGNET, and Phase II of the Chief Digital and AI Office’s (CDAO) Joint AI Test Infrastructure Capability (JATIC) program, which includes NRTK, we are creating resources and establishing best practices for evaluating AI systems across different tasks, datasets, and modalities.
If you’re evaluating an AI system and aren’t sure whether your current testing is giving you the full picture, we’re happy to talk through it with you. Contact our Team to discuss your test and evaluation needs.
Disclaimers and Acknowledgments:
Effort sponsored by the U.S. Government under the Tradewind Project Agreement. The U.S. Government is authorized to reproduce and distribute reprints for Governmental purposes notwithstanding any copyright notation thereon. The views and conclusions contained herein are those of the authors and should not be interpreted as necessarily representing the official policies or endorsements, either expressed or implied, of the U.S. Government.
This material is based upon work supported by the Defense Advanced Research Project Agency (DARPA) under Contract No. HR001125CE017. Any opinions, findings and conclusions or recommendations expressed in this material are those of the author(s) and do not necessarily reflect the views of the Defense Advanced Research Project Agency (DARPA).