Advancing the Science of Foundation Model Performance at ICML 2026
As foundation models are increasingly deployed to critical workflows, knowing what a model can do is only part of the challenge. Researchers also need to understand when those capabilities will manifest and why a model may succeed in one situation but fail in another.
Addressing this unpredictability requires a stronger connection between theory and empirical evaluation. This was the focus of the “Combining Theory and Benchmarks: Towards a Virtuous Cycle to Understand and Guarantee Foundation Model Performance” workshop, which was held on July 10 in Seoul, South Korea. The workshop was one of just 44 selected from 247 proposals (17.8% acceptance rate) for the International Conference on Machine Learning (ICML), one of the premier academic conferences for machine learning and artificial intelligence research. Kitware’s Brian Hu served as an organizer.
The workshop brought together researchers from mathematics, statistics, machine learning, and industry. Together, they explored how theory and structured evaluation can contribute to a more predictive science of foundation model performance.
Understanding the “Jagged Edge of Intelligence”
It remains difficult to predict when a frontier model will succeed or fail, where a model can sometimes perform very complex tasks while also failing at seemingly simple ones. This challenge is known as the “jagged edge of intelligence.”
For industry, this unpredictability is particularly important as foundation models are deployed to critical workflows. Guaranteeing their performance is necessary.
The workshop centered on three research challenges that can help build that understanding. One is quantifying model capabilities across different scales and levels. Another is establishing foundations for understanding model generalization and composition. The third is developing reliable and structured empirical evaluations.
These questions extend beyond understanding individual models. They have direct implications for large-scale deployment, evaluation pipelines, and red-teaming practices in industry.
Connecting Theory With Model Evaluation
Better collaboration between theorists and experimentalists was a central theme of the workshop. Connecting these perspectives can help researchers better understand model performance.
Mathematical theory can contribute to understanding and predicting model generalization. At the same time, researchers are examining how to design empirical and open-world evaluations. Discussions also addressed uncertainty quantification and multimodal AI for real-world workflows.
Together, these areas can help researchers better understand model performance and strengthen the connection between theoretical predictions and empirical evaluation.
The workshop featured four keynote presentations, four selected oral presentations, and more than 100 accepted papers presented as posters.
Kitware’s Contributions Through AIQ
The connection between theory and empirical evaluation is also central to Kitware’s work on the DARPA AIQ program. As a Technical Area 2 (TA2) performer, Kitware leads test and evaluation efforts through the development of the Mathematical Assurance and Generative AI Network Evaluation Toolkit (MAGNET).
Within AIQ, Technical Area 1 (TA1) performers develop mathematical theories and models that predict the outputs of AI transformer models with formal guarantees. MAGNET evaluates these theories at scale. It empirically validates their predictions using relevant datasets and full-scale AI models.
Mathematical theories can provide predictions about model behavior. Empirical evaluation can then test whether those predictions hold at scale.
Several workshop organizers are also involved in the AIQ program, creating another opportunity for collaboration. Kitware’s work was highlighted during a keynote by DARPA Program Manager Patrick Shafto. The keynote included recent work using AI to autoformalize theories with the Lean language. This can help verify theories and create foundational building blocks for new ones.

From Research to Reliability
Building a more predictive science of foundation models will require continued collaboration between theorists and experimentalists. The workshop demonstrated how theoretical insights and structured evaluation can work together to better understand and guarantee model performance.
Kitware is contributing to this effort through the DARPA AIQ program and MAGNET, our open source system for evaluating guarantees on frontier model capabilities.
Contact us to learn more about Kitware’s work in AI test and evaluation.
Contact Us