You can access the Network best practice document here.
Advanced AI systems are being built and deployed across borders, languages and industries. To develop and deploy AI well, robust ways to evaluate and understand it are needed more than ever. Trust in AI and adoption both hinge on a shared understanding of why AI systems behave as they do. As AI diffuses globally, so too must our approach to evaluation.
That’s why the International Network for Advanced AI Measurement, Evaluation and Science (NAAIMES) – formerly the International Network of AI Safety Institutes – continues to drive forward the science of AI evaluation. Established in 2024, the Network brings together Institutes globally, in Australia, Canada, the European Union, France, Japan, Kenya, the Republic of Korea, Singapore, the United Kingdom and the United States.
This year, on the margins of the 2026 International Conference on Machine Learning, in Seoul, the Network gathered, alongside industry leaders and civil society organisations. Network members discussed the latest in the science of AI evaluation: showcasing evaluation tools (such as AISI’s newly released Engineering Playbook), raising open questions and challenges encountered in evaluating AI agents, and mapping out current and future best practices.
Sharing best practice
The event in Seoul coincides with the completion of the Network’s first best practice guidance document.
Network best practices are designed to support the growing ecosystem of third-party evaluators, who play a crucial role in promoting trusted adoption of AI.
High-quality measurement science is key to governments’ and businesses’ ability to assess AI for a wide spectrum of needs. Agreeing this between institutes with leading expertise in the field should enhance comparability between evaluations and so strengthen the science of evaluations.
This guidance builds on and is intended to complement existing best practice publications – most notably NIST AI 800-2 – while providing specific recommendations to third-party evaluators, such as on capability elicitation. It details:
- The importance and process of carefully defining evaluation objectives and selecting the appropriate benchmarks.
- How to ensure comparability between evaluations with attention given to both the generation of outputs (such as inference settings) and the analysis of outputs (e.g. scoring criteria).
- How to iterate on capability elicitation in order to best measure how the model performs at a relevant to your evaluation objectives.
- The processes and issues to work through in conducting evaluations and tracking results, from appropriate logs to common debugging tips.
Addressing open questions in AI evaluations
Around the themes of the Network, members also have published on the following topics, building on the open questions raised by our February 2026 blogpost:
- The Singapore AI Safety Institute authored ‘How should evaluators test AI systems (as opposed to models)?’, discussing how to best develop testing that mirrors real-world interaction with an entire AI system (rather than just the model layer). It examines identifying the relevant risks, building realistic test environments, constructing representative datasets, choosing metrics that reflect both outcomes and how they are reached, and combining rule-based, LLM-based and human evaluation. You can read the Singapore AI Safety Institute’s blog in full here.
- The Canadian Artificial Intelligence Safety Institute authored ‘What information should AI evaluators share?’, a case for maximum transparency. The piece addresses the careful balance in clearly stating an evaluation's intent and methodology to build trust and comparability, while maintaining ‘selective disclosure’ - withholding sensitive details such as CBRN (chemical, biological, radiological, or nuclear hazards) or cyber datasets or adversarial jailbreak insights to avoid misuse and data contamination. Read the full blogpost here.
- The French National Institute for AI Evaluation and Security (INESIA) published ‘What should evaluators prioritise when resources are limited?’ In managing limited time, compute, access and expertise, the paper recommends first determining decision-relevance: being clear what decision the work will inform, for whom, and what results would actually change the next step. Starting with lightweight testing allows quick determination of whether signals exist to go more in-depth. Limitations should also be made clear in any reporting. You can read the full blog here.
Looking forwards
As AI is developed and deployed globally, sustaining that foundation will require continued coordination across Institutes, standards bodies and jurisdictions – precisely the collaboration the Network exists to advance.
After gathering in Seoul, Network members will align on further best practices that address common issues at the forefront of evaluation science, with particular attention given to agentic capabilities.