Four words that often get blurred together

An AI company can say a system was thoroughly evaluated and still leave an obvious question unanswered: evaluated how? A quick benchmark, a security review, a user study and a field trial are all different things, even when they are filed under the same reassuring word.

The National Institute of Standards and Technology, or NIST, is trying to make that difference more explicit. Its new draft is called the TEVV-Athlon Framework. TEVV stands for test, evaluation, verification and validation.

The name is a little unwieldy. The underlying point is not. A useful assessment should say what was measured, why it was measured and how the result connects to the system a person will actually use.

What the draft proposes

NIST’s initial public draft describes a four-stage way to build customised AI assessments. It uses the term TEVV-Athlon for an assessment made from Events and Tools that produce data about selected measurement concepts, which the document calls Blocks.

The agency says this is meant to be extensible, adaptable and customisable. It names conventional statistical models, large language models, multimodal systems and agentic systems as examples of the kinds of AI that may need different evaluation approaches.

That flexibility is deliberate. A single score does not tell a hospital, a public buyer, a researcher and a consumer app team the same thing. The draft is trying to give those groups a shared way to describe their method without pretending their risks are identical.

Why a benchmark is not the whole story

Benchmarks are useful. They make comparison possible and can expose a clear weakness. But a benchmark result alone does not show whether a system meets an organisation’s goals in the setting where it will be deployed.

NIST frames TEVV as evidence that a system can effectively meet individual or organisational goals while minimising negative impacts. Its draft says the framework can help organisations produce information about system performance and measure real-world impact and outcomes.

That is not a promise that the framework will catch every failure. It is a request to connect technical checks with the conditions that matter outside a test set, including the context, the people affected and the limits of the evidence.

A draft is not a rulebook

NIST has not created a new binding requirement. The document is an initial public draft, and the agency is asking for comments on its definitions, scope, gaps and practical usefulness until 6 October.

The questions NIST asks are revealing. It wants to know whether the framework covers different kinds of AI evaluation well enough, whether it helps with novel systems, and which concepts need clarification or revision. Those are signs of a framework that is still being shaped in public.

For organisations, the immediate value may be modest but real. It offers a common prompt before a launch or a procurement decision: which claim are we testing, which evidence would count, and what outcome would change our mind? That is harder to hide behind than a generic claim of rigorous evaluation.

What is confirmed, what NIST proposes, and what is open

Confirmed: NIST announced the TEVV-Athlon initial public draft on 7 August 2026 and opened a 60-day comment period that ends on 6 October. The published draft presents a four-stage framework for customised AI assessments.

NIST’s proposal: the framework can be applied across a broad range of AI systems to produce meaningful performance information and help organisations assess real-world impact and outcomes. That is the agency’s stated aim, not a demonstrated result from broad deployment.

Open questions: which definitions will survive public comment, how much work a credible TEVV-Athlon will require, whether organisations will use it in practice, and how well the final framework handles the fast-changing behaviour of large language models and agents.

Sources

  1. NIST — The TEVV-Athlon Framework for Evaluating AI SystemsPrimary NIST announcement, dated 7 August 2026. Source for the framework’s scope, stated purpose and public-comment deadline.
  2. NIST AI 200-2 (Initial Public Draft) — The TEVV-Athlon Framework for Evaluating AI SystemsPrimary draft report. Source for the four-stage process, Events, Tools and Blocks terminology, and the draft’s technical framing.