Find the failure before writing the rule

An AI support agent should not retrieve one customer's details for another. That sounds like a simple requirement. In practice, the request can arrive directly, under a claim of authority or after several turns of conversation. Microsoft published run-assert-eval on September 24 to connect three jobs that teams often handle separately: finding such risks, testing whether they happen and putting a rule at the point where the agent acts.

The open-source workflow starts with Clarity, Microsoft's agent threat-modeling project. A person selects which identified risks deserve testing. ASSERT then builds evaluations for those specific behaviors. The proposed control is written in Agent Control Specification, or ACS, so it can be applied during the agent's operation. The developer still has to review the generated policy and its placement before running it.

The same test matters

Microsoft's example is a billing agent assigned to one account. In its baseline run, the agent exposed another customer's information in 12 of 40 applicable conversations, or 30%. The team put a deterministic rule before a tool call to reject a request for any other account. It also checked the tool's result before that data could return to the model.

The governed agent was then measured with the same behavior definition, test cases and judging approach. Microsoft reports two cross-customer violations in 34 applicable conversations, or 5.9%. That is a useful demonstration of a narrower failure rate, not a clean bill of health. The denominators differ because the rates count applicable conversations, the samples are small and two violations still occurred. The company also reports no legitimate-request violations in this sample; that cannot establish how the rule would behave for every real customer.

Why this is more than another safety score

A safety score alone does not stop a tool from reading a record. The more concrete idea here is to turn a measured failure into a rule at the tool boundary, then check whether that rule blocks the bad action without also blocking useful work. It is possible to make an agent appear safe by refusing everything. Microsoft tracks impermissible actions and permissible requests separately to make that trade-off visible.

The public repositories contain the skill, the underlying projects and a worked billing-agent example. The tools are available for developers to inspect, but applying them well still requires choosing meaningful risks, reviewing policy and testing the actual agent configuration. Microsoft's post describes this as an early practice it wants to make into a repeatable release gate.

What remains unproved

This was a controlled demonstration supplied by the toolmaker. It did not measure a deployed billing service, independent human review of every outcome or the cost of running larger evaluations. Its automated judge may make mistakes, and risks omitted during discovery can remain outside the test entirely.

For a team adopting an agent, the practical question is whether it can name a forbidden action, show the test that exposes it, point to the control that prevents it and rerun the same test after a change. This release gives developers one way to do that. It does not remove the need to keep looking for failures the first test missed.

Sources

  1. Microsoft Command Line: Introducing run-assert-evalPrimary September 24 release, workflow, billing-agent results, limitations and repository links.
  2. Microsoft: Agent Control SpecificationPrimary documentation for runtime interception and policy contract.
  3. Microsoft Responsible AI: ASSERT repositoryPrimary open-source repository containing the evaluation toolkit and run-assert-eval skill.