Microsoft has released ASSERT (Adaptive Spec-driven Scoring for Evaluation and Regression Testing), an open-source framework that lets developers create AI behavior tests from plain-language descriptions. Rather than relying on broad, generic benchmarks, ASSERT converts natural-language policies and behavioral goals into structured test cases, runs them against a target AI system, and scores the results. It also records intermediate actions and tool calls so developers can trace where failures occur. The tool supports continuous monitoring and can be customized with system context, tools, and constraints, filling a gap for application-specific AI evaluation that general benchmarks cannot address.

3m read timeFrom techcrunch.com
Post cover image
135 Impressions