18 August 2026

AI testing shifts from models to full system performance

  • Researchers are building testing frameworks that measure entire AI systems, not just individual models, including how tasks route between components and overall cost.
  • Hamel Husain released an eval-skills plugin demonstrating this approach. Agent Arena tested it against 1.7 million real-world task sessions.
  • These frameworks track practical outcomes like total completion cost and whether systems break tasks into steps correctly, rather than abstract benchmark scores.

How it was covered

Latent Spaceswyx & Alessio

The field is moving from model-level evals to harness-level measurement covering routing, decomposition, memory, verifier loops, and total completion cost, as exemplified by Hamel Husain's eval-skills plugin and Agent Arena's cost-per-task filters based on 1.7M real-world sessions.