18 August 2026

AI testing shifts from models to full system performance

  • Researchers are building testing frameworks that measure entire AI systems, not just individual models, including how tasks route between components and overall cost.
  • Hamel Husain released an eval-skills plugin demonstrating this approach. Agent Arena tested it against 1.7 million real-world task sessions.

How it was covered