18 August 2026

New benchmark tests AI models on learning hidden rules through exploration

  • Researchers created DiG-bench, a test of 70 text-based games measuring whether AI systems can figure out unstated rules by trying things out.
  • Anthropic's Claude Opus 5 and a model called Fable 5 performed best. Most current leading AI models failed the hardest challenges.

Where they differ