18 August 2026
New benchmark tests AI models on learning hidden rules through exploration
- Researchers created DiG-bench, a test of 70 text-based games measuring whether AI systems can figure out unstated rules by trying things out.
- Anthropic's Claude Opus 5 and a model called Fable 5 performed best. Most current leading AI models failed the hardest challenges.