18 August 2026

Researchers release benchmark testing AI's ability to discover hidden rules

First reported

Import AI ran this on .

  • DiG-bench is a set of 70 text-based games measuring whether AI can figure out unstated rules through trial and error instead of being told.
  • Anthropic's Claude Opus and a model called Fable 5 outperformed other AI systems, but only these two solved any of the hardest difficulty tasks.

How it was covered