17 August 2026updated 18 August
New benchmark tests AI's ability to discover hidden game rules
- Researchers created Dig.bench, a test with 70 text-based games where the rules are not explained upfront.
- The benchmark measures whether AI agents can figure out unknown rules through trial and error, like humans do.
- Current AI models struggle with the hardest games while humans solve them, showing a gap in this capability.
How it was covered
TLDR AITLDR editorial team
Dig.bench is a benchmark with 70 text-based games measuring whether agents can experiment to discover unknown rules. Humans solve even the hardest games while the best models struggle with top-tier games.