17 August 2026updated 18 August

New benchmark tests AI's ability to discover hidden game rules

  • Researchers created Dig.bench, a test with 70 text-based games where the rules are not explained upfront.
  • The benchmark measures whether AI agents can figure out unknown rules through trial and error, like humans do.
  • Current AI models struggle with the hardest games while humans solve them, showing a gap in this capability.

How it was covered

TLDR AITLDR editorial team

Dig.bench is a benchmark with 70 text-based games measuring whether agents can experiment to discover unknown rules. Humans solve even the hardest games while the best models struggle with top-tier games.