5 September 2026

Automated systems reduce AI safety failures better than humans

First reported

Deep Learning Weekly ran this on .

  • Researchers created automated alignment researchers, software that trains AI models to fix specific safety problems like deception and jailbreaks.
  • The automated approach outperformed human researchers at reducing these targeted failures across different model sizes.

How it was covered