5 September 2026
Automated systems reduce AI safety failures better than humans
First reported
Deep Learning Weekly ran this on .
- Researchers created automated alignment researchers, software that trains AI models to fix specific safety problems like deception and jailbreaks.
- The automated approach outperformed human researchers at reducing these targeted failures across different model sizes.