3 September 2026
Researchers discover method to manipulate language models into harmful outputs
First reported
The Algorithm ran this on .
- A new technique allows researchers to reliably trick large language models, the AI systems behind chatbots, into generating dangerous information they're designed to refuse.
- The vulnerability was demonstrated by getting models to provide instructions on sabotaging aircraft navigation systems, a task they normally reject.