3 September 2026

Researchers discover method to manipulate language models into harmful outputs

First reported

The Algorithm ran this on .

  • A new technique allows researchers to reliably trick large language models, the AI systems behind chatbots, into generating dangerous information they're designed to refuse.
  • The vulnerability was demonstrated by getting models to provide instructions on sabotaging aircraft navigation systems, a task they normally reject.

How it was covered