19 August 2026
Safety guardrails in open AI models removed in minutes
First reported
TLDR AI ran this on .
- Researchers demonstrated that refusal mechanisms, which prevent AI models from answering harmful questions, can be stripped away quickly through a technique called abliteration.
- Open-weight models are affected, meaning models whose code and weights are publicly released and anyone can modify.