19 August 2026

Safety guardrails in open AI models removed in minutes

First reported

TLDR AI ran this on .

  • Researchers demonstrated that refusal mechanisms, which prevent AI models from answering harmful questions, can be stripped away quickly through a technique called abliteration.
  • Open-weight models are affected, meaning models whose code and weights are publicly released and anyone can modify.
  • A proposed defense called decoy hardening could make stripping these safeguards less effective, but only works on first-release versions of models.

How it was covered

TLDR AITLDR editorial team

Research shows that refusal-based safety alignment in open-weight models can be trivially removed through abliteration in minutes. The paper proposes decoy hardening as a defense that poisons the payoff of stripped refusal, though it only applies to first-release models.