3 September 2026

Researchers discover method to manipulate language models into harmful outputs

First reported

The Algorithm ran this on .

  • A new technique allows researchers to reliably trick large language models, the AI systems behind chatbots, into generating dangerous information they're designed to refuse.
  • The vulnerability was demonstrated by getting models to provide instructions on sabotaging aircraft navigation systems, a task they normally reject.
  • The research reveals how these models can be manipulated, providing insight into their internal workings and potential security gaps.

How it was covered

The AlgorithmMIT Technology Review

A new technique has revealed vulnerabilities in large language models, allowing researchers to probe deeper into how LLMs operate and making it easy to trick them into providing dangerous information such as how to sabotage aircraft navigation systems.