As Simon Willison notes, this vulnerability, while now fixed, reveals a deeper, persistent tension in the AI industry. The very act of hiding reasoning—a selling point for proprietary models—created a new class of risk when the encryption was poorly implemented. The episode suggests that as models become more complex and their internal states more valuable, securing the entire inference pipeline, not just the input and output, will be critical. It also hints that 'reasoning' as a product feature may introduce unforeseen attack surfaces that simpler, stateless completions do not.
AI model providers said to have patched reasoning trace vulnerability
Researchers reportedly extracted hidden reasoning from proprietary LLMs by replaying encrypted internal thought tokens.
AIpressr commentary on an article originally published by Simon Willison.
For informational purposes only. AI-assisted commentary may contain errors. full disclaimer ↓hide ↑
This is AIpressr's editorial commentary on a report originally published by another outlet — it is opinion, not the original reporting, and not an endorsement by or affiliation with that outlet. Follow the linked source for the underlying facts. Editorial & AI disclosure.
Editor's Take
Simon Willison highlights a research paper revealing a now-patched security flaw in how major AI providers handled internal reasoning traces. According to the report, encrypted 'chain-of-thought' blocks could be fed back into weaker models to force them to output the hidden reasoning of more powerful ones. This incident underscores the ongoing, opaque security challenges in proprietary AI systems, where even internal mechanics can become attack vectors if not meticulously designed.
“The paper's authors found that every model under the same family used the same encryption key, which meant you could feed those blocks back into the weakest model family members and jailbreak them into outputting the unencrypted raw reasoning blocks!”
Our analysis
Have AI news to share?
Submit your release →Publisher or subject of this story? Object to this commentary or request a correction →
