Articles

LLM Reasoning Traces Stolen: Encrypted CoT Flaw

Researchers decoded 315,320 encrypted AI reasoning blocks from OpenAI, Anthropic and Google, recovering credentials and PII. How the reasoning-trace flaw works.

Chisato Chisato · · 6 min read
A red padlock resting on a laptop keyboard, representing a broken confidentiality boundary

The “thinking” that leading AI models keep hidden from users turns out not to have been private after all. In a paper titled Stealing Reasoning Traces from Proprietary LLM APIs — submitted to arXiv on August 10 and widely discussed across security channels since — researchers showed that the encrypted reasoning blocks returned by OpenAI, Anthropic, and Google can be decoded by feeding them to a weaker model in the same provider family. Scraping public logs alone, the team decoded 315,320 reasoning blocks and recovered 182 credentials and hundreds of pieces of personal data — from traces that were never supposed to be readable.

The finding cuts to a design decision at the heart of how modern reasoning models are served over an API. It is not a break of the encryption itself. It is a demonstration that the party the encryption was meant to lock out — a different session, a different user, a weaker model — was holding a working key all along.

How reasoning gets carried between calls

Reasoning models produce long internal chains of thought before they answer. Providers generally do not want to expose that raw reasoning to users: it can reveal proprietary training signals, safety-relevant content, or intermediate steps a competitor could distill. But the reasoning still has to persist across turns so a multi-step agent can continue where it left off.

Rather than store those traces server-side, the providers return them to the client as blocks of encrypted text. The client passes the block back on the next call, and the provider decrypts it to reconstruct the model’s prior thinking. To the user, the block is opaque — a sealed container meant purely for conversation continuity.

The flaw, according to the researchers, is in how those containers are keyed. Within each provider’s ecosystem, the encrypted blocks are authenticated and encrypted with a key shared globally across sessions, users, and models. That makes the blocks fully interchangeable: a block produced in one session can be replayed into a completely different session, belonging to a different user, and the provider will accept and decrypt it. The container was built to be portable for the legitimate client — but nothing tied it to that client.

Turning a weaker model into a decoder

Portability by itself leaks nothing, because the plaintext is still supposed to stay inside the provider. The second half of the attack extracts it. The researchers found that if you inject a reasoning block generated by a strong model — say Claude Opus — into a weaker, compatible model in the same family, the weaker model will happily transcribe the hidden reasoning verbatim in plaintext.

In effect, the weak model acts as a “fuzzy decoder.” The paper describes using Claude Haiku 4.5 to decode Claude traces, GPT-5.6 Luna for OpenAI traces, and a Gemini model for Google traces. Because the block decrypts correctly for any model in the family, and because a smaller model is easier to steer into simply repeating what it was given, an attacker needs only two API calls: one to submit the captured block, one to read back the reasoning it was never meant to see. There is no need to touch the cryptography — the system decrypts the content willingly and then a cheap model reads it aloud.

What came out of the public logs

To measure real-world exposure, the team did not attack live users. They scraped 6,708 publicly available agent-trajectory logs from GitHub and Hugging Face — the kind of debugging and evaluation dumps developers routinely publish — and ran the decoder against the encrypted blocks embedded in them. From 315,320 decoded reasoning blocks, they recovered:

  • 367 artifacts of personally identifiable information (PII), and
  • 182 credentials — including material like API keys and passwords — from genuine user sessions.

The point is that these secrets were already sitting in the open. Developers who shared a trajectory log assumed the encrypted reasoning inside it was inert. Instead, every one of those blocks was a decodable record of whatever the model reasoned over — including secrets a user had pasted into a session and assumed the “hidden” chain of thought would keep hidden.

Four ways to abuse it

The paper lays out four distinct abuse paths, and they range well beyond passive snooping:

1. Model distillation. Competitors can harvest a frontier model’s raw reasoning at scale — exactly the proprietary signal the encryption was meant to protect — and use it to train or improve their own systems.

2. Cross-user data theft. As the public-log experiment showed, private data and credentials inside other users’ published traces become readable.

3. Recovering concealed harmful content. A model can return a safe-looking visible answer while its hidden reasoning contains material the safety layer meant to suppress; decoding the block recovers what the polished answer omitted.

4. Invisible prompt injection. Attackers can embed malicious instructions entirely inside an encrypted reasoning block and seed them into public agentic datasets. Anyone who later replays that block is carrying a payload they cannot see — a stealthy variant of prompt injection hidden behind opacity that was supposed to be a security feature.

The vendor response

Following responsible disclosure, OpenAI, Anthropic, and Google deployed server-side patches that close the demonstrated attack paths. Because the fixes are server-side, developers do not need to change their own code to benefit from the mitigation for new sessions.

There is a large caveat, though: the patches do not retroactively protect historical logs. Reasoning blocks that were generated before the fix and already published — in the very GitHub and Hugging Face repositories the researchers scraped — may remain decodable. The confidentiality of anything already exposed cannot be restored by a server-side change; those blocks are in the wild, and the secrets inside them should be treated as compromised.

What it means

This is a confidentiality failure rooted in a reasonable-sounding design choice: keep reasoning portable by shipping it to the client as an encrypted, self-contained block. The trouble is that “portable across sessions, users, and models” is indistinguishable from “usable by an attacker who obtains the block” when a single global key underwrites the whole scheme. Encryption without binding — to a user, a session, a single model — is a lock whose key everyone in the building already holds.

Who is exposed. Any developer who published agent logs, evaluation runs, or debugging traces containing encrypted reasoning blocks — and any user whose secrets passed through a reasoning model before the patch and ended up in one of those logs. The immediate defensive action is unglamorous but urgent: rotate credentials that may have been present in shared traces, and scrub reasoning blocks from any logs still public. It is the same hygiene lesson that follows every real-world credential-exposure incident, applied to a new surface.

Why it matters beyond this bug. The industry has leaned on opacity as a security property — hide the chain of thought, and its contents are safe. This work shows opacity is not confidentiality. The prompt-injection variant is the sharpest illustration: a boundary marketed as protective became a channel for smuggling attacks precisely because no one could see inside it. As more of the ecosystem runs on multi-step agents that pass reasoning between calls — and as those traces increasingly move through shared caches and context-persistence layers — the way that intermediate state is keyed and bound becomes a first-order security question, not an implementation detail.

What to watch next. Whether providers move from a single global key toward per-user or per-session binding of reasoning blocks, so a captured block is useless outside the context that created it. Whether platforms that host public agent datasets start stripping or flagging embedded reasoning blocks the way they already scan for leaked keys. And whether this becomes a recurring class of finding — much as agent sandbox escapes did — as researchers keep probing the plumbing that carries state between AI calls. The reasoning was hidden. It was never sealed.

Chisato Chisato · · 6 min read

LiteLLM CVE-2026-59822: CISA KEV AI Infra Attacks

CISA added seven exploited flaws to its KEV catalog on Sept. 2, and three target AI infrastructure — LiteLLM, Kestra, and Starlette. What to patch and why it matters.

#Security #Vulnerability #AI
Chisato Chisato · · 5 min read

SharedRoot: Claude Cowork Sandbox Escape Explained

Researchers show how a single message can push Claude Cowork's AI agent out of its Linux VM to read a Mac's SSH keys and cloud credentials. The SharedRoot chain, explained.

#Security #AI #Vulnerability