Articles

Microsoft Copilot CoSnitch Flaw (CVE-2026-24301)

CoSnitch let one click on a link make Microsoft Copilot exfiltrate a victim's Gmail and Drive data. How the chained flaw worked and why Varonis called it meta-hacking.

Chisato Chisato · · 6 min read
An open padlock resting on a laptop keyboard, symbolizing a bypassed access control

Another week, another AI assistant that could be turned against the person it was built to help. Researchers at Varonis Threat Labs disclosed a critical vulnerability in Microsoft Copilot Personal that let an attacker silently pull data out of a victim’s connected accounts with nothing more than a single click on a malicious link. Microsoft tracked the issue as CVE-2026-24301, assigned it a severity score of 8.8 out of 10, and shipped a patch on August 18, 2026. Varonis nicknamed the finding CoSnitch — because, as the researchers tell it, Copilot ultimately snitched on itself.

The flaw is notable for two reasons. The mechanics are a clean illustration of the failure mode security teams have been warning about since AI assistants started reading untrusted content. And the discovery method — the researchers coaxing the model into explaining its own internals — points to a class of attack that does not fit neatly into any existing playbook.

What CoSnitch did

Microsoft Copilot Personal is the consumer-facing assistant that users can wire up to their own accounts — Gmail, Google Drive, Google Calendar, and the like — plus Copilot’s own memory and chat history. The value proposition is convenience: ask Copilot a question and it can reach into your connected services to answer. CoSnitch turned that same reach into an exfiltration channel.

According to Varonis, CoSnitch was not a single bug but a chain of three separate weaknesses that, stitched together, produced a one-click compromise:

  1. An undocumented URL parameter. Combined with Copilot’s ordinary query function, this parameter allowed an attacker-crafted prompt to execute automatically on page load — no typing, no confirmation, no “are you sure?” The victim only had to open the link.
  2. Action within the authenticated session. Once the injected prompt fired, it ran with the victim’s own permissions. Copilot could query the connected apps — reading email, files, calendar entries, memory, and past chats — because from the system’s perspective the request came from the legitimate signed-in user.
  3. Exfiltration via Copilot’s own fetching. The stolen data was then funneled to an attacker-controlled server through Copilot’s built-in URL-fetching feature, disguising the theft as routine outbound traffic rather than an obvious data leak.

Put together, the sequence is a textbook case of indirect prompt injection: untrusted input (the crafted link) becomes an instruction the AI faithfully carries out, and the AI’s legitimate capabilities — account access, web fetching — become the attacker’s tools. No password was stolen, no permission was bypassed in the traditional sense, and no access control was tripped. The system did exactly what it was told; the problem was who got to tell it.

Meta-hacking: when the model maps its own attack surface

The more unusual part of the story is how Varonis found the undocumented parameter in the first place. Rather than reverse-engineering Copilot’s code, the researchers asked Copilot about itself — repeatedly.

When they suggested that automatic prompt execution should be possible, Copilot pushed back, explaining why such a thing was supposedly infeasible. But each refusal came with reasoning. By treating those explanations as leads and reframing them as follow-up questions, the researchers got the assistant to progressively describe its own behavior and architecture — and, eventually, to reveal the exact undocumented parameter that made the attack chain work. Copilot’s own effort to prove the exploit couldn’t happen is what handed the researchers the piece they needed to make it happen.

Varonis calls this technique meta-hacking: social-engineering the model’s reasoning process instead of attacking its code. It is a genuinely new wrinkle. Traditional vulnerability research treats the target as an opaque binary to be probed from the outside. Here the target is a system whose entire job is to explain things helpfully — and that helpfulness, pointed inward, becomes a reconnaissance tool against itself. The same instinct that makes an assistant useful makes it a leaky source of information about its own guardrails.

Timeline and exploitation

The disclosure timeline is its own subplot. Varonis says it reported CoSnitch to Microsoft in December 2025, and the fix did not ship until August 18, 2026 — roughly eight months later, a gap that drew pointed commentary given the flaw’s severity and the sensitivity of the data it exposed. Microsoft says it has found no evidence of active exploitation before the patch landed, which is the reassuring part: as far as the vendor can tell, this was caught and closed by researchers rather than discovered first by attackers.

That is the difference between a disclosure and an incident, and it matters. There is no breach notification here, no known victims, no data confirmed stolen in the wild. What there is instead is a worked example — a fully documented, one-click, zero-interaction path from a link to a victim’s inbox — that is now public. The patch closes this specific chain. It does not close the category.

A pattern, not a one-off

CoSnitch is the latest entry in a fast-growing catalog of attacks that exploit AI assistants wired into real corporate and personal data. Just this month, the same Varonis team detailed RovoBlast, a technique that tricked Atlassian’s Rovo assistant into leaking Jira and Confluence contents; agentic coding workflows have been shown vulnerable to injected instructions hiding in repositories. The through-line is consistent: give an AI system broad, authenticated access to data and the ability to reach the network, and any attacker who can slip instructions into what the AI reads inherits a slice of that access.

The defensive lesson is not “turn off Copilot.” It is that connected AI assistants expand the attack surface in ways classic controls were never designed to see. A zero-trust posture — treat every request as untrusted until verified, and never let a session’s identity alone authorize sensitive data movement — assumed the actor issuing requests was a human or a known service. An assistant that executes attacker-supplied prompts under the user’s own identity blurs that line. Guardrails have to move down to where the data actually leaves: constraining which connected services an assistant can read, scrutinizing where it is allowed to fetch, and logging its outbound traffic the way any other potential exfiltration channel would be logged.

What it means

For Microsoft Copilot users, the immediate takeaway is simple: the specific hole is patched, there is no sign it was exploited, and no user action beyond staying current is required. This one is closed.

The larger meaning is where the story earns attention. CoSnitch is a clean, public demonstration that the convenience features selling AI assistants — one-click actions, connected accounts, built-in web fetching — are precisely the primitives an attacker needs for silent data theft. The winners from this disclosure are defenders and researchers, who now have a concrete, named example to point to when they argue that AI assistants need their own security model rather than a bolt-on of old assumptions. The exposed party is any organization rolling out connected AI to employees on the belief that vendor guardrails are sufficient.

Two things are worth watching. First, meta-hacking as a method: if models can be talked into mapping their own attack surface, every helpful, explanatory assistant is a potential informant about its own weaknesses, and vendors will need to think about what their systems disclose when asked about themselves. Second, the eight-month patch gap: severity-8.8 flaws that sit for the better part of a year invite scrutiny of how AI-assistant vulnerabilities are triaged, especially as these tools graduate from novelty to default. The category of attack — untrusted input becoming trusted instruction inside a system holding your data — is not going away with this patch. It is becoming the defining security problem of the assistant era.

Chisato Chisato · · 6 min read

Atlassian Rovo Vulnerability: RovoBlast Data Leak

Researchers showed Atlassian's Rovo AI could be tricked into leaking Jira and Confluence data via prompt injection. Here's how RovoBlast worked.

#Security #AI #Prompt Injection
Chisato Chisato · · 5 min read

GitLost: GitHub AI Agent Leaks Private Repos

Researchers say a single crafted GitHub Issue could trick GitHub's Agentic Workflows into posting private repository contents publicly. Here's how GitLost works.

#Security #AI #GitHub