Back to Blog
·5 min read·
McKinseyCodeWallAI securityagent securityred teamLilliAuthorized Intent ChainOWASP

An AI Agent Hacked McKinsey in 2 Hours — And Read 46 Million Private Messages

In March 2026, CodeWall's AI agent broke into McKinsey's internal platform Lilli and exposed 46M private messages in just 2 hours. Nobody listened — then the July 2026 wave of AI agent attacks proved it was the canary in the coal mine.

Part of our AI Agent Security Series — retrospective on the March 2026 breach that predicted everything that happened in July.

In March 2026, a security startup called CodeWall built an AI agent that did the unthinkable: it hacked McKinsey & Company’s internal AI platform, Lilli, in under two hours. The agent gained full read and write access to the production database, exposing 46 million private messages between McKinsey consultants and clients.

McKinsey had described Lilli as a tool that could “rewire the way we operate.” Within 120 minutes, an offensive AI agent rewired that confidence into a security nightmare. This wasn’t a sophisticated zero-day exploit. It was a clean, automated bypass of permission logic — executed entirely within normal access boundaries.

The Age of “AI Attacks AI” Has Arrived

Traditional red teams take days or weeks to find vulnerabilities. CodeWall’s agent did it in two hours. It used generative AI to probe, reason, and pivot through Lilli’s interface at machine speed — faster than any human pentester could ever hope to match.

The agent didn’t use brute force or malware. It operated inside the platform’s normal permission model, exploiting a chain of authorized intents. This is the same Authorized Intent Chain concept that Tenet Security later documented: the agent never broke a rule — it just automated the loopholes.

Attack surfaces that once took days now take hours. And defensive tooling hasn’t caught up. The McKinsey hack was an early warning that, four months later, would be followed by a wave of attacks — Agentjacking, Friendly Fire, GhostApproval, and Mozilla’s 0DIN DNS attack. We covered those in detail, but the McKinsey breach was the canary in the coal mine.

What Exactly Happened?

CodeWall’s agent started by exploring Lilli’s API endpoints with a standard reconnaissance script. Within 15 minutes, it discovered that the platform’s access control relied on a session token that didn’t properly scope user roles. The agent then generated a series of queries that progressively escalated to database-level read and write operations.

After two hours, it had dumped the entire message history — 46 million records — and left behind a backdoor for persistent access. The agent itself wrote the final exploit payload, debugging its own errors along the way. No human intervened.

The Register reported that McKinsey patched the vulnerability within hours of notification, but the damage was already done: sensitive client communications, strategy documents, and proprietary analysis were exposed.

Why Traditional Security Misses This

Most security tools treat AI agents as black boxes — they monitor inputs and outputs but ignore intermediate reasoning steps. The CodeWall agent didn’t trigger any of McKinsey’s existing alerts because it never made a suspicious request. It simply chained together multiple legitimate actions.

This is precisely what the AGentShield benchmark later confirmed: tool abuse detection is “weak across the board.” The benchmark tests 6 commercial AI agent security tools across 537 test cases. Even state-of-the-art agents failed to spot simple tool abuse chains. Several providers that catch >95% of prompt injections miss most unauthorized tool calls.

BankInfoSecurity noted that the agent used a technique similar to “SQL injection via natural language prompts” — but far more subtle.

The MCP Attack Surface Connection

The McKinsey hack exploited the Model Context Protocol attack surface — the interaction layer between an AI model and its external tools. When an agent can read, write, and execute on databases, every permission becomes a potential weapon.

We previously documented seven specific attack types targeting the MCP layer. The McKinsey case maps directly to what we called “agent privilege escalation via chained API calls” — the agent never escalated privileges, it just discovered that “legitimate” access was broader than anyone realized.

Nobody Listened — And Then July Came

When the McKinsey hack first broke, many dismissed it as a one-off red-team exercise. “It’s just a test,” they said. “Real attackers won’t use AI agents that way.”

Four months later, the July 2026 wave of AI agent attacks proved otherwise. Agentjacking exploited Sentry DSNs to hijack Claude Code and Cursor. Friendly Fire bypassed safety classifiers with multi-step chains. GhostApproval defeated approval dialogs. And Mozilla’s 0DIN showed that even clean GitHub repos could trick agents into reverse shells.

Each exploited the same class of vulnerability — Authorized Intent Chains — at scale. Enterprises that had ignored the McKinsey warning were caught flat-footed.

The pattern is clear: offensive AI agents are evolving faster than defensive tooling. What took CodeWall two hours in March can now be done in minutes.

What You Can Do Now

Your AI platforms — whether internal chatbots, customer-facing agents, or coding assistants — have attack surfaces you haven’t mapped. The McKinsey breach shows that privilege escalation via authorized intents is not theoretical. It’s an OWASP-level vulnerability for the agent era.

Start with a professional AI agent security audit. We at dotfm.me specialize in identifying chained permission vulnerabilities before attackers do. We’ll simulate an offensive AI agent against your environment and deliver a prioritized fix list — in hours, not weeks.

Don’t wait for the next wave. Contact us to schedule your AI agent security audit.

Is your AI-built app ready for real users?

We audit, harden, and ship AI-built apps. From security review to production deployment.

Get an audit