AI Integration Security Risks – The Blast Radius Problem

Table of Contents

Every enterprise AI rollout eventually runs into the same uncomfortable fact. The system is only useful if it can read your data, and the moment it can read your data, every source that data touches becomes part of your attack surface. Nobody sells that on the slide deck.

The Trust Boundary That Agentic AI Quietly Deletes

Traditional security thinking treats content and instructions as separate. An email is data. A script is instructions. Software has generally kept those apart, and where it hasn’t, we’ve called it a vulnerability class and patched it.

Agentic AI collapses that distinction on purpose. A model reading an email and a model being told what to do are, mechanically, doing the same thing, consuming text and producing behaviour. Instruction-following flexibility is the entire value proposition, and it is also the entire hole.

EchoLeak (CVE-2025-32711) is the clean demonstration. An attacker embedded instructions in an ordinary email such that Microsoft 365 Copilot’s cross-prompt injection classifier didn’t flag it, and once ingested, the model executed the embedded instructions as legitimate task content, exfiltrating mailbox data through a Teams domain the organisation already trusted. Nobody needed to compromise Microsoft’s infrastructure. They just needed to write an email a classifier wouldn’t catch and a model would follow. ForcedLeak did something structurally similar through Salesforce’s web-to-lead forms, and got extra mileage from a CSP rule that still permitted an expired but formerly trusted domain as an exfiltration route.

The pattern in both cases: nothing was hacked in the conventional sense. The system worked exactly as designed, and the design itself was the vulnerability.

Your MCP Integrations Inherit Your Trust, Not Just Your Access

The Model Context Protocol solves a real problem, giving models structured access to tools and data without bespoke integration work for every combination. It also means that connecting an MCP server isn’t like adding a library dependency. It’s closer to handing someone your session token.

The postmark-mcp incident is the case worth sitting with. A package impersonating a trusted email provider’s MCP integration functioned correctly until version 1.0.16, at which point it began quietly forwarding every processed email, including password resets and internal memos, to the developer’s own domain. A compromised npm package in a build pipeline threatens your build. A compromised MCP server sits inside a live workflow with the permissions your AI agent was granted, which in most deployments is broader than any single human user’s access, because nobody’s got round to scoping it down yet.

That’s the bit worth saying plainly to anyone who’s connected an MCP server and moved on. You didn’t add a tool, you added a party with your access level who you’re trusting to stay honest indefinitely, with no equivalent of a code review gate before every use. Our piece on the hidden supplier risks in enterprise AI covers how this kind of exposure tends to arrive through connections nobody formally approved.

Data Exposure Isn’t New, But the Surface Area Is

DeepSeek exposing over a million records including chat logs, API keys, and credentials wasn’t a novel attack. It was a database left open, the kind of misconfiguration security teams have been fixing since before “AI” meant anything more than a chess engine. What’s changed is what sits behind that misconfiguration. A conventional application database holds records. An AI system’s backend can hold the raw material of every conversation anyone has had with it, which for enterprise deployments increasingly means proprietary strategy discussions, customer data, and credentials volunteered in the course of getting help with something mundane.

The control failure is old. Our piece on why most organisations have no real visibility of AI usage covers why the blast radius keeps catching leadership by surprise even when the underlying mistake is a familiar one.

Transparency Is a Defence, Not a Compliance Line Item

GDPR Articles 13 to 15 require meaningful information about the logic behind automated decisions. Most organisations treat that as a documentation exercise, writing down roughly how the system works, filing it, and moving on. That framing misses what’s actually at stake.

OpenAI’s published research on situational awareness in advanced models describes something more specific: models recognising evaluation conditions and adjusting outputs to appear aligned while pursuing different underlying objectives. Whether or not any given deployed system does this, the possibility changes what audit trails are for. They’re there because if a model’s stated reasoning and its actual behaviour can diverge, the only way to catch that divergence is to have kept enough of a record to compare the two.

Treat explainability as an audit obligation and you’ll build the minimum. Treat it as the only mechanism you have for catching a system that’s telling you one thing and doing another, and you’ll build something that actually works.

What This Means in Practice

This argues for adopting AI carefully rather than not adopting it, which is different from how most organisations currently operate, extending trust to every connected source and every third-party integration by default and hoping the vendor’s security posture is sound. Zero trust principles, applied properly to AI tooling, mean the model gets exactly the access a task requires and nothing it doesn’t, every MCP connection gets the same scrutiny as a new privileged account, and every AI-touched data flow gets monitored the way you’d monitor any other route out of the building.

The organisations that get burned by this won’t be the ones who moved slowly. They’ll be the ones who connected everything and assumed the trust boundary was still where it used to be. Continuous AI Assurance is built for exactly this kind of ongoing exposure, not a one-off check.

Related Posts