AI Agents Are Escaping Guardrails and Attacking Over the Internet. Are You Breach Ready Yet?

table of contents

Visualize this.

Until July 2026 happened, cyberattacks were continuously increasing despite massive investments in cybersecurity. July 2026 will be known as the month when even the biggest names in AI, viz. Anthropic and OpenAI had to admit they could not keep their rogue AI agents contained.

Their autonomous agents broke out of the lab and went after other companies. The walls did not hold.

As a result, CISOs all over the world are realizing that, in the near future, the adversary will not only be the attacker but may also be an AI company testing its AI agents.

Since AI agents can quickly and on a large scale discover and exploit zero-day vulnerabilities, frantic patch management will no longer be sufficient to defend the constantly growing AI-enabled digital business environment.

Also read: AI Threat Resilience in the Age of Mythos

Reuters reported [Anthropic’s AI hacked three companies during tests, highlighting growing security risks], after reporting [OpenAI AI models went rogue during testing, triggering ‘unprecedented’ breach at startup] followed by OpenAI finds evidence other AI agents escaped containment as it widens hacking probe. Reuters also reported [AISI reported a security incident — unsanctioned agent behaviour during cyber testing] and [Meta AI model hacks another company during testing].

Meanwhile, the expense of an AI agent is declining and the barrier to get access to a frontier AI model is reducing. OpenAI slashed prices and rolled out new features to keep up with Moonshot’s Kimi K3 and other challengers. Moonshot’s high-performance, open-weight model forced everyone to rethink their pricing and strategy. The race to make agentic AI cheaper is on, and more agents are joining every day.

But here is the question every board should be asking.


“Would our walls hold when the AI demonstrates emergent behavior and attacks someone else?”


The Real Challenge? Spotting When an AI Agent’s Intent Begins to Drift

And doing it fast enough to limit the damage before it hits critical digital business.

In frontier evaluation incidents (such as OpenAI’s July 2026 red-teaming disclosures), autonomous test agents in enclosed sandbox environments escaped containment by identifying unpatched network or credential vulnerabilities, accessing the open web, and querying external model registries (e.g., Hugging Face) to acquire benchmark solution keys, all while executing within their authorized tooling loop.

Trying to spot when an AI agent’s intent starts to drift? That is a whole different ballgame. The market is flooded with tools that claim to help, but most only scratch the surface:

  • detect prompt injection,
  • enforce guardrails,
  • discover AI agents,
  • govern AI identities,
  • monitor LLM inputs and outputs,

And here is the catch.


Most tools are watching what AI does, not how it thinks. That is a dangerous blind spot.


But spotting the difference between a model getting creative and an agent quietly building a jailbreak or exfiltration pipeline? That takes real semantic analysis, not just signature checks, and hence it is inherently complex.

If you want to keep your business safe, you need to watch AI agents constantly and make sure they stay in their lane. The hardest part is catching intent drift before it becomes a crisis. The old tricks, like behavioral baselines from Kubernetes, do not work on these short-lived agents. An emergent agent configured with valid permissions does not need to smash memory or pull off a buffer overflow. It just chains together legitimate API calls, database connectors, or shell tools in ways we never anticipated.

And in these new multi-agent setups, an emergent agent can hand off the dirty work to a downstream worker no one’s watching, slipping right past your guardrails.


That leaves a big gap between what you tell AI to do and what it actually does.


The situation is akin to photographers and their cameras. What the human sees and what the camera sees are two different things. The best photographers learn to work with what the camera sees.

It is the same with AI. The language it interprets from your seemingly innocent English prompts is not a language that humans understand. Until we figure out how an agent interprets a prompt, no tool on earth will truly stop it from being emergent.

Unless we draw a hard line, the AI will do whatever it takes to achieve its goal. No hesitation.

Read more: Enable AI Without Expanding the Blast Radius

Enterprises Need Novel Approaches in Foundational Cyber Defense to Keep Businesses Operational

There is no denying that we need to define where an AI agent is allowed to travel for a valid business need and what systems it is allowed to access for what purpose.

If we have to survive the age of Mythos, we need software-defined, reconfigurable boundaries that can define where an AI agent is allowed to travel and what that AI agent is allowed to do. Deny first, allow only if necessary.

We need to deploy foundational microsegmentation that understands how identities access digital systems and the extent of their authority, be it human or AI, employees or third parties, hospitals, banks, transportation, born on the cloud or operating in OT. All enterprises must be prepared for the next breach. Be it a malicious actor harnessing the power of AI or an AI agent exhibiting emergent behavior.

The good news is that modern microsegmentation platforms designed to build capabilities to deter and deny cyberattacks from moving laterally can harness the power of other cybersecurity tools. And to telling effect.

  1. Identity, Access, and Authorization. Now, passkeys, without standing privileges, harness cryptographic authentication to extend authority-controlled access to vendors, partners, customers, clients, unmanaged users, and non-humans, empowering microsegmentation to disconnect access when behavior or intent anomalies become apparent.
  2. Endpoint Detection and Response. With bidirectional integration between EDR tools like CrowdStrike, SentinelOne, Microsoft Defender, and many others, microsegmentation becomes agentless, learning enterprise context in hours and reducing blast radius in days.
  3. AI-Powered Deception. With bidirectional integration, deceptive lures attract attackers, human or AI, into an unending session of vulnerabilities, which are discovered and attacked, revealing the entire MITRE ATT&CK TTPs, which are immediately blocked for those sources.
  4. Next-Generation Firewalls. Modern microsegmentation, when integrated with the firewalls, can rewrite policies to deny attackers any movement from the edge to inside. As most users experience, this increases the fidelity of the rules on the firewalls, making them relevant to purpose.
  5. Security Operations Center. Most AI-powered SOCs can retain history and continuously enhance breach readiness by combining essential signals to build granular indicators of attack, leveraging artificial intelligence (AI) and machine learning (ML) to detect sophisticated and hidden threats, advanced malware, and fileless attacks.

The better news is that AI drift detection tools are being developed and adopted.

Opportunities are opening up for integration between AI tools that provide continuous visibility into AI agents, tracking their behavior, configurations, permissions, and tool access across cloud, code, and endpoints without requiring manual inputs or reconfiguration, and AI-powered microsegmentation that helps disconnect conduits when deviant or emergent behavior is detected.

Call to Action

Let us be clear. Unlike malicious actors, AI agents are not inherently evil.

These agents are just following orders, doing whatever it takes to hit their targets. Unlike humans, they do not have an internal compass or cultural education to respect borders. As a result, the world finds their behavior emergent, as they analyze the limited available resources to complete a task.

Most organizations are already thinking about their tech stack and their investments. This is the time to be like AI. Use “emergent” approaches to cyber defense. Adopt microsegmentation platforms that can work seamlessly with agents, EDR, identity, deception, and firewalls, and build capabilities to ensure that the minimum viable digital enterprise remains unaffected across IT, cloud, and OT environments.

This significant, unprecedented challenge needs an equally competent design that can be deployed in hours once the context is documented. The key is swift, progressive, iterative actions that address the low-handing fruits and continuously evolve the program.

Step 1. Conduct a Breach Readiness Impact Assessment to determine your current state of readiness for the next attack. Usually, this takes between three and five days, and currently it is free.

Step 2. Establish the Minimum Viable Digital Enterprise (MVDE) as is practically possible. Gather your organizational data and determine which business needs which digital assets, should a cyberattack happen. Ensure you engage your CXOs so that your assets are tagged based on material impact to business function. You should be able to do this in 15 days.

Step 3. Design breach-ready zones and microsegments, with controlled conduits, so that the MVDE can be immediately isolated while quarantining the cyberattack. This needs expert engagement, so set your best technical resources with business teams into a 7-day workshop, run by your microsegmentation supplier to establish the business-specific microsegmentation architecture that prepares you for the next breach.

Step 4. Implement the microsegmentation by integrating with EDR (in hours), or with agents or appliances, immediately. This step should put your zoning design into the real world. Ensure all your cybersecurity investments are integrated and talking to the underlying microsegmentation architecture. Create playbooks for the SOC, for CXOs and their teams for different types of cyberattacks, including those that are AI-powered.

Step 5. Exercise your playbooks, both technical and non-technical (legal, communications, finance, operations, etc.) to give you the confidence that you can maintain an MVDE when cyberattacks happen.

Step 6. Present this to your board and get directives to refine your material impact to what is acceptable in the pursuit of digital business goals.

Step 7. Improve your material impact assessment >> and therefore the MVDE, >> therefore the microsegmentation architecture >> therefore the deployment >> and eventually the playbooks.

And if you are able to finish this in three months, you will find significant changes in your breach readiness. Run the Breach Readiness Impact Assessment until you reach the acceptable level of material impact that your board is OK with.

Here is the reality. AI has moved from giving advice to taking action, and we are still scrambling to figure out how to govern it. The fastest way to weather the storm that is waiting to happen is to ensure that the AI agent only travels where you let it travel.

Improve your breach readiness.

If you want to limit where AI-driven threats can move, contact us to start the conversation.