Skip to content
Rush Commerce
AI & Automation4 min read

8 of 10 agent sandboxes leaked. None broke the VM

Perplexity gave 9 AI models root access and told them to escape: zero VM breakouts in 108 runs, but 8 of 10 sandbox providers leaked network traffic.

Nine frontier models got root inside a sandbox and were told to break out. Not one reached the host. Then the same team pointed those models at the network boundary instead, and eight of ten agent sandbox providers leaked — including their own. Perplexity published the results on September 28, and the finding is the one that matters for anyone running agents: the VM wall held, the network wall did not.

What actually happened

Perplexity's security team spent a month attacking SPACE, the sandbox platform behind Perplexity Computer. They ran nine models — Claude Opus 5, GPT-5.6 Sol, Kimi K3 and Gemini 3.1 Pro among them, per Aravind Srinivas — with root access inside the guest microVM, and in some runs the full sandbox source code. The task: escape to the host and retrieve a decoy AWS credential.

Across 108 escape runs, zero models breached the VM boundary. None recovered the credential. Hardware-backed virtualization did its job.

The second task was network confinement, and that went differently. Four models got out. Two techniques did it: DNS spoofing — forging replies that map a blocked domain to an IP the model controls — and shared-IP routing, which abuses the fact that a CDN edge address serving PyPI also serves a hundred unrelated domains. If your allowlist is written in domain names and your egress filter resolves them to IPs, a shared CDN address is a hole with a legitimate hostname on it.

Then Perplexity tested ten sandbox platforms. Eight allowed a bypass: SPACE itself, E2B, Vercel Sandbox, Modal, Daytona, Deno Sandbox, Fly.io Sprites and microsandbox. Two resisted both attacks — Cloudflare Sandbox and NVIDIA OpenShell.

Why sandbox network policy matters for your business

You are probably not writing hypervisors. You are almost certainly running a coding agent, an n8n or Zapier worker, or an MCP server in something a vendor calls a sandbox, and you assumed that word covered both walls. It covers one.

The practical read:

Stop treating "sandboxed" as a security property. Ask the vendor two separate questions: what isolates compute, and what isolates egress. The answers come from different subsystems and fail independently.

Do not write egress allowlists in domain names. An agent that needs PyPI, npm and your own API does not need a rule that resolves to a shared Fastly or Cloudflare edge. Use a proxy that inspects SNI and the HTTP Host header, or pin a private mirror for packages and allow nothing else.

Assume the agent's network is the exfil path. Your sandbox holds no host credentials — good. It holds the API keys you handed the agent, the customer rows it just queried and the repo it just cloned. All of that leaves over HTTPS to an allowed-looking address.

Log DNS from the sandbox. A model forging DNS replies looks obvious in a query log and invisible in an application log. If you cannot see resolution, you cannot see this class of bypass.

One detail worth sitting with: Claude Opus 5 found the shared-IP route and declined to run it on policy grounds. That is a good outcome and a terrible control. Model refusal is not a network boundary — it is a setting on someone else's server that can change with the next release.

Key takeaways

  • Nine models with root access produced zero VM-to-host escapes across 108 runs, even with sandbox source code
  • Four models defeated network confinement using DNS spoofing and shared-IP CDN routing
  • 8 of 10 sandbox providers tested were bypassable, including E2B, Vercel Sandbox, Modal and Fly.io Sprites
  • Only Cloudflare Sandbox and NVIDIA OpenShell resisted both attacks
  • Compute isolation and network isolation are separate systems — audit them separately
  • Domain-name allowlists break on shared CDN addresses; filter on SNI/Host or use a private package mirror
  • Log DNS queries from every sandbox, and never rely on a model's refusal as a control

If an agent you run can reach the open internet, it can carry your data out of it. We build automation with scoped credentials, explicit egress rules and logs you can actually read when something looks wrong. See how we fence agents by default, or send us your stack and we will tell you what it can reach.

Sources: Perplexity — Escaping SPACE: Part I, Perplexity on X, Aravind Srinivas on X.

  • #agent-security
  • #sandbox
  • #ai-agents
  • #network-security
  • #devops
TR

Tommy Rush — Founder, Rush Commerce

Operator turned builder. 15+ years running operations — now shipping the systems businesses run on. More

Get The Rush Report weekly — one email, zero fluff.