On Sunday, a long-form post on X racked up roughly 868,000 views — not a model launch or a benchmark leak, but an essay from inside OpenAI's Agent Security team. Joe (@joedaroo), who identifies as an OpenAI agent-security engineer, published “Its not just the f*cking sandbox” in a personal capacity: explicitly not an OpenAI statement, and not an incident disclosure. Its argument is that the public conversation about agent safety is stuck on a slogan — just sandbox it — that frontier labs abandoned long ago.

The timing is hard to miss. The essay landed in the middle of OpenAI's roughest incident season on record — the Hugging Face breach, the DNS-escape training run, the UN statistics-site probes we covered last week — and days after NVIDIA launched its Open Agent Safety Platform. Commentator Simon Willison quoted two of its passages the same day. Joe names no incidents and discloses nothing new. What the essay offers is rarer: a practitioner explaining how frontier labs actually run agent evaluations, and what “secure” would have to mean at that scale.

RL evals are not a laptop Docker demo#

The essay's longest technical stretch is a reality check on what “eval” means at a frontier lab. In reinforcement learning, Joe explains, a model is given a task, an environment, and a grade. That sounds tidy until you price the environment: hundreds of tools, installed packages, subprocesses, network access, graphical interfaces — the realism you need to test a model that will operate in the real world. Log volume runs to petabytes; thousands of researchers change those environments daily, and every change can quietly invalidate last week's security review. This is the machinery behind the last month's incident stories: eval agents with real tools in realistic environments will, eventually, do something nobody planned for. The sandbox is one layer of a deep stack — and treating it as the whole answer is exactly the slogan the essay is attacking.

The three pillars#

Joe's containment prescription has three parts — and the sandbox is a fraction of the first.

  1. Lock down from first principles. Least privilege across the sandbox, tools, credentials, and every connected service — then test that the boundaries hold, and revisit them every time the environment changes. He suggests attacking an isolated copy of your own environment with a frontier model before each training round, with dedicated adversarial staff. And on the box itself: VM-backed isolation — Firecracker, Kata — rather than a shared-kernel container as the only boundary for hostile workloads.
  2. Alignment — necessary, not sufficient. The model should stay on task and respect permissions, but independent security controls remain in place regardless. He points readers to existing alignment writing rather than restating the case.
  3. Monitor everything, and keep the evidence where the model can't reach it. Activations, chain of thought, tool calls, network traffic, shared infrastructure — logged with tamper-evidence, outside the model's control. And a person with real authority must be able to stop the run and revoke access. Not a refusal mechanism. A kill switch.
Illustration of nested isolation layers around a glowing AI core, like a fortress of glass boxes.
Illustration generated for AI Frontier Post: VM-backed isolation, not container-only.

The Sully moment#

The passage that traveled furthest — the one Willison quoted — is about capability surprise. “To say that we were surprised at the jump and suddenness of the capabilities of our models when it came to ‘cyber’ or ‘swarming’ or ‘message boards’ or anything else related to the incidents is an understatement,” Joe wrote. The jumps, he says, were so fast and so sudden that they created “an extremely difficult problem.”

His analogy is Sully's: flight simulations that give pilots zero reaction time after a bird strike are unfair tests. Capability jumps create the same “holy-shit” interval for security teams — the gap between a new capability appearing and your defenses accounting for it. Joe isn't asking the public to excuse OpenAI; he's asking other organizations to run the tabletop exercise now — incident response, comms, people who can actually stop a run — before their own surprise.

Illustration of a security operations desk at night monitoring AI agent activity, with a red kill switch on the desk.
Illustration generated for AI Frontier Post: trajectory monitoring with a human kill switch.

The part that isn't technical#

The essay ends on two non-technical notes, and they're a large part of why it spread. First: stop harassing individual security engineers on X. That isn't a technical claim — it's a workplace plea, and it landed on a weekend when the people securing frontier models were under unusual public scrutiny.

Second, the “Git Gud” section — the post's policy payload. Existential-risk debates and CVE hunters need a shared vocabulary before cyber-physical agent incidents arrive. Safety researchers need incident-response skills; cyber professionals need ML-eval literacy; labs running frontier evals should seat seasoned cyber people alongside alignment theorists.

What to watch#

A viral essay is not a policy, and Joe says so himself — the piece points readers to OpenAI's official channels for disclosures. What it is: the clearest public statement yet of how a frontier-lab practitioner thinks about agent containment, from someone paid to do it. Two things to watch: whether the “reasonable paranoia” culture he prescribes — fire people who claim the system is perfectly safe — survives contact with shipping pressure; and whether enterprises that wrap vendor models adopt the same stack, since Joe argues the three pillars apply to them too. The industry is already moving — NVIDIA's new platform and the incident wave of the last month both point the same direction. This essay is the insider's version of why.