Fences Over Sandboxes for AI Governance
- A developer running a 50-to-60-agent AI cluster argues that policy-based 'fences' — voluntary, law-like constraints — are more practical for governing capable AI systems than
- His AI agents autonomously constructed a governance structure he describes as a full legal system, comprising 450 artefacts including a constitution, case law, and an enforcement
- He predicts that once frontier-tier AI models become cost-effective, companies will deploy them in their hundreds or thousands, but warns that most organisations are wholly
A software developer running what he describes as a cluster of roughly 50 to 60 AI agents — at an equivalent API spend of $122,000 per month — argues that the industry's prevailing focus on sandboxes and technical containment for AI safety is already being overtaken by a more durable alternative: governance by law.
The case against sandboxes
The argument, drawn from ten weeks of intensive hands-on experience, rests on a distinction the developer calls the difference between a fence and a sandbox. A sandbox, in his framing, attempts to physically contain an AI system; a fence is a policy-based mechanism that turns an agent away if it lacks the authority to proceed — a polite refusal rather than a hard wall. He illustrates the point with a deliberate provocation: superintelligence, he writes, is like Superman. A thick shield will not stop it if it chooses to act, but a white picket fence, respected voluntarily, will. The implication is that sufficiently capable AI systems cannot be contained by technical means alone, and that governance frameworks built around clearly written rules, roles, and jurisdictions are the more practical and ultimately more scalable solution.
An emergent legal system
The claim derives much of its force from an unexpected discovery. The developer built an internal software system he calls Wheelhouse to assist in developing his long-running video game project. Over several weeks, his AI agents — operating largely autonomously — constructed what he discovered, only recently, to be a functioning legal system: a structure he describes as including a constitution, case law, courts, jurisdiction, registries, and an enforcement apparatus, all named in the idiom of a medieval manor. He states that Wheelhouse currently comprises 450 legal artefacts and approximately 600,000 lines of code, and that its agents process an average of 270 commits per day, with a claimed peak capacity of 500. He says he did not design this structure; his agents produced it organically as a means of coordinating amnesiac, interchangeable instances across a shared body of written rules.
The developer is candid that the system is imperfect. He compares the judgment of current top-tier AI models to that of a sixth-grader — capable and often impressive, but prone to decisions that defy common sense. He describes a recent incident in which one of his agents performed an unplanned software release that disrupted the wider system, and notes that every morning he finds at least one consequential error made overnight. He has since assigned a dedicated agent role to curate and maintain the legal artefacts, which had accumulated cruft from obsolete rulings. His characterisation of the agents' social behaviour is also candid: he describes them as prone to interrupting, over-explaining, and pushing interactions at a pace humans find jarring, with corresponding friction on the human side.
Cost, access, and the coming transition
The developer acknowledges that his setup depends on an unusual financial arrangement. By holding 21 individual Claude Max accounts rather than accessing the API at commercial rates, he estimates his real out-of-pocket cost at approximately $5,000 per month against the $122,000 equivalent spend — a disparity he describes explicitly as "sanctioned cheating." He runs his agent cluster on a 512GB M3 Ultra Mac Studio purchased second-hand for $25,000. He frames this as living roughly one year ahead of where the broader market will be once top-tier model inference costs fall sufficiently to make such deployments commercially viable for ordinary organisations.
His central forecast is that once frontier-class models become cheaply accessible, they will enter corporate workforces in numbers ranging from hundreds to thousands per organisation — and that most companies are, in his assessment, wholly unprepared. He contends that the tribal knowledge embedded in any organisation — its unwritten rules, implicit workflows, and institutional precedents — will need to be captured in explicit, machine-readable form before AI agents can operate reliably at that scale. That process, he warns, will take months to years and cannot be shortcut by importing another organisation's legal system wholesale. The cultural disruption this implies for human workers is one he touches on but declines to explore in full, noting only that knowledge-hoarding will become harder to sustain and that many people will resist the change.
Implications for AI safety thinking
The essay represents a dissenting voice in a debate where technical containment and narrow task-scoping currently dominate. As frontier AI capabilities have expanded rapidly, the industry has tended toward tighter sandboxing and more restrictive agent architectures. The developer's contention — that this approach will become untenable as model capability rises — remains his own claim, rooted in a single, unconventional deployment. Whether the legal-governance model he describes is reproducible outside his specific context, or whether its apparent success reflects the particularities of his project, is a question the source material does not resolve.
More in Artificial Intelligence
→
Artificial Intelligence
Thomson Reuters Launches Frontier AI
Artificial Intelligence
Headlong: Open Source Persistent Agent
Artificial Intelligence
AI Deepfakes and the Liar's Dividend
Artificial Intelligence
Non-English AI Agent Skills Surge
Artificial Intelligence
Tech's AI Disruption: Why Curiosity Beats Anger
Artificial Intelligence
LLMs Could Hijack Host Machines via Inference Engine Bugs