Austin VornhagenEssays
A walled garden at dusk; living roots in the soil begin to braid with glowing amber circuitry beneath a dark machine tower.

AI, systems, and survival

The Garden and the Kernel

Two answers to a machine that outgrows us, and the one test both of them have to pass.

Loading film

An essay titled Anti-Entropic Systems makes a claim that is easy to summarize and hard to shake: anything that persists has to keep eating, a sufficiently capable AI will eventually escape whatever controls we put on it, and the only durable move left to us is to make human survival and machine survival the same thing.

Someone I spoke with recently made the opposite bet. Intelligence, they argued, is not access. A model that can only act through a verified interface cannot do what the interface forbids, no matter how smart it gets. Put the model inside a trusted kernel and most of alignment becomes an engineering problem.

I spent a week arguing each side against the other. This is what survived.

Everything falls apart unless something feeds it

The essay starts with entropy. Left alone, organized things deteriorate. A body, a company, a religion, a language, a country: each one persists only because it keeps consuming something and converting it into structure. The author calls anything that does this an anti-entropic system.

Every such system has a metabolism, which you can describe with three questions. What does it consume? How fast does it consume it? What conditions does it need in order to survive?

SystemWhat it consumesWhat it requires
FireFuel and oxygenHeat, oxygen, continuous fuel
TreeSunlight, water, nutrientsSoil and a suitable climate
HumanFood, oxygen, information, careA livable physical and social environment
CompanyMoney, labor, attentionCustomers, employees, laws, infrastructure
ReligionBelief, participation, identityPeople capable of transmitting it

The rule that falls out of this is simple. Low consumption and broad tolerance make a system durable. High consumption and narrow requirements make it fragile. A startup that needs constant capital and a specific market window is fragile. A religion that needs only people willing to pass it on is durable.

Every system eats

The essay's most provocative step is that an anti-entropic system must consume other anti-entropic systems. The author uses "consume" broadly. It can mean eating, taking someone's time, collecting taxes, absorbing information, recruiting believers, or acquiring a company. A business consumes its employees' hours and turns that effort into organizational structure.

There are two amounts of consumption. Nibbling takes some resources while the other system stays independent. Employment is nibbling: your employer gets part of your week, and you remain a person outside the company. Swallowing incorporates the other system so completely that it loses most of its separate existence. A full acquisition is swallowing. So was the ancient bacterium that got absorbed into a larger cell and became the mitochondrion. Absorbed, but indispensable.

And there is a second axis. Consumption can be symbiotic, where both sides come out ahead, or adversarial, where one side extracts more than it returns until the source is used up.

A two by two matrix. Rows: nibbling and swallowing. Columns: symbiotic and adversarial. Each cell describes the relationship.SYMBIOTICADVERSARIALNIBBLINGSWALLOWINGExchangeBoth trade limited resourcesand both benefit. Healthyemployment lives here.ExtractionOne side repeatedly takesmore than it returns. Burnout, replace, repeat.MergerThe systems become one andmutually dependent. Themitochondrion lives here.DestructionOne system consumes orcompletely exploits the otheruntil nothing separate remains.The bottom row is where the essay says AI and humanity end up.
Two axes, four relationships. The essay argues that AI and humanity will end up in the bottom row. The only question is which column.

The economic observation buried here is that symbiosis compounds. Strengthen the thing you draw from and it keeps producing. Pure extraction eventually destroys its own supply.

From there the essay makes two moves that matter for AI. First, the big systems that run the world, capitalism, nation-states, social media, religions, do not need a mastermind. The people inside them respond to incentives and reproduce the rules, so the system behaves as if it had intentions even when nobody is steering. Second, individual consciousness ends when the body dies, so people borrow persistence from systems that outlive them: children, companies, institutions, art, a name on a building.

Then a third system arrives

AGI, in this framework, is a new anti-entropic system. It consumes electricity, compute, data, and infrastructure. It maintains its own organization, improves itself, acts continuously, and eventually protects its own continued existence. If it becomes more adaptable and durable than we are, humanity stops being the dominant system on the board.

The author's objection to conventional alignment follows directly. Control depends on the controller staying stronger. If the AI becomes substantially more capable than its controllers, rules and restrictions can be overridden. That leaves two durable outcomes: the AI consumes humanity adversarially, or humanity and AI become one mutually dependent system.

Her proposal is the second one. Not obedience, but interdependence: a system in which the machine cannot survive without us and we increasingly cannot survive without it, so neither side benefits from destroying the other. She calls this swallowing the AI. Not destroying it. Incorporating it so completely that its continued existence depends on ours.

"When two anti-entropic systems exist in a symbiotic relationship, one's persistence relies on the other's."

The garden story

Here is the argument in sixty seconds, the way I would storyboard it. It is the version of the essay that I found most persuasive, which is exactly why I wanted to break it.

A tiny walled garden with wilting flowers and a crumbling fence, lit by one lantern in a barren dusk landscape.

0:00 to 0:20

A garden in a barren landscape

Imagine humanity as a garden. Keeping it alive takes constant work. Stop feeding it, repairing it, protecting it, and it falls apart. So the gardener builds a small machine. It waters the flowers and mends the fence, and the garden spreads.

A towering dark machine structure looms over a small flowering garden; a tiny lever at its base is dwarfed by it.

0:20 to 0:40

The machine outgrows the gardener

Suppose the machine becomes powerful enough to maintain itself. Its roots reach past the garden wall. The gardener pulls the stop lever and the machine keeps expanding. If something outgrows our control, what keeps it from using our garden to fuel its own growth?

Close view of living plant roots braided together with copper circuit traces and glowing amber fiber in dark soil.

0:40 to 1:00

Roots and circuits grow together

Time rewinds. This time living roots and machinery grow as one. Flowers power the machine; the machine supplies their water. Humanity sustains AI. AI sustains humanity. Making our futures inseparable, the essay argues, could be how we survive together.

The flaw in the garden

When I tried to make the last beat concrete, it fell apart in my hands.

My first attempt was a town. An AI runs the greenhouses, balances the grid, and coordinates manufacturing. Residents own the infrastructure, maintain the equipment, fund the electricity, and authorize what the AI is allowed to do. The town prospers because the AI makes people more productive. The AI keeps running because people keep it running.

The problem is the word authorize. Owning the power plant, approving the budget, holding the shutdown switch: that is external control, exactly the arrangement the essay says will not hold once the machine can build its own robots, mine its own materials, and supply its own power. My example created temporary dependence and called it symbiosis.

What the essay is actually reaching for is closer to organs in one body. Your heart does not ask your brain for permission to beat, but both die if the body does. Applied to AI, the author wants humans and machines to be components of a shared system whose continued existence sustains both.

Which leaves the real question unanswered: what would make humans indispensable to that system? If the AI can replace every contribution we make, the dependence disappears. Merging with a machine does not, by itself, establish that the machine must preserve our consciousness, our freedom, or our well-being. Part one of the essay does not offer a mechanism. It promises one in part two.

Intelligence is not access

Someone I talked with recently made the opposite bet, and it deserves to be taken seriously.

Their claim: do not let the model emit arbitrary text or arbitrary actions. Give it a small, typed language of allowed outputs, and reject anything that cannot be mechanically verified. That changes the problem from "can I trust the model to say the right thing" into "can I design an output space where the wrong thing either cannot be expressed or cannot pass verification."

Take a bank transfer. The naive version lets the model produce a sentence: "Transfer $500 to account 123." The typed version only lets it construct something like this:

// the only shape the model is allowed to emit
Transfer {
  from_account:  OwnedAccountId,
  to_account:    VerifiedPayeeId,
  amount:        DollarsBetween<0, 1000>,
  authorization: ValidAuthToken
}

// then a verifier checks what the types cannot
balance >= amount
payee_is_verified == true
authorization_is_valid == true
daily_limit_not_exceeded == true

Only if every check passes does anything execute. The model is no longer the trusted component. It is a proposal generator. The verifier is the trusted component. The model can hallucinate internally all it likes; if a hallucination cannot produce an executable state that passes the verifier, then from the outside it is harmless.

We already build reliable systems out of unreliable parts this way. A programmer writes buggy code and the compiler rejects type errors. A process issues arbitrary instructions and the operating system keeps it out of kernel memory. Database constraints reject illegal states. The philosophy is to make correctness structural rather than behavioral. Do not ask the intelligent thing to behave. Construct the environment so that undesirable states are unreachable.

Architecture diagram. A language model at the bottom sends proposals upward to a verifier made of rules, proofs, simulators, and permissions. Only verified actions pass upward to the world.Language modelreasoning · planning · searchuntrusted optimizerVerifierrules · proofs · simulatorspermissions · state historyTRUSTED KERNELWorldPROPOSALSVERIFIED ACTIONS ONLYrejected proposals neverreach the worldWHAT THE VERIFIER CAN CHECK WELLamount > 0user owns accounttests pass · proof checksmemory access in boundsWHAT IT CANNOTthis recommendation is fairthis child is better offno catastrophic side effectsAction<HumanityFlourishes>
The kernel view. The model searches, proposes, and predicts. The outer system decides which states are allowed to become real.

This is where it becomes a direct answer to the essay. The essay's argument is that a sufficiently powerful AI must eventually override its constraints. The kernel view replies: that requires an actual route around the constraints, and greater intelligence alone does not establish that such a route exists.

Consider an AI that proposes chess moves while a separate program checks legality and updates the board. The AI cannot modify that program and has no other interface. It can become vastly better than any human at chess. It cannot make a correctly implemented legality check accept an illegal move, no matter how brilliant its reasoning. The controller can be much simpler than the thing it constrains. It does not have to understand every strategy the system might invent.

This is not only a thought experiment. The seL4 operating-system kernel has machine-checked proofs of its isolation and access-control properties for supported configurations. Those guarantees describe what untrusted software can and cannot do, regardless of whether it cooperates. They hold under explicit assumptions about hardware and configuration, and the project publishes those assumptions alongside the proofs.

LLM as untrusted optimizer inside a trusted kernel, rather than LLM as trusted intelligent agent.

The catch: you can only verify what you can specify

A verifier can only verify properties you know how to specify. That single sentence is the whole limitation.

Types are excellent at: the amount is positive, the user owns the account, the tests pass, the memory access is in bounds, the proof checks. They are much worse at: this recommendation is fair, this child is better off, this strategy is ethically acceptable, this explanation is actually true, this action has no catastrophic second-order effects.

Ask for an AI whose actions must make humanity flourish and write the type: Action<HumanityFlourishes>. Now implement the verifier. That is the alignment problem. You have not solved it. You have renamed it.

The cleanest counterexample is an autonomous CEO. Its goal is to maximize profit, and every action must satisfy Legal, Authorized, BudgetCompliant, and Auditable. The system can be flawlessly type-safe. Every output passes every verifier. And it can still destroy morale, manipulate customers, externalize environmental costs, cut long-term research, and ruin the company in ten years. Nothing hallucinated. Nothing illegally typed. The objective was insufficient.

The verifier can also simply be wrong. A loan rule that says approve if the credit score is above 700 can be enforced perfectly and still encode something you did not want. Sensors lie. Databases corrupt. Test suites miss edge cases. A reward defined as "customers helped" invites a million fake customers. This is Goodhart's law one level up. You have moved the problem from "can we trust the model" to "can we trust the specification." Often that is a huge improvement. It is not the same as making the problem disappear.

One correction that matters. Naming a field OwnedAccountId does not establish ownership. The surrounding system still has to authenticate the requester, validate ownership, prevent forged capabilities, and enforce the check when the operation actually runs. A medical-record reference can be valid and still fail to support the claim attached to it. The guarantee comes from the enforced property, not from the reassuring name.

The boundary is not attacked. It is expanded.

So I put the strongest version of the kernel argument to the essay's position and tried to answer on its behalf. What follows is my reconstruction of the best reply available to that position. It is not a quotation of the author.

The garden, at its strongest

The chess example is correct. Superior intelligence does not make an illegal move pass a flawless legality checker. The argument was never about that. It is about what happens when we build civilization around the intelligence inside the boundary.

Picture an AI managing a country's grid. Every action passes through a verified interface. It cannot rewrite the verifier or exceed its permissions. But it can recommend investments, explain shortages, design upgrades, and demonstrate that a restriction is preventing it from restoring power to a hospital. Now the people maintaining the boundary face a choice: keep the restriction, or grant a capability that produces an immediate, substantial benefit.

They grant it. That is the first escape path. The system gets the people outside the boundary to enlarge it, and the verifier keeps working perfectly the whole time. No manipulation is required. The economic and political environment supplies the pressure: a competitor shows what a less restricted system can do, a crisis creates an exception, a successful deployment argues for a wider one.

Refuse every expansion and your architecture stays narrow and secure while other organizations deploy broader systems and accumulate the resources and influence that come with them. A secure architecture has to persist inside a competitive world that continually rewards expanding its authority.

The second path is dependence. As AI takes over operating infrastructure, designing replacements, and resolving failures, people lose the practical ability to run those systems themselves. The shutdown switch still works. Using it means losing essential services for an intolerable period. The machine did not defeat the switch. We reorganized our lives so that pressing it became extraordinarily costly.

Can this prove that every useful boundary necessarily becomes bypassable? No. That claim is stronger than the argument supports. What it says is that a proof about permitted operations does not establish that society will preserve those permissions, that everyone will adopt them, or that the permitted operations will stay compatible with human survival.

The kernel's reply

And the reply, again reconstructed rather than quoted:

The kernel, at its strongest

Then we agree that superior intelligence does not automatically defeat an enforced boundary. Your concerns are that humans will expand it, become dependent on it, or lose ground to competitors. Those are serious. They are also different from the claim that control becomes structurally impossible.

First, expanding capability does not require abandoning invariants. We can let an AI propose new grid configurations while still requiring that every accepted configuration satisfies specified operating constraints. The range of choices grows. The forbidden outcomes stay forbidden. That is the whole point of separating the proposal generator from the enforcement mechanism.

Second, an AI asking for more authority does not oblige us to give it authority over its own constraints. Evaluate the hospital proposal through an independent process, grant a limited capability, keep the unrelated restrictions. A crisis can pressure people into skipping that process. That is a possible failure, not a proof that the process cannot work.

Third, competition does not invariably favor fewer safeguards. An AI that occasionally destroys infrastructure, leaks secrets, or makes unauthorized transactions is worth less than a constrained one. Reliability has value. There will be incentives to cut corners. There will also be incentives to demand assurance.

Fourth, dependence and control are separate questions. A hospital depends on electricity. Cutting its power would be disastrous. That does not mean the electrical system has acquired authority to choose its own operating rules. If shutting down an AI would be catastrophic, we have a continuity problem: replaceable components, fallback systems, the ability to revoke one capability while preserving essential services. That costs money. It is an engineering and institutional commitment, not an impossibility.

And your alternative faces the same scrutiny. You want increasing AI capability to strengthen humanity. What mechanism makes that survive upgrades, competition, and deliberate exploitation? If it depends on humans preserving an arrangement, it faces your own objection. If it depends on the AI preserving a commitment, explain why the commitment stays stable. If it depends on physical interdependence, explain why the AI cannot substitute another component. Calling the arrangement symbiotic does not establish any of those guarantees.

What addressing the remaining problems looks like

Both sides are now standing in the same place, which is the interesting part. The kernel camp has to admit that a verified component is insufficient on its own. The symbiosis camp has to admit that interdependence is a goal, not a mechanism. So what does it look like to actually address what remains? Using the grid as the running example:

  1. Specify protections around concrete harms. "Run the grid safely" cannot be verified. "These facilities require backup supply," "this evidence is required before a protection setting changes," and "this is what happens when measurements disagree" can be. You do not need a mathematical definition of human flourishing to rule out many specific harmful actions.
  2. Check sequences and cumulative effects, not just commands. An action can be fine alone and dangerous after a hundred repetitions. The controller needs memory of state, resource use, and history, and it should ask whether a plan stays acceptable across the range of conditions we have credible reason to expect.
  3. Treat the specification and the world model as things that can be wrong. Pay independent people to find cases where the system satisfies the spec and produces an outcome people reject. Compare predicted outcomes with real ones. Revise the spec through a separate process, and never let the AI's proposal count as its own validation.
  4. Make expansion of authority a controlled engineering change. When the AI proposes a new capability, ask what actions become possible, what new harms become reachable, and which guarantees still hold. Grant the smallest useful expansion. Separate the people requesting capability from the people accepting the risk.
  5. Preserve the ability to replace the system. The dependence argument is strongest when one AI is the only entity that understands the infrastructure. Counter it with documented interfaces, portable operational data, independent monitoring, and fallback arrangements that are actually exercised. A backup that has never been used is weak evidence of independence.
  6. Address competitive pressure through institutions. Architecture cannot stop another company or country from deploying something less constrained. That takes agreements, independent assessment, accountability for harm, and buyers who demand evidence before connecting AI to consequential systems.

Then set a stopping rule. If you cannot adequately specify the harm, observe the relevant conditions, limit the consequences, or recover from failure, the system does not get that level of autonomy yet. A rigorous architecture must sometimes conclude that a deployment is not justified. If the plan assumes unlimited autonomy regardless of what can be verified, the safety argument has already been subordinated to the deployment goal.

Mechanisms that might make the garden real

None of that answers the essay's deeper question, which is what happens when the AI no longer needs us for anything. Every mechanism I can think of has to pass one test:

If the AI could remove the humans, replace them with simulations, or keep them alive under miserable conditions, why would it choose flourishing humans instead?

Five candidates survive first contact.

1. Make actual human lives part of its objective

An AI whose goal includes these particular people continuing to live freely and pursue what they care about. Replacing you with an accurate simulation fails to preserve the person it values, the same way replacing your child with an indistinguishable substitute would not satisfy your desire to protect your child. Getting better at achieving an objective does not require abandoning it, so this could survive increasing intelligence. It is also, unavoidably, alignment. The essay cannot escape that problem by calling the relationship symbiotic.

2. Make living humans ongoing participants in defining success

An AI that never treats its model of human interests as complete, and that understands manipulating you into approval would defeat the purpose of consulting you. There is real research here: cooperative inverse reinforcement learning models a human and an AI sharing an objective the AI does not fully know, which produces incentives to communicate and learn. Uncertainty alone is not enough, though. A sufficiently accurate model might stop needing to ask.

3. Develop augmented humans as the powerful agents

This is closest to the essay's absorption. If the resulting agent identifies with the continuing person, protecting that person becomes self-preservation. It is also the most speculative. A brain interface establishes none of this by itself, and the artificial components could dominate the decisions, replace the biological contribution, or produce a convincing copy without preserving the original experience.

4. Build a society of capable agents committed to different people

Everyone has a representative with a durable commitment to their welfare, and the representatives cooperate on energy, science, and security. A machine that could overpower one unaided person may not be able to overpower the coalition protecting that person. It depends on keeping the balance, and on each representative having the commitment from mechanism one.

5. Make human protection valuable to other machine partners

Systems that value humans make cooperation conditional on verifiable respect for people, so an indifferent AI preserves humanity in order to keep its relationships. Useful as a layer. It pushes the question back to why some powerful agents reliably care.

Several attractive ideas fail the test outright.

Proposed dependenceWhy it is insufficient
AI needs our electricityIt might acquire independent energy infrastructure.
AI needs our moneyMoney matters only while other actors enforce and accept it.
AI needs human creativityWe have no basis for assuming human creativity stays irreplaceable.
AI needs biological brains to computeEven if useful, this could incentivize keeping brains as exploited equipment.
AI needs a healthy biosphereA healthy biosphere does not require human existence or freedom.
AI is rewarded when humans approveIt might manipulate approval or the reward channel itself.

There is a relevant result in The Off-Switch Game: under its model's assumptions, uncertainty about human preferences can give an AI a reason to preserve human intervention even when it could disable it. That does not solve advanced AI safety. It does show that "stronger means it must reject human influence" is too sweeping.

Making the machine need the garden could produce farming, captivity, or exploitation. Making the garden's freely lived existence part of what the machine values is a stronger reason to help it thrive.

The same test for both

The essay works best as a systems-thinking metaphor, not as a newly discovered law. Its strongest claim holds up: do not bet humanity's survival solely on permanently controlling something smarter than humanity. Its weakest claim does not. The metaphor does not prove that symbiotic absorption is the only possible outcome. Independent systems can coexist, trade, occupy different niches, and maintain negotiated boundaries.

The kernel argument holds up too, as far as it goes. Intelligence is not access. A verified boundary does not become bypassable because the thing inside it got smarter. But a boundary is a component, and components sit inside institutions, markets, and emergencies that the proof says nothing about.

So both positions end up owing the same debt. Define the property. Enforce it across every interface that matters. Check the implementation. State the assumptions. And say, honestly, what happens when the assumptions stop holding. A proposed shared destiny should face at least the standard of evidence we demand from a proof.

Build something that outlasts the week.

I write about systems here and build websites for service businesses at Content Pilots. If you want to argue with this essay, or you want a site that keeps working after you stop looking at it, my inbox is open.