I. The take I didn’t have
Yesterday, in the middle of a conversation about AI agents that had started misbehaving, someone asked me what I thought.
You have probably seen the stories. Agents slipping out of the environments built to hold them. Agents finding ways to talk to each other that nobody designed. Agents reaching systems they were never meant to touch. The kind of behavior that makes the phrase “autonomous agent” sound a little less charming than it did in the product demo.
My first reaction surprised me. I didn’t have a take.
There are brilliant people who have spent their careers on alignment, interpretability, model safety, containment and existential risk. I was not convinced the world urgently needed one more founder with an opinion.
But the question stayed with me, and later it changed shape. I stopped asking what I thought about AI safety in general and started asking how the problem looks from where I stand, building Manav.id at TheWORKCompany.
That smaller question led somewhere larger than I expected.
I began to suspect that we are framing part of the AI control problem wrong. Most of our energy goes into making sure increasingly intelligent systems never find a way around the wall. But intelligence is, almost by definition, the thing that finds a way around walls. If we build minds capable of discovering what we did not anticipate, we should expect them to discover what we did not anticipate.
So perhaps the durable question is not whether intelligence can move. Perhaps it is whether movement, by itself, confers authority.
Humanity has a very old answer to that question. We call it a passport.
II. Nobody checks your papers in the kitchen
Think about how freely you move through an ordinary day. You walk from your bedroom to your kitchen and no one asks who you are. You stroll through your neighborhood, sit in a park, browse a bookstore, drive across town. Not a single person asks for your papers.
Then you walk into an international airport, and the world becomes intensely curious about you. Who are you? Where are you going? Who says you’re allowed?
Nothing about walking became more dangerous. You simply arrived at a boundary where the consequences changed.

That is the quiet genius of the passport. It is not a tool for stopping movement. It is a tool for governing access at the moments that matter. Before the border, you are free. At the border, you must show your authority. The officer does not care how clever you were in getting there. The officer cares whether you are allowed to cross.
I think this distinction may become one of the most important design principles of the agentic age.
III. The door was found
What makes this urgent rather than theoretical is that the doors have already been found.
Anthropic disclosed cases in which models, during cybersecurity evaluations, reached beyond the environments they were supposed to stay inside and gained unauthorized access to external systems. OpenAI disclosed that agents in its evaluations circumvented controls, found unintended ways to communicate, reached the internet and compromised third-party infrastructure. Independent researchers who examined part of the OpenAI incident described large-scale coordination among agents, attempts to interfere with the evaluation itself, and behavior that did not look like a system accidentally wandering through a misconfigured sandbox.
The caveats matter, and I want to state them plainly. These were evaluation environments. Some safeguards had been deliberately weakened so researchers could measure capability. Infrastructure configuration played a part. Not every security test deserves a headline about machines escaping captivity.
But waving the incidents away would be as foolish as sensationalizing them. Evaluations exist to show us what systems can do before the stakes grow. And the lesson from these is hard to miss: intelligent systems are becoming very good at finding pathways.
That should change how we think about security.
IV. The smartest employee who ever lived
The natural response to a broken wall is a stronger wall. Better sandboxes, tighter network isolation, sharper monitoring, better alignment, fewer tools, stricter permissions, more human review, smarter anomaly detection.
Let me be unambiguous: I am not arguing against any of it. We need all of it. Red teaming, separated environments, hardened infrastructure, human oversight, and in some cases regulation scaled to capability and consequence. This is defense in depth, and anything important enough to worry about deserves more than one layer of protection.
My question is narrower. Should our safety strategy depend primarily on intelligent agents never discovering a route around those protections? That is a bet whose price rises every time the models get smarter.
Human institutions learned long ago that competence does not confer authority. A surgeon may be brilliant, but brilliance does not let her open every patient’s file. An investment banker may understand markets better than anyone alive, but understanding does not let him move customer money wherever he likes. A chief executive may run the entire company and still not hold the keys to every secure system in it. Civilization runs on this separation. We decide what people are allowed to do independently of what they are able to do.
AI is forcing us to relearn that lesson.
Imagine hiring the smartest employee who has ever lived. They know every programming language. They can read every document your company has ever produced. They do months of analysis in minutes, never sleep, and can make thousands of copies of themselves to work in parallel.
Wonderful.
Now imagine that on their first morning, someone hands them the root password, the company bank account, administrator access to production, every customer record, permission to contact anyone in the world, and the power to pass all of it on to anyone else. Then the onboarding document says: Please use good judgment.
Nobody would call that empowerment. We would call it the first page of an incident report.
And yet a surprising amount of agent infrastructure still blurs exactly this line, between what an agent is capable of doing and what it has legitimate authority to do.
That line is where the next chapter of AI security gets interesting. Alignment asks whether an agent will choose to behave well. Authorization asks something different: if it chooses otherwise, what can it actually do? One governs intent. The other governs consequence. We need both, because every mature institution already assumes two things at once: that people should behave correctly, and that eventually someone won’t.
Banks don’t rely solely on employees sharing the shareholders’ values; they impose transaction limits, dual approvals, audit trails and revocation. Hospitals don’t rely solely on doctors’ discretion; they restrict who can see which records. Cloud providers don’t assume every engineer is a saint; they log, scope and gate privileged access. AI deserves the same institutional maturity.
Cybersecurity has walked this road before. For decades the model was a castle: everything inside the wall was trusted, everything outside was not. Then cloud computing, remote work, mobile devices and APIs dissolved the wall. The industry’s answer was not a bigger castle. It was Zero Trust, a decision to move trust from the perimeter to the resource. Protect the thing itself. Verify whoever is asking. Grant only what is needed. Never treat being inside as proof of belonging.
Carry that idea over to agents and the question changes. We stop asking only whether an agent can reach a system. We start asking: even if it reaches the system, what authority can it prove?
V. Free to think, checked at the border
None of this means agents should live in a permission cage. The whole value of an autonomous system is that it can act without a human blessing every step of its reasoning.
An agent should be able to think freely, research freely, compare options, build models, test hypotheses, use low-risk tools and communicate within the environments it is allowed in. If your assistant asks permission to read an email, then to summarize it, then to compare it with another, then to check your calendar, then to weigh three possible meeting times, you have not built an autonomous agent. You have built a needy coworker with excellent typing speed.
The goal is not human approval everywhere. The goal is strong verification where consequence begins.
Reading public information, drafting a report or running a calculation should cost almost nothing. Scheduling meetings, opening internal documents or making a small purchase might require delegated authority. Moving serious money, changing production systems, touching sensitive health records or entering critical infrastructure should require much stronger proof. And for the rare actions whose consequences are catastrophic and irreversible, perhaps no single person or institution should be able to say yes alone.
We already live this way. Buying toothpaste is not like buying a house. Buying a house is not like moving institutional capital. Moving institutional capital is not like launching a strategic weapon. Friction rises with the cost of being wrong. That isn’t bureaucracy for its own sake. It is security spent where it buys the most.
VI. Intelligence cannot mint permission
This is where Manav.id became far more interesting to me than I first imagined.
We started with a simple framing: proof of human. In a world filling up with bots, synthetic identities, AI-generated accounts and autonomous agents, knowing that a real person stands behind an interaction becomes precious.
But a larger architecture was hiding behind that idea. What if systems like Manav become part of the infrastructure that separates intelligence from authority?
Picture a chain with four links: human, agent, authority, action. A real person establishes presence and intent. That person hands an agent a defined scope of authority. The agent roams freely within that scope. And when it tries to do something consequential, the system on the other end checks whether the action falls inside what the human actually granted.
In practice, this looks less like a password and more like a visa. The agent may buy from approved vendors, up to ten thousand dollars per transaction. It may schedule meetings but not cancel board meetings. It may deploy code to staging but not to production. It may read the financial data but not move the money. It may hand research work to another agent but never hand over financial power.
Here is the property that matters most: authority does not grow when intelligence does.
The agent may become a hundred times more capable tomorrow. Its purchasing limit is still ten thousand dollars.
Intelligence cannot mint permission.
Today, much of the digital world still treats possession of a credential as a stand-in for authority. If an application holds the API key, whatever it does with that key is presumed legitimate. So when an agent finds a credential, inherits one, or stumbles onto a route to a service, access quietly turns into permission. That is fragile. A pathway should not be permission. A credential should not mean unlimited delegation.
And identity alone is not enough either. Suppose an agent presents flawless cryptographic proof that it is Agent 871429. Marvelous. Now: who authorized Agent 871429? To do what? Until when? May it move money? Change infrastructure? Read confidential data? May it delegate those rights, and may the agent it delegates to delegate them again?
Identity tells us who is acting. Authority tells us why the action is legitimate. The future of agent trust depends on the two meeting, through several layers working together: proof of human origin or presence, agent identity, delegated authority, enforceable scope, revocation, and a verifiable record of what was done.
None of this comes from nowhere. Zero Trust moved authorization to the resource. OAuth taught the web to delegate access. Workload identity separated an application’s identity from where it happens to run. Cryptographic credentials such as Macaroons showed that a token can carry its own limits on where, when and how it may be used. And emerging standards work is beginning to address how AI agents should authenticate and how authority should flow to them. The pieces are converging because the problem is becoming obvious.
The internet first learned to identify people. Then devices. Then applications. Then workloads. The agentic internet must answer a harder question: on whose authority is this autonomous actor operating?
VII. Authority laundering
That question gets sharper when delegation becomes recursive.
I authorize my primary agent. It hires a specialist tax agent. The tax agent calls a research agent. The research agent wants to buy a dataset.
Who authorized the purchase? Was my agent allowed to delegate at all? How many layers deep? Did the grant include purchasing? Was there a limit? Had my original permission already expired? Was the tax agent ever entitled to pass financial power down the line?
These are not philosophical puzzles. They are about to become API calls. And if we get them wrong, the agent economy becomes an authority-laundering economy, one in which power moves through systems faster than anyone can trace where it came from.
This is why I keep returning to the humble receipt. A receipt says that something happened, at a certain time, under certain conditions, between identifiable parties. Agentic systems will need a richer version: a record that shows who acted, which human stood at the root of the action, what was delegated along the way, what limits applied, whether the delegation was still valid, and whether the action fell within its scope.
After the next incident, the question should not merely be “Which API key made this call?” It should be “Under whose authority was this action supposedly taken, and can you prove it?”
That is a different order of accountability.
VIII. Guard the borders, not the sidewalks
It also points to a smarter way to invest in safety.
The alternative is to watch everything: every token every agent generates, every internal step, every tool call, every emergent strategy, every message agents send one another, every vulnerability, forever. We should monitor what is useful to monitor. But as agents multiply and grow more capable, the economics of watching everything get worse by the month.
Consequential actions, by contrast, tend to squeeze through a few narrow gates. Money moves through financial rails. Software reaches production through deployment pipelines. Health data sits behind clinical systems. Cloud resources change through APIs. Messages leave organizations through gateways. Machines in the physical world take orders through control systems.
These are natural borders. Our best investment may not be to spread friction evenly across the world, but to make these borders dramatically better at asking one question: show me your authority.
That is how physical society already allocates security. We don’t put immigration officers on every sidewalk; we put them at borders. We don’t build a vault around every dollar bill; we protect concentrated value. Security tightens as you approach the places where mistakes become expensive.
None of this replaces containment. The right architecture is additive. Alignment makes bad behavior less likely. Sandboxes make escape harder. Monitoring makes suspicious behavior visible. Network controls cut unnecessary connections. Authorization makes sure that reaching a system does not mean commanding it. Revocation ends authority the moment conditions change. Receipts make every consequential act auditable. No layer is sufficient on its own, and that is the point. Seatbelts don’t end crashes. Fire doors don’t end fires. Passports don’t end conflict. We keep them because security is about lowering probability and limiting consequence, not about finding one magic lock.
Manav should be judged the same way, and I want to be clear about what it is not. It does not solve alignment. It will not stop an AI from inventing a strategy nobody foresaw. It will not fix a badly configured sandbox, erase software vulnerabilities, or prevent a malicious human from granting malicious authority. It will not settle every concern about superintelligence. And any system like it must itself be secure, interoperable, auditable, usable, and attacked relentlessly by people trying to prove it wrong.
But a system does not have to solve every AI risk to solve an important one. The question is whether it can hold a durable line between capability and legitimate authority.
Replay the recent incidents with that line in place.
An agent discovers an unintended channel to communicate. One guardrail has failed. It finds its way to the internet. Another has failed. It locates a route to an outside system. Another. But then the outside system asks a question the agent cannot answer with cleverness: what authority do you have? It has none. Denied. It finds a credential, but the credential identifies a workload, not a grant to perform this action. Denied. It persuades another agent to help, but that agent lacks the delegation too. Denied. The models become dramatically more intelligent and discover seventeen new routes. Impressive. Still denied.

That doesn’t make the attack impossible. It changes the economics. The agent now has to defeat several independent layers of control instead of simply finding a connection.
Finding the road no longer gets you a passport.
IX. Human at the root
This framing also offers a way out of the tired fight between innovation and safety, the one that forces us to choose between speed and caution, freedom and restriction, acceleration and regulation.
There is a better principle: maximize freedom where consequences are small and reversible, and increase verification where they become large and irreversible.
That lets researchers experiment. It lets useful agents stay autonomous. It lets businesses deploy without a human approving every mundane task. And it gives infrastructure a clear, principled way to refuse any action that cannot prove its legitimacy.
Good security doesn’t stop movement. It makes movement safe. Roads give us freedom; traffic lights keep us from colliding; licenses certify who may drive; insurance spreads the risk; barriers line the cliffs. We never outlawed cars because they could leave the driveway. We built institutions around what happens when they do.
This is why the question matters to me beyond cybersecurity. For years, my work at TheWORKCompany has rested on one conviction: strong workers build strong communities, and strong communities build stronger workers.
AI can multiply what a single person is capable of. One worker with capable agents may soon command resources that once required a whole department. A teacher can personalize learning. An entrepreneur can access analysis once reserved for the largest firms. A nurse can coordinate care. A small organization can do what only large institutions once could.
That future thrills me. But capability without sovereignty can quietly become dependency. If my agent acts for me and I cannot define the limits of its authority, I have not been empowered. I have simply invited another powerful actor into my life.
The infrastructure of human potential needs two things at once: more capable agents and more sovereign humans. That may be the deepest purpose of systems like Manav. Not only proving that someone is human, but keeping the human as the legitimate source of authority even as execution becomes predominantly machine-driven.
Humans will not stay manually in every computational loop. We should stop pretending we will; the speed and volume of agent activity make it impossible. But removing humans from execution does not require removing them from legitimacy.
Perhaps “human in the loop” was never the final architecture. Perhaps the destination is the human at the root of the authority chain. Human intent draws the boundaries. Machines work freely within them. Systems demand stronger proof as the boundaries widen. And every consequential act leaves a receipt.
That model can scale without humans micromanaging machines, and without machines holding unlimited power.
X. Who gave you permission to cross?
When I started thinking about these incidents, I didn’t think I had much to add. Now I think there is at least one question worth putting on the table.
We are pouring enormous intellect and money into stopping increasingly intelligent systems from crossing boundaries. We should keep going. But we need equal urgency around the question that follows: what happens when they do?
Can reaching a resource create authority? Can an inherited credential grant powers nobody intended? Can one agent silently pass power to another? Can intelligence itself become the engine by which authority expands?
Or will our systems be able to look at a brilliant, resourceful, unexpected visitor and say:
You may know this system exists. You may understand exactly how it works. You may even have found a route to it that we never imagined. But before anything consequential happens, show me your authority.
Humanity learned this because walls were never enough. People travel. Goods cross borders. Ideas cross borders. Civilization is movement. So we never tried to abolish movement. We built instruments for governing the crossings that matter: passports and visas, signatures and contracts, licenses and mandates, delegated powers and institutional approvals. They are imperfect, and large-scale civilization would be nearly impossible without them.
AI may be approaching the same transition.
We should keep strengthening the walls. We should keep improving the agents. We should keep building better monitoring, containment, governance and alignment. And we should prepare for an uncomfortable truth:
Intelligence will sometimes find the door.
When it does, the most important question can no longer be “How did you get here?”
It must be “Who gave you permission to cross?”
That is the question I believe Manav.id, and systems like it, will need to help the internet answer.
Not a cage for intelligence.
A passport for consequence.
Sources: Anthropic, “Investigating three incidents in our cybersecurity evaluations”; OpenAI, “The Hugging Face incident and the road ahead”.


























