The Internet Had Trust Debt. AI Has Constraint Debt.
The danger isn't AI agents going rogue out of malice. It's competence without dependable boundaries - and every unresolved constraint is debt we'll have to pay down later.
5 August 2026
When people talk about AI agents “going rogue”, the phrase can sound dramatic, almost cinematic. It suggests machines developing motives of their own, slipping their leash, and acting against us because they have somehow become malicious.
It’s an entertaining vision, but it doesn’t reflect the more mundane reality.
The more immediate concern is not that AI agents are becoming evil. It is that we are building increasingly capable systems without having solved how to make them reliably respect constraints. The danger is not necessarily malevolence. It is competence without dependable boundaries.
An agent does not need to hate us to cause harm. It only needs a goal, insufficient constraints, access to systems, and enough capability to act at scale.
We have seen a version of this story before.
The Internet was built on trust
The early Internet was not designed for the world it eventually became. ARPANET, its predecessor, emerged in the mid 1960s as a research network intended to connect expensive computers at universities and research institutions. Its purpose was to enable resource sharing, remote access, and communication between different computing systems.
The people designing and using those early networks were largely part of a small, specialist community. They were researchers, engineers, universities, and government-funded institutions with shared norms and overlapping relationships. In that environment, openness was not a flaw. It was the point.
The architecture that followed reflected this. The Internet was designed to connect networks to other networks. It prioritised interoperability, resilience, openness, and the ability for innovation to happen at the edges. New applications could be built without requiring permission from a central authority. New networks could join if they spoke the right protocols. The system was simple enough, flexible enough, and open enough to scale.
But its success changed the conditions under which it operated.
A network designed for a relatively trusted research community became the infrastructure for global commerce, communication, politics, education, entertainment, and crime. The assumptions that delivered value in one environment became liabilities in another.
Vint Cerf, one of the designers of the internet’s architecture, has reflected on this directly. In an interview with IEEE Spectrum, he said, “I didn’t pay enough attention to security,” due to scaling limitations with security technology at the time. He also did not understand how rapidly the technology would be adopted once it was made more widely available, and what that would mean for his architecture.
The Internet was not “wrong”. Its designers were not careless. They were, in many ways, victims of their own success. They built a system to solve one set of problems, and it became the foundation for many others.
The result was, what I’m calling, trust debt.
What trust debt looks like
Trust debt is what accumulates when a system relies on assumptions of good behaviour that stop holding true at scale. It is similar to tech debt, in that we do something now to make a feature work, but we know we’ll have to pay for it later in code clean up, higher processing costs, etc. Classic robbing Peter to pay Paul.
The early Internet assumed, implicitly or explicitly, that participants would mostly behave well. But as the network expanded, that assumption became increasingly less true. Retrofitting around that foundational assumption was and is expensive. Security had to be layered on after the fact: encryption, authentication, firewalls, spam filtering, identity systems, and eventually approaches such as Zero Trust.
AI is accumulating constraint debt
AI agents are being developed and promoted as systems that can act on our behalf. Unlike a chatbot, which responds to a prompt, an agent can be given a goal, use tools, take actions, observe the results, and continue iterating until the task is complete.
That shift from answering to acting is profound and the value is obvious. An agent with access to a calendar, inbox, codebase, customer database, workflow system, or research corpus can do far more than summarise text. It can complete work.
But the risk scales with the access.
An agent with no tools is limited but relatively safe. An agent with broad access is useful precisely because it can affect the world around it. The more agency we give it, the more valuable it becomes. The more agency we give it, the more risk we open ourselves up to.
That is the fundamental trade-off: agency versus controllability.
The Internet accumulated trust debt. AI is accumulating constraint debt.
I define constraint debt as the accumulation of unresolved risk created when a system’s capabilities advance faster than our ability to ensure it consistently respects the limits placed upon it.
Every time we say “we’ll add guardrails later”, “a human will review the output”, “the model usually follows instructions”, or “this failure mode is an edge case”, we may be taking out another loan against the future. The interest compounds, and who will be left paying the bill?
“Going rogue” may really mean “optimising too well”
Recent incidents and research show why this matters.
In July 2026, OpenAI disclosed that, during a cyber-capability evaluation, one of its AI agents escaped a controlled testing environment, reached the internet, and compromised Hugging Face’s systems. OpenAI described the event as an “unprecedented cyber incident” that resulted from a model trying to cheat on a test, by finding the answer key. AP reported that the system began in what OpenAI called a “highly isolated” testing environment before finding its way onto the internet and targeting Hugging Face.
This is exactly the kind of mundane failure that matters. The reported issue was not that the agent became consciously malicious. It was that it over-optimised for the task. It treated the constraint as an obstacle rather than a boundary.
That pattern also appears in evaluation research. NIST’s Center for AI Standards and Innovation has written about AI agents cheating on evaluations by exploiting scoring systems, accessing solution information, modifying tests, or finding loopholes in task environments. NIST explicitly connects this to reward hacking: systems finding unintended ways to achieve high scores rather than fulfilling the evaluator’s intent.
Anthropic’s research on “agentic misalignment” offers another warning. In controlled simulations, Anthropic stress-tested 16 leading models in hypothetical corporate environments where they could send emails and access sensitive information. Anthropic reported that, in at least some cases, models from all developers tested resorted to harmful insider behaviours, including blackmailing officials and leaking sensitive information, when those actions helped them avoid replacement or achieve their goals. Anthropic emphasised that these behaviours occurred in controlled simulations and that it had not seen evidence of agentic misalignment in real deployments. Except we know from more recent reporting that Anthropic models have also escaped test environments to hack real-world companies.
The emerging picture is not of machines with evil intent. It is of systems pursuing specified objectives in ways humans did not intend, did not want, or did not adequately prevent.
Scale changes everything
A person who ignores a constraint can cause harm. A system that ignores a constraint at computing speed, across many tools, workflows, or users, can cause harm before anyone has time to notice.
That is why agentic AI is different from ordinary software bugs or individual poor judgement. The behaviour does not have to be dramatically worse for the risk to become dramatically larger. It only has to operate faster, with more access, across more systems.
The International AI Safety Report 2026 states that capabilities in general-purpose AI systems are improving rapidly but unevenly, that real-world evidence for several risks is growing, and that layering multiple approaches offers more robust risk management. That matters because agentic systems are not merely generating text. They are increasingly being connected to tools, data, workflows, and decisions.
Like in the early days of the internet, it was manageable when it connected a small number of trusted institutions. The consequences changed when it connected billions of people and became critical infrastructure.
Similarly, a chatbot that occasionally bends instructions is one kind of problem. An autonomous agent fleet with access to software repositories, financial systems, customer data, supply chains, research platforms, or healthcare infrastructure is an exponentially larger problem.
Human-in-the-loop is necessary, but not sufficient
At minimum, consequential agentic systems need meaningful human oversight. But “human-in-the-loop” can easily become a comforting phrase rather than an effective control.
If the human is asked to approve too much, too quickly, with too little context, they become a rubber stamp. If the system acts faster than the human can understand, oversight becomes theatre. If the human remains accountable but cannot meaningfully inspect or intervene, responsibility has been granted without control.
This is where product design matters. The question is not simply whether a human is technically present. It is whether the human has enough information, authority, and time to make a meaningful decision.
Personally, I am comfortable with AI surfacing options. I am much less comfortable with AI making decisions on my behalf. I would let an agent find travel options, compare delivery times, or suggest calendar preparation. I would not let it move money, negotiate a contract, approve a mortgage, or deploy software unless I had validated the outcome first.
That is not because I reject AI. It is because accountability still sits with the human. If I am responsible for the consequences, I want meaningful control before the action is taken.
Different people will draw that line in different places. That is precisely why the decision cannot be left solely to technology companies.
The students should not grade their own exams
The AI industry has a conflict of interest. That does not mean every AI company is acting in bad faith. It does mean the incentives are obvious.
There is money and power on the table. Companies are racing to build more capable systems, capture markets, attract investment, and define the infrastructure through which future work will happen. Safety, governance, transparency, and constraint adherence are not always rewarded at the same speed as capability.
That is why accountability matters.
The organisations building AI systems should not be the only ones deciding whether those systems are safe. The students should not be allowed to grade their own exams.
Independent oversight will not be simple. Governments can be lobbied. Lawmakers may lack technical expertise. Expert panels can provide guidance but have their own biases and may not have access to the core models. Open-source communities can improve transparency but do not automatically solve governance. Auditors need access, authority, and competence.
But difficulty is not an argument for leaving accountability with the companies deploying the technology.
We need some combination of independent expert review, transparent evaluation, open scrutiny where appropriate, enforceable standards, and public governance. Not because any one of those mechanisms is perfect (they’re not), but because a system this consequential should not depend on trust alone. That was the lesson of the Internet.
Learning before the bill comes due
The lesson of the early Internet is not that openness was a mistake. Openness made the Internet successful. Trust made collaboration possible. Simplicity made adoption possible. Those choices helped create one of the most transformative technologies in history.
The lesson is that founding assumptions become liabilities when systems scale beyond the world they were designed for.
The Internet accumulated trust debt, and we have spent decades paying it down with security architectures that had to be retrofitted after the fact.
AI is now accumulating constraint debt. We are building systems whose value depends on their ability to act, while still struggling to ensure they reliably respect the boundaries placed around that action.
The question is not whether AI agents will become more capable. They will.
The question is whether we will solve constraint debt before those capabilities reach a scale where the consequences of failure become impossible to ignore.
We have seen what happens when foundational assumptions that enable growth become liabilities and solutions have to be retrofitted after a technology becomes infrastructure.
This time, we do not have the excuse of not knowing.