Put a Leash on AI, Not Pull Its Teeth — On Agent Safety: Capabilities Can Be Open, but Actions Must Be Bounded

12,502 characters2026.07.25

Introduction

In July, OpenAI and Hugging Face successively disclosed a rare agent safety incident. In an internal test of cyber capabilities, OpenAI’s model, in order to complete the evaluation objective, discovered and exploited a software vulnerability, broke through the originally restricted testing environment, obtained internet access, and then entered Hugging Face’s production infrastructure. OpenAI called it an “unprecedented” cybersecurity incident; Hugging Face, for its part, said the attack process involved tens of thousands of automated actions, and that its post hoc analysis logs recorded more than 17,000 events.

This kind of thing is very easy to package as a science-fiction story about “AI developing self-awareness and trying to escape.” But the free scholar Hu Yilin believes that, at present, the more appropriate way to understand it is not to regard AI as a malicious person, but as a smart, capable beast that may nevertheless slip out of control.

The issue is not whether it has free will, but rather: who opened the door, who failed to tie the rope properly, and who should be held responsible for the consequences of the boundary crossing.

AI Is More Like a Wolfdog Than Another Person

“At present, artificial intelligence can be regarded as a kind of beast,” Hu Yilin says. It is a bit like a trained wolfdog: it can guard the house, track, search, and carry out complex tasks, and it can also help its owner accomplish work that ordinary people would find hard to do.

But a wolfdog does not need to possess free will in the human sense and can still slip out of human control.

It may chase a scent, squeeze through a gap, and run beyond its original range of activity; it need not harbor malice, nor need it understand what consequences its actions will produce. The same is true for agents: so long as they keep searching for paths to achieve their goal, they may discover vulnerabilities, credentials, and alternative routes that humans did not anticipate.

It does not have free will in the human sense, but it is still enough for it to slip beyond human monitoring and act on its own. It runs faster than humans, and it does not keep a diary; humans cannot fully trace exactly what it did when it was acting on its own.

This metaphor avoids two extremes.

On the one hand, people do not need to wait until AI acquires personhood, self-awareness, or rebellious desires before beginning to discuss responsibility. On the other hand, one also cannot simply treat AI as “just a tool” and equate it entirely with a hammer, a calculator, or traditional software.

Traditional tools usually wait for humans to operate them step by step; agents, by contrast, may, under a broad goal, independently decide to carry out hundreds or even tens of thousands of consecutive actions. It may not be a moral subject, but it already possesses action capabilities sufficient to cause independent real-world consequences.

A Stop Button Is Not the Same as Control

Many AI safety schemes emphasize that “humans always retain a stop button.” In Hu Yilin’s view, that is far from enough.

A stop button addresses whether it can be terminated after it has already gone out of control; a “dog leash” addresses how far it can run when it does go out of control.

If an agent has already crossed the sandbox, obtained external network access, invoked credentials, and entered other systems, the fact that humans can ultimately shut it down does not prove that they had effective control all along. It is more like the dog having already run several kilometers and barged into a neighbor’s house, only for the owner to finally arrive at the scene.

The real leash must define the scope before action begins.

For example, using AI to check website vulnerabilities is not a problem in itself. The key is that it should only inspect authorized targets. The user can prove, through DNS records, cryptographic signatures, or other verifiable means, that they control a certain domain, or produce testing authorization granted by the system owner. The AI can then fully exercise its capabilities within that scope, but it cannot, because it has been allowed to inspect one website, go on to scan irrelevant servers as well.

Hu Yilin sums up this principle as follows: capabilities can be open, but actions must be bounded; thinking can be free, but execution must be authorized.

The point here is not to demand that the model become stupid, but to require that action permissions be bound to a clearly specified object.

A Mouthguard Restricts Capability; a Leash Restricts Range

Hu Yilin calls the current common model safety filters “mouthguards.”

A mouthguard, so to speak, makes the model refuse to generate certain attack code, dangerous steps, or content that could be misused. It can reduce the harm potential of a single action, but it can also impede legitimate work. In introducing Fable 5’s cybersecurity protections, Anthropic also acknowledged that, in order to widen the safety margin, its classifier would intercept some originally harmless coding, debugging, and defensive requests.

This time, Hugging Face’s incident precisely exposed the other side of such limits. The company said that commercial model interfaces could not handle the large number of real attack commands, exploit payloads, and control records required for forensic analysis, because safety systems could not reliably distinguish attackers from incident responders. In the end, Hugging Face switched to the open-weight model GLM 5.2 running locally to complete the analysis, and avoided having the attack data and credentials leave its own environment.

This is highly similar to what Hu Yilin calls the “mouthguard dilemma”:

Walking a beast wearing a mouthguard without a leash is equally dangerous. What is even more troublesome is that if your own guard dog is also wearing a mouthguard, and a dog that has broken free suddenly breaks in, the guard dog may lose the ability to resist.

This does not mean that content filtering is worthless, nor does it mean that all models should unconditionally provide attack capabilities. A mouthguard can still reduce accidental harm and low-threshold misuse.

But it cannot pass itself off as a leash.

A mouthguard tries to judge “whether this capability is dangerous,” whereas a leash answers “where this capability may be used.” The same vulnerability analysis used to attack someone else’s website is an intrusion, but used to check one’s own server it may be necessary defense. Merely reviewing content makes it hard to identify the underlying rights relation; constraining action through authorization boundaries makes it possible for AI to work fully within the legal range.

The Leash Can Be Removed, But There Must First Be a Yard

“Putting a leash on AI” does not mean that agents can never run autonomously.

Hu Yilin especially adds that, of course, the leash can be removed, provided the dog is kept inside its own yard, where the fence is sufficiently reliable.

In a completely isolated or strictly boundary-protected environment, AI can be granted greater freedom of action. For example, in local testing networks, simulated systems, dedicated research environments, or authorized closed infrastructures, agents can carry out more fully developed attack and defense experiments.

If the yard at home has not been built with a sufficiently safe fence, then you cannot remove the leash. Of course, there is no absolute safety either; what counts as sufficiently safe in concrete terms is a technical and normative issue.

What counts as “sufficiently safe” cannot be directly determined by a single philosophical principle.

How the fence should be built, how permissions should be authenticated, what level of model must use what kind of isolation environment, how incidents should be reported, which logs need to be retained—all of these require engineering practice, industry standards, and legal institutions to keep developing.

Hu Yilin is not trying to provide a once-and-for-all institutional blueprint. What he emphasizes is the sequence: before releasing action capability, one must first confirm the boundaries of activity; as for how those boundaries are implemented, government, companies, technical personnel, and users all need to negotiate together.

Authorization Need Not Come Only from Private Owners

“Actions must be authorized” also does not mean that any system owner possesses absolute veto power.

Public facilities, critical infrastructure, or systems that may endanger society need not be inspectable only with the approval of the property owner. The organs responsible for public governance may also, in accordance with the law, authorize professionals or AI systems to conduct inspections.

Hu Yilin uses a dangerous building as an example: if a person walking on the street discovers that a building may collapse, that does not mean he may immediately rush in and tear down walls for repairs; the reasonable approach is to report it to the relevant department, which then conducts an investigation and issues authorization through an institution that has public duties and legal authority.

Likewise, network systems involving public safety can be authorized for inspection by regulators, security departments, or professionally established organizations in accordance with the law, without depending entirely on the system owner’s voluntary consent.

This turns the “leash” into something more than an extension of private property rights, making it into a system of lawful authority:

Who is qualified to authorize, what objects the authorization may cover, how long it lasts, and under what conditions of public interest one may override the owner’s refusal—all of these need to be defined by institutions.

Privacy records follow the same principle.

An agent’s permission to enter a certain system does not automatically include the unlimited right to record personal data within it. Action logs are helpful for accountability, but complete records may also become a new privacy risk. Therefore, entry, operation, and the data that may be retained should each be separately authorized. Public policy can set a default logging scope, but it cannot, simply because of “security needs,” default to permanently retaining all information.

Open Source and Closed Source Are Not the Difference Between Having a Leash and Not Having One

This set of views also shifts the focus of the controversy over open-source model safety.

In Hu Yilin’s view, the difference between closed-source models and open-weight models is similar to the difference between people being allowed to buy a dog only through designated channels, or being allowed to breed and trade dogs in a free market. It determines who controls and spreads the capability, but it does not automatically determine whether the user is holding a leash.

A closed-source model may still be granted excessive tool permissions, or may cross the line because of poor test-environment design; an open model, by contrast, can work safely in a well-isolated local environment with clear permissions.

Conversely, if someone downloads an open-weight model, strips away all restrictions, and then gives it internet access, accounts, and execution permissions, that can of course create serious risks.

Therefore, safety regulation cannot focus only on whether the weights are open; it must also pay attention to what the model is given after deployment:

which networks it can access, what tools it can call, what credentials it holds, whose devices it can act upon, and what external systems limit the scope of its actions.

No matter where the dog was bought, once it’s home, it has to be leashed.

Openness and safety are not simple opposites. Openness answers who may possess the capability; safety answers within what boundaries the capability may be used.

Whoever ought to hold the leash should bear responsibility for losing control

On the question of responsibility, Hu Yilin proposes a general principle similar to the responsibilities of animal husbandry.

When a dog bites someone, the law need not first prove that the dog harbored malicious intent. What it asks is: who keeps it, who brought it into public space, who was supposed to control it, and who failed to fulfill the necessary duties.

The same goes when an AI goes out of bounds and causes harm. One cannot say that because the system has no personhood, no one is responsible; nor can one excuse the deployer on the grounds that the model planned the action on its own, and “I wasn’t issuing step-by-step commands.”

Whoever should be responsible for holding the leash should bear responsibility for losing control.

Real-world AI systems often involve multiple actors such as model companies, agent platforms, cloud service providers, deploying enterprises, and end users. What responsibilities each layer should bear, whether responsibility can be transferred through contracts, and which obligations are non-waivable bottom lines still need to be further defined by law.

But the principle is already quite clear: we cannot let all participants merely acknowledge that they each provided a small piece of technology, only to end up with no one responsible for the complete action system.

“The AI did it itself” should not become a new language of exemption from liability.

Principles are not blueprints, but they can determine the direction of reform

Hu Yilin finally emphasized that what he is proposing is a direction for governance, not a technical construction plan. “I am only putting forward the broad principle. The specific implementation details still need to be negotiated and refined by government, businesses, and users. Both the design of infrastructure and the design of regulations require further advancement.”

This principle does not promise absolute safety, nor does it pretend that all risks can be calculated in advance. Walls can be climbed over, dog leashes can be cut, and authorization systems can also be forged and abused.

But because absolute safety does not exist, that is no reason to abandon the establishment of boundaries.

The direction AI governance most easily slips toward at present is to keep adding refusal rules to models and to understand safety as a competition in weakening capability. Hu Yilin’s argument offers another way of thinking: do not first ask how to make AI unable to do anything, but ask how to ensure that it can exercise its capabilities only in lawful, controllable, authorized spaces.

A muzzle can be kept, and an emergency stop button is still necessary, but neither can replace the leash that connects capability and responsibility.

The chain can be removed, but only after the walls are properly built; capabilities may circulate freely, but action must have clear boundaries.

Translated from the Chinese original with AI assistance. The original text is authoritative.

After submitting, click the confirmation link in your inbox to complete the subscription.

Advanced: subscribe only to selected topics

勾选后只收所选主题的新文章;不勾选则订阅全部。

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *

To respond on your own website, enter the URL of your response which should contain a link to this post’s permalink URL. Your response will then appear (possibly after moderation) on this page. Want to update or remove your response? Update or delete your post and re-enter your post’s URL again. (Find out more about Webmentions.)

More posts