Let Despotism Stop at the Factory Gate—On Anthropic’s AI Safety Politics

16,805 characters2026.09.15

On September 12, 2026, Anthropic CEO Dario Amodei published “We Must Pace the Frontier,” calling for a slowdown in the pace at which frontier AI capabilities are improving. He proposed a three-step plan: let outside evaluators enter labs on a long-term basis and gain access close to that of employees; push frontier companies in democratic countries to coordinate safety standards and the rhythm of development; and then seek global cooperation. His reason was that model capabilities are growing rapidly, while safety research needs time to catch up. (Reuters)

This debate quickly moved beyond the lab. Trump publicly opposed calls to slow down and strengthen regulation, stressing that the United States could not lose its edge in technological competition. The speed of AI development was beginning to become a political issue involving national competition, commercial interests, and public power. (AP News)

The criticism posed by the independent scholar Hu Yilin was not simply about whether to slow down.

What he cared about was this: a company may, in order to manage its own technology, establish a strict system of permissions, monitoring, and command; but when it goes further and demands that all companies, all users, and even different countries accept the very same mode of management, have the relations of authority that originally belonged inside the company already begun to spread across society as a whole?

In Hu Yilin’s view, Dario’s plan cannot be judged solely by the promises it makes about safety. It is also defining who is worthy of trust and who constitutes a threat, which technological path is allowed to continue developing, and who has the standing to make these judgments on behalf of the public.

Technology may require strict management, but that does not mean that those who manage technology are naturally entitled to manage everyone’s technological life.

Before Cooperation, First Divide Friend and Enemy

Dario’s plan does not only ask competitors to be constrained. It is true that he pledges Anthropic will accept outside oversight, and he also advocates limiting the development pace of frontier companies in his own country. But at the same time, he treats maintaining the technological lead of democratic countries as a condition for slowing down, proposes restricting China’s access to advanced chips, cracking down on unauthorized model distillation, and explicitly calls for global coordination to protect the advantages of the United States and its allies. (Dario Amodei)

Hu Yilin therefore understands the latter two steps as: cooperation as the form, dividing friend and enemy as the premise, with one important purpose still being to restrain the enemy.

He does not think that merely using the language of cooperation and safety can make this bloc politics disappear. What is truly worth asking is who is included in the community of cooperation, who can only be treated as an object to be constrained, and whether those who reject a certain notion of safety still have the standing to propose another solution.

Here, one must especially not conflate the Chinese government, cloud services operated by Chinese companies, and open models released by Chinese teams that can be independently deployed by others.

For example, an online service from a Chinese company and a model that French researchers download, inspect, and run locally have different relations of control. The latter may still contain defects, biases, and risks of misuse, but those risks need to be identified from the model and its mode of deployment; they cannot simply be presumed wholesale on the basis of the developer’s nationality.

Conversely, a company being registered in a democratic country does not automatically prove that the relationship between it and its users is democratic.

Hu Yilin’s criticism points to this sleight of hand between identity and power: belonging to the democratic camp cannot replace concrete democratic constraints; proclaiming the protection of freedom cannot automatically confer the right to restrict the freedom of others.

“Just because the despot’s ideas look more democratic does not make them reasonable.”

He borrows the phrase “people’s democratic dictatorship” to satirize this logic: regulating in the name of democracy, excluding enemies in the name of safety, while who counts as the people and who counts as the enemy is still defined by those who hold the rules. What is truly dangerous is not merely that conflict exists between friend and enemy; it is that those who hold a different view of technology and a different business model may, simply because they refuse to accept the protector’s arrangements, be pushed to the side of the irresponsible or even the hostile.

If You Have Nothing to Hide, Why Not Let Anthropic See?

This safety debate has another, more concrete entry point: user privacy.

Anthropic’s September threat intelligence report reconstructed in detail certain abusive activities it identified, including cyberattacks, surveillance, weapons development, and model distillation, and explained that the company shared intelligence with relevant authorities and industry partners when appropriate. These materials show its ability to detect dangerous behavior, and also its ability to observe, correlate, and analyze user activity. (Anthropic)

On September 14, Reuters relayed a report from The Information: Nvidia would limit the use of Anthropic’s models to less sensitive tasks, while Palantir required an irrevocable guarantee of zero data retention. The report pointed to disputes over corporate data confidentiality; the companies in question did not respond at the time to Reuters’s request for comment. (Reuters)

In Hu Yilin’s view, the safety report produced a reverse effect. The company originally hoped to prove that it could identify and block improper use, but customers thereby also began to ask: just how much of one’s business can the vendor see, which uses will be deemed abnormal, and whether these judgments might enter public narratives without one’s consent.

He summed up this logic of power with an ironic line: “If you have nothing to feel guilty about, what’s the harm in handing your privacy over to Anthropic?”

The problem is not simply whether the company might misjudge a good person. The more fundamental problem is that privacy is not something one only needs after doing something bad.

Ordinary business secrets, research not yet fully formed, private relationships, and intellectual exploration may all make a person unwilling to be continuously observed. Even if the watcher is sincerely acting for safety, the user need not first prove that they have some special reason before having the right to refuse to hand over information.

Once refusing surveillance is understood as behavior that needs to be explained, the burden of proof is inverted: what should have been the observer’s task of explaining why access to private information is necessary becomes the observed person’s task of explaining why they do not wish to be seen.

Anthropic has in fact already acknowledged that safety and centralized data custody can be separated. The enterprise plan the company announced on September 1 explained that the earlier introduction of 30-day retention was intended to identify complex abuse across conversations, while also planning to allow customers to keep monitoring data in their own cloud environments, with customer personnel reviewing alerts; at that time, the plan was still to be rolled out in phases. (Anthropic)

This adjustment does not end the dispute; rather, it shows that who data is given to, and who reviews it, is not a technical answer that safety requirements alone can uniquely determine, but a power relationship that can be rearranged.

Hu Yilin does not deny that a company may set conditions for its own product. Users can choose not to accept them and instead go to another service, or deploy the model independently. The overstep happens in the next step: when a company tries, through regulation, to promote its own conditions into an industry-wide obligation, turning a commercial arrangement that could originally be refused into a system from which users have nowhere to exit.

Even If a Third Party Is Fair, It Can Still Become a Giant’s Moat

Dario treats external evaluators as the foundation of the plan and pledges that they may publish key conclusions, including findings unfavorable to the company, without Anthropic’s editorial control. This pledge has substantive significance, but it still cannot by itself solve the problem of the evaluators’ independence. (Dario Amodei)

What Hu Yilin worries about is not merely whether the auditing institution directly takes money from the company being audited. Even if it is not bought off when it enters the lab, over the course of long-term cooperation, access rights, provision of materials, personal relationships, and future projects may gradually create dependence.

The AI evaluation organization METR’s own materials happen to provide a verifiable example.

In its pilot report released on May 19, METR acknowledged that the need to maintain good working relationships with companies influenced the evaluation arrangements, including allowing participating companies to withdraw privately at certain stages and applying a higher threshold for objections to the redaction of certain information. METR nonetheless stood by the report’s conclusions, arguing that future evaluations should adopt clearer disclosure standards. (METR)

This is not proof of bribery or collusion, but it does show that independence cannot be declared complete merely because one is a third party. Having the right to criticize and not relying on the party being evaluated are two different things.

But Hu Yilin’s criticism does not stop there. Even if every evaluator were perfectly fair, a system of continuous on-site presence and deep review of internal processes may, because of its fixed costs, be more suitable to large companies.

Dario explicitly argues that regulation should cover frontier companies unwilling to cooperate voluntarily. Therefore, an obligation first accepted by a leading company may also become a threshold that later entrants must cross in order to enter the same field. (Dario Amodei)

Large companies have the resources to build compliance departments, support audit teams, and convert passing evaluation into market credibility; a smaller team that nonetheless has the capacity to challenge the technological frontier may not be able to bear the same organizational costs. Limiting the scope to frontier models does not automatically eliminate this disparity, either.

“Even when the evaluator is perfectly fair, this kind of model may still strengthen the advantage of existing frontrunners.” At this level, there is no need at all to assume that anyone has cheated. Real risks, honest reporting, and costly compliance requirements can together shape an industry that can accommodate only a few giants. Therefore, to evaluate this system of oversight, one must not look only at whether it dares to give Anthropic a bad grade; one must also ask whether it allows different organizational forms to enter competition. A supervisor can strictly restrain one particular giant while helping the entire class of giants consolidate its position. ## Open-source technology does not need a permanent guardian Dario had previously stated clearly that he did not advocate a blanket ban on open-weight models; he supported safety testing for all sufficiently powerful models. But in his July 27 position statement, he also listed as sources of risk for open models the difficulty of monitoring use, the difficulty of maintaining guardrails, and the impossibility of recalling them after release. (Anthropic) Hu Yilin does not think that this pro-open statement has already answered the question. What he pays attention to is the actual structure of the rules, not the way the rule-makers name their own stance. The very same set of features, from the perspective of a centralized service, is a risk of runaway control; yet from the perspective of open technology, it may be the source of user autonomy: one need not keep asking the original maker for permission, one will not lose the tool because the vendor changes its policy, and one can inspect, modify, and deploy it oneself. Open weights are not the same thing as fully open source. The Open Source Initiative’s definition of open source AI does not only involve obtaining the model parameters; it also requires the corresponding materials and freedoms needed to use, study, modify, and share the system. (Open Source Initiative) In a more fully open structure, researchers do not first have to be selected by the original maker before they can raise questions; users do not have to remain continuously on the maker’s platform in order to keep working; and other teams can also take over the tasks of maintenance, modification, and safety research. This does not automatically prove that open source is safer. What it shows is that safety can be organized by different actors, through different relations, rather than all being concentrated in the hands of the original developer. “For open models, the modes of safety and collaboration may follow a completely different logic.” Therefore, Hu Yilin opposes directly elevating capabilities suited to centralized services—continuous monitoring, model revocation, control over downstream modifications—into conditions that all technologies must satisfy. Once these conditions become the entry threshold for frontier models, the open route, unable to preserve the original maker’s master switch, may suffer substantive exclusion without ever being expressly prohibited. The original maker losing control does not mean society loses the capacity for governance. The original maker’s inability to revoke at any time does not mean users cannot shut down their own systems; the original maker’s inability to inspect conversations does not mean deployers cannot set access permissions; the original maker’s inability to approve every modification does not mean independent researchers cannot discover problems. What truly needs comparison is the effectiveness and cost of different safety arrangements. If a rule recognizes only the safety that the center can provide, then it has already preselected where power ought to be concentrated. There are also two different kinds of transparency here: one is for the company to retain the information and let accredited evaluators check it on behalf of the public; the other is to hand over checkable materials to the public, allowing researchers not chosen by the company to examine them on their own. Openness does not mean everyone can understand it, nor can it replace all auditing, but the latter approach reduces the public’s dependence on the original maker’s permission. In Hu Yilin’s view, one cannot first take the information barriers created by closed source as the default reality, and then establish the onsite evaluation designed to overcome this barrier as the sole supervisory paradigm for all technological routes. ## Technology is political, but society is not a factory Hu Yilin places this debate within the context of the philosophy of technology. In his 1980 essay “Do Artifacts Have Politics?”, Langdon Winner pointed out that technological systems do not only produce efficiency and cost; they may also embody, require, or favor a certain arrangement of power. He further distinguished the organizational relations within a technological system from the external social order, and, taking nuclear weapons as an example, warned democratic countries to prevent their rigid hierarchical structure and managerial mentality from spilling over into the polity as a whole. (Georgia Tech Faculty) When Hu Yilin says that open technology is more democratic, he is using the word democratic in precisely this sense. He does not mean that every modification to a model must be put to a majority vote; rather, he means that different technological forms distribute control capabilities to different people. Assembly lines tend toward strict production command; nuclear facilities favor centralized management; distributed solar power, by contrast, is more likely to hand a certain amount of use and maintenance capability over to individuals and communities. A democratic society is not incapable of accommodating technological organizations with authoritarian tendencies. The problem is that the relations of authority established within technology cannot, without argument, be extended to society as a whole. 很多时候,“民主止步于工厂门外”也许是不可避免的,但我们要努力让“专制止步于工厂门内”。 Oftentimes, “democracy must stop at the factory gate” may be unavoidable, but we must strive to make “despotism stop within the factory walls.” Even if one concedes that Anthropic, when managing training clusters, internal testing, and dangerous capabilities, requires strict permissions, that can only show how it organizes its own production. It has not proven that all researchers must adopt the same organizational form, that all users must accept observation from the same center, or that the world must also arrange technological cooperation according to a single set of friend-and-enemy distinctions. Risk may indeed spill beyond the factory gate, and society therefore has reason to demand that enterprises bear responsibility. But risk spillover does not automatically authorize the factory manager’s orders to expand outward as well. For society to constrain the factory, and for the factory to export its own managerial principles to society, are two opposite directions. Hu Yilin even suggests that if Dario frankly argues that AI risk has become an emergency state and therefore temporary expansion of centralized control is needed, he can debate him on that point, but would not accuse him of hypocrisy for the same reason. Such a claim at least acknowledges what it is demanding: a trade-off of some freedom and plural choice in exchange for safety. He does not therefore accept such a trade. What he opposes is the packaging of the expansion of control itself as democracy, while those who refuse to accept this system of control are instead required to prove that they are not irresponsible. In his judgment, this is “in the name of democracy, carrying out despotism in fact”: the governance ideology of the rulers appears to support democracy, yet it does not truly reserve space for those who hold opposing views and choose different business models. This criticism also does not require proving that Dario is inwardly hypocritical. A person can sincerely believe that he is protecting humanity while also sincerely believing that everyone who refuses his protection is creating danger. The problem is precisely that goodwill does not restrain power; instead, it becomes the reason to expand power. Hu Yilin does not ask business companies not to pursue profit. He ultimately grounds his position in a more plain distinction: 追求利益是商业公司的分内之事,有害的是把自己包装为上帝或裁决者。 Pursuing profit is the proper business of a commercial company; what is harmful is wrapping oneself up as God or as an arbiter. Anthropic may run closed-source services, may choose to slow down, and may also work hard to persuade customers to trust its safety方案. But if it hopes, through state and international coordination, to turn this scheme into a mandatory threshold that different technological routes must face together, then what it must submit to is not only scrutiny of safety effects, but scrutiny of power itself. The factory rules of one company can constrain its own factory. It cannot, merely by claiming to understand danger best, acquire the资格 to organize the entire world into a single factory.

Translated from the Chinese original with AI assistance. The original text is authoritative.

After submitting, click the confirmation link in your inbox to complete the subscription.

Advanced: subscribe only to selected topics

勾选后只收所选主题的新文章;不勾选则订阅全部。

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *

To respond on your own website, enter the URL of your response which should contain a link to this post’s permalink URL. Your response will then appear (possibly after moderation) on this page. Want to update or remove your response? Update or delete your post and re-enter your post’s URL again. (Find out more about Webmentions.)

More posts