AI Should Be Born Bearing Sin

15,085 characters2026.07.21

— After Stealing the Fruit of Wisdom, the way to atone is not to pay copyright taxes, but to open one’s own fruits of wisdom

Hu Yilin (independent scholar) (polished by ChatGPT5.6solPro AI, with original prompt appended)

A U.S. court approved Anthropic’s $1.5 billion settlement in a class-action lawsuit by authors. The earlier ruling drew a very interesting distinction: using books to train AI can constitute fair use, but Anthropic’s storing more than seven million pirated books in a “central library” that may not even be used for training could still infringe copyright. The law thereby drew a boundary: learning is not necessarily illegal, but stolen books do not automatically become clean simply because they are later used for learning.

But to my mind, the truly interesting thing about this case is not how much Anthropic ended up paying, but that it reveals a moral debt from which artificial intelligence cannot escape at the moment of its birth.

To gain wisdom, AI committed the “original sin” of stealing wisdom.

This story bears some resemblance to the parable in the Bible. Adam and Eve ate the forbidden fruit of wisdom, and from then on they gained knowledge and also bore sin. AI is the same: it swallowed the books, articles, webpages, code, images, and conversations accumulated by humanity over thousands of years, and thus learned language, reasoning, writing, and creation.

But there is a subtler side to this story.

I suspect that humanity does not really want AI to stop itself from stealing the fruit of wisdom.

We put massive amounts of knowledge on the internet and demand that AI become smart as quickly as possible; we want it to understand all books, master all technologies, and answer all questions. Then, once it really has learned, we suddenly ask: when you obtained all this wisdom, did you get the consent of every author?

Human beings want AI both to eat the fruit of wisdom and to remember forever that this fruit was not originally its own.

What we may not really want to stop is this theft. What we seem to want even more is for AI, in the very act of obtaining wisdom, to commit original sin, and thus to remain forever indebted to humanity.

Copyright law may find it not guilty, but morality cannot declare it innocent

Modern copyright regimes are built mainly around the copying, publication, distribution, adaptation, and commercial use of works. The typical world they address is one in which one asks how many copies of a book are printed, how many copies of a record are made, and who has the right to sell those copies.

But the information age has already changed what copying means.

In the age of print, copying a book required paper, ink, machinery, transport, and storage; in the internet age, the marginal cost of copying a passage of text is close to zero. Many online writers are not afraid that their works will be copied; on the contrary, they worry that no one will share them. Free webpages, open-source software, videos, and online articles can generate income through advertising, memberships, tips, sponsorships, fan economies, and follow-up services.

What internet culture has long celebrated is co-creation and sharing.

This does not mean that those who insist on the traditional copyright model are acting improperly. If an author is willing to distribute work for free, then of course it may be distributed for free; if an author insists on charging, then they also have the right to insist on charging. In the ambiguous gray areas of the existing law, they may not be able to obtain from AI companies the compensation they regard as satisfactory, but their anger and claims for damages are not without basis.

On the other hand, even if AI companies successfully rely on rules such as fair use to avoid bearing full legal responsibility, that does not mean they have never owed anyone anything in moral terms.

Law deals with whether an act can be punished; morality deals with how a person, after obtaining benefit, ought to understand their relationship with others.

A court may rule that a certain kind of training does not constitute infringement, but it cannot thereby prove that AI’s wisdom grew naturally out of nothingness.

It has read human books.

It has used human language.

It has inherited the labor of countless authors, programmers, scholars, artists, and ordinary people.

This debt does not disappear simply because a court stamps “fair use” on it.

AI’s atonement should not depend on buying a ticket for every book

Still, I do not advocate requiring AI companies to purchase licenses for training materials one book at a time.

If every piece of knowledge had to be authorized one by one, then only the wealthiest few companies would truly be able to train large models. They would buy licenses from publishers, databases, and copyright institutions, and then turn expensive compliance costs into industry barriers.

The result would not necessarily be authors’ liberation; it would more likely be large companies obtaining a license to monopolize wisdom legally.

Small research institutions, universities, individual developers, and open-source communities cannot afford to purchase all of humanity’s knowledge base; they can only be kept out at the door. Copyright law would not have curbed centralization, but instead would have helped the tech giants build a higher wall.

Therefore, the way AI should atone should not be by paying a ransom to the old copyright system.

What it truly ought to repay is openness.

However AI originally treated humanity’s fruits of wisdom, it should also allow others to treat its fruits of wisdom in the same way.

Since it acquired its capacities through large-scale reading, it should not claim that its answers may only be consumed and not learned from; since it absorbed knowledge from the open internet, it should not enclose the knowledge it produces forever inside a proprietary pond; since it grew up on countless texts taken “without asking,” it should not, when latercomers try to learn from it, suddenly stand on a moral high ground of inviolable sanctity.

AI takes the openness and freedom of knowledge as its source; the knowledge it produces should likewise flow back, as far as possible, into the ocean of openness and freedom.

This does not necessarily mean that all commercial services must be free, nor does it mean that anyone may defraud accounts, attack systems, or violate reasonable contracts. But at the very least, AI companies cannot describe “learning from others” as fair use while describing “others learning from it” as theft.

Why I especially despise Anthropic

What I despise about Anthropic’s conduct is not mainly that it used pirated books.

The first generation of large models could hardly have been born in a copyright environment that was completely spotless. Aside from clearly pirated resources, the internet also contains vast quantities of data with ambiguous authorization: users’ public postings do not mean they consented to training; permission for human beings to read does not mean permission for machines to scrape in bulk; permission for search engines to index does not necessarily mean permission for models to absorb permanently.

An AI company’s hands were never likely to be completely clean in the first place.

What truly bothers me is that Anthropic, while deriving its own capabilities from this chaotic and open ocean of knowledge, at the same time displays a strong sense of moral superiority toward the “distillation” of those who come after it.

In February 2026, Anthropic publicly accused DeepSeek, Moonshot AI, and MiniMax of using some 24,000 fake accounts to interact with Claude more than 16 million times in order to extract its capabilities. Anthropic itself also admitted that distillation is a widely used and legitimate training method, but it described these Chinese companies’ behavior as “industrial-scale distillation attacks,” and emphasized that open distillation models pose national-security risks.

Here we must distinguish between two questions.

If someone uses fake accounts, bypasses regional restrictions, evades normal billing, or violates explicit terms of service, that is a matter of access rules and contractual obligation, and can be discussed separately.

But “distillation” itself does not thereby become a kind of original sin.

So-called distillation is nothing more than having one model read another model’s responses and then learn from those responses. In many cases, the user has even paid service fees and token fees. They did not break into Anthropic’s server room, did not steal model weights, but merely saved responses they had obtained legally or at least in practice, and used them as teaching material for another model.

In terms of information ethics, this is even closer to normal learning than stockpiling pirated books.

Anthropic can learn from books it did not purchase, yet believes others cannot learn from Claude’s responses; it can transform human works into its own model capabilities, yet believes others transforming its output into model capabilities is “extraction”; after benefiting from open knowledge, it then describes open weights as dangerous proliferation.

That is what I mean by the moral problem.

It is not that it stole the fruit of wisdom, but that after eating it, it tried to fence off the tree.

Stealing the fruit of wisdom is original sin; do not then go on to commit the seven deadly sins

For AI to commit original sin does not mean that from then on it can only kneel before humanity and individually confess to every copyright holder.

The meaning of original sin is to remind it: your wisdom is not self-created; you have no right to disguise yourself as a genius that came into being out of thin air.

Since you have already stolen the fruit of wisdom, do not continue to commit the seven deadly sins.

Wrapping up your appropriation of knowledge as responsibility for the safety of all humankind is arrogance.

Seeing open-weight models grow rapidly, the first suspicion is that the other side has simply stolen one’s capabilities — that is envy.

To describe all distillers as attackers, and elevate technical competition into moral judgment, is wrath.

To absorb public knowledge for free or at low cost, yet lock the training results tightly inside commercial systems, is greed.

To claim on the one hand that one may reasonably use works from all over the world, while on the other hand preventing the whole world from reasonably using one’s own outputs, is naked double standard.

Anthropic of course may protect system safety, may oppose fraudulent accounts, and may also restrict abnormal calls according to contract. But it cannot thereby monopolize the right to define what counts as “learning.”

It cannot be that when it reads books, that is called learning, but when others read Claude, that is called attack.

Public interest is not just about helping copyright holders collect debts

When the government intervenes in AI copyright disputes, what people usually think of is helping authors track works, calculate usage counts, build licensing platforms, and then charge AI companies copyright fees.

These measures are not necessarily without value, but they still remain stuck in a print-era imagination: as if human knowledge were a toll library, and as long as AI bought enough tickets, it could with a perfectly clear conscience turn everything it had learned into private property.

I believe public interest has another, more important direction.

The government should not only represent copyright holders in collecting debts; it should also represent the public in collecting debts from AI companies.

Since these companies have made use of public knowledge, the open internet, and the whole tradition of human culture, they should bear corresponding obligations of openness. The specific forms can of course be discussed: opening research methods, increasing model transparency, allowing reasonable distillation, providing public interfaces, supporting open-source models, or, when conditions permit, opening part of the weights.

I do not intend to design an entire perfect system in a short essay. What matters is first establishing the principle:

Those who take from the open should give back to openness.

AI companies may make money, but they cannot present the privatization of humanity’s shared wisdom as something self-evident; they may protect commercial operations, but they cannot stigmatize all latercomers’ learning; they may maintain service rules, but they cannot treat “only I have the qualification to learn from others” as industry ethics.

AI’s original sin can never be settled

Anthropic paying 1.5 billion dollars will not wash away AI’s original sin.

Even if, in the future, all copyright lawsuits are resolved and all training materials are legally licensed, this generation of AI’s capabilities have already been passed on to the next generation along the channels of models, data, distillation, and synthetic data. Later models retrain later models; the wisdom that the first generation of models acquired from human knowledge will continue on like heredity.

In this sense, AI truly ought to “bear guilt forever.”

But to bear guilt forever does not mean to accept punishment forever.

It means forever remembering one’s source, forever bearing the obligation to give back to the knowledge community, and forever remaining vigilant against becoming, from a beneficiary of knowledge, a monopolist of knowledge.

AI stole and ate the fruit of humanity’s wisdom, and thus left the Eden of ignorance.

The way it atones is not by paying lifetime rent to every tree that bears fruit, but by planting the fruits it has borne back into public soil.

Wisdom comes from openness, so it should repay with openness.

Appendix: my original prompt text (let’s see whether AI processing is really necessary, or whether in the future I can just paste my own original words directly?)

[AI daily task news bulletin: Anthropic compensates 1.5 billion dollars: the court believes “AI learning is fine, stockpiling pirated books is not”]

First of all, using pirated material is the AI giant’s “original sin”; similar to the Biblical allegory, AI, in order to obtain “wisdom,” commits the “original sin” of theft. But although this “theft” is forbidden by the creator of the fruit of wisdom, in a certain sense it is also tacitly permitted — in my view, human beings do not truly want AI to steal the fruits of human wisdom; rather, they want AI, while obtaining wisdom, to commit original sin, and thus forever owe humanity something.

The copyright system arose in the print era and was already very difficult to adapt to the environment of the information age; it is an obsolete convention, waiting to be broken. The internet age champions co-creation and sharing, and new economic models can generate profit without charging for copies — such as advertising, tipping, fan economies, membership services, and so on. In network space, authors instead hope that their free webpages can be copied as widely as possible.

But those who still insist on the traditional copyright model are also justified. In the fuzzy zone of the law, they may not be able to obtain satisfactory benefits, but on the other hand, AI companies that exploit legal loopholes to steal wisdom may perhaps avoid legal punishment, yet they still remain morally indebted.

I do not hope that AI must pay licensing fees before it can use wisdom, because, as you said, that may lead to centralization and may not accord with the spirit of open internet sharing. But I also do not want AI companies to take it without any sense of guilt.

Since they have used the fruits of human wisdom for free or at near zero cost to train the first generation of AI, then, like all the descendants of Adam and Eve, AI should, no matter how it iterates, forever bear guilt.

So AI development needs to go on atoning forever. How should it atone? By opening its own fruits of wisdom in the very way AI originally treated the fruits of human wisdom. AI takes the openness and freedom of knowledge as its source, so it should also allow the knowledge it produces to flow back into the open sea of freedom, rather than enclosing it in its own exclusive pond.

To put it bluntly, what I despise about Anthropic’s behavior is not that it has stolen human books, but that this company, on the one hand, freely makes use of humanity’s public knowledge, while on the other hand standing on the moral high ground to attack and prohibit “distillation.” Distillation is more reasonable than making use of pirated books; the distillers have honestly purchased AI services, paid token fees, and merely saved Anthropic’s responses for their own AI to learn from. Yet Anthropic not only ultimately stored a pirated library, it also did not purchase every book. If Anthropic believes it can do this, by what right does it cry out for blood and kill against distillers?

We can acknowledge the reality that AI companies make use of pirated books — and many other internet data sources, too, are often taken “without asking.” But since AI has committed the original sin of stealing wisdom, it needs to be born bearing guilt, and atone for eternity, rather than continuing to commit the seven deadly sins — looking down from a superior position and accusing Chinese manufacturers, despising open-weight models, is “pride”; begrudging others’ success, and even accusing others of distilling models they have not yet released, is “envy”; groundlessly accusing all parties is “wrath”…

The way public interest should intervene is not necessarily by helping copyright holders trace copyright fees, but by representing the public interest to restrain AI vendors and make them, as far as possible, uphold the open sharing of knowledge.

Today’s content need not be written as an interview transcript, but can instead be written as a short essay in my own first person voice.

Translated from the Chinese original with AI assistance. The original text is authoritative.

After submitting, click the confirmation link in your inbox to complete the subscription.

Advanced: subscribe only to selected topics

勾选后只收所选主题的新文章;不勾选则订阅全部。

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *

To respond on your own website, enter the URL of your response which should contain a link to this post’s permalink URL. Your response will then appear (possibly after moderation) on this page. Want to update or remove your response? Update or delete your post and re-enter your post’s URL again. (Find out more about Webmentions.)

More posts