Actual title: "Unsealed Briefs in Authors’ Case v. Microsoft/OpenAI: Top Execs Knew Their Mass Book Piracy Was Illegal And Would Put Authors Out of Work"<p>Twitter is also mentioned:<p>"356. In July 2020, OpenAI employee Ryan Lowe assessed the risk of continuing to use LibGen for the book-summarization project. Lowe wrote that he thought “there’s a >80% chance that we have some exchange of the form: ‘where did you get the books data?’” and “‘we can’t say’[.]” Nelson Decl. Ex. 325 at -315. Lowe estimated “a further ~40% chance that that leads to a moderate-sized Twitter kerfuffle that negatively affects the external perception of our work.” Id.
Lowe added: “if we’re fully okay with these potential outcomes, then I’m comfortable continuing using Libgen for the project.” Id."<p><a href="https://authorsguild.org/app/uploads/2026/09/Class-Plaintiffs-SUF-9.17.26.pdf" rel="nofollow">https://authorsguild.org/app/uploads/2026/09/Class-Plaintiff...</a>
I am so tired of people dismissing claims that these are plagiarism devices. The creators literally know they essentially stole this work.<p>This is literally a conversation where they’re deciding if they’re willing to take on the risk, then determining “yes.“<p>And the worst part? They were absolutely right. They have not suffered any real consequences. And when people try to bring up this flagrant plagiarism and theft, they are shouted down by AI evangelists.
The “hacker” community (for lack of a better term) has a longstanding and well documented skepticism towards the concept of intellectual property in general. “Information wants to be free” and all that.<p>Let he who has not downloaded from Annas Archive cast the first stone
I think we, as humans (not executives, a different creature altogether if you ask me) make a distinction between a hobbyist hacker, someone doing something for their own curiosity or someone who frees something for others to use freely as well (e.g., F/LOSS licenses, Creative Commons, etc.), vs. someone who uses the "Hacker Ethos" and then promptly builds their own moat where they solely can profit and excludes others from the freedoms they themselves enjoyed.
No, I think the overwhelming majority of copyrighted material should already be public domain. We should all be free to build on it as we wish. I don't begrudge OpenAI et al. for doing it at all. I begrudge copyright lobbyists for pushing for the rest of us to be unable to do the same. Because my position is consistent instead of being based on whether I like someone.<p>If it takes giant AI companies to show the extreme economic loss caused by maintaining the copyright farce, and to make it clear that we simply <i>cannot</i> continue to do so or we risk becoming economic vassals, then good. Them flouting the law is a good step toward reforming or dismantling it.<p>Aside: "it's fine if you're an individual but not if you're large" in this case sounds awfully self-serving. If you truly believe it's "theft" (I don't), then individuals stealing is still wrong. I don't see how that argument doesn't directly excuse e.g. retail theft or other antisocial behavior.
I'm all for copyright no longer being a thing. Until that's the case, I don't want the world in which AI companies mass-violate it but individuals still get punished for that.<p>I want the world in which fanfiction is completely legal and the best of it is sold in bookstores. I want the world in which projects emulating macOS in the cloud, <i>on non-Apple hardware</i>, are widely used and legal. I want the world in which the many video game decompilation and enhancement projects are 100% legal. I want the world where every single creative project someone wants to build that draws upon the work of others is legal.<p>And until we have that, if AI is going to mass rip off all our work and use it to compete with us, I hope copyright is one of <i>many</i> tools used to burn it to the ground.
Are they even using it to compete with us? There's plenty of people talking here about how it doesn't make any economic sense to run a local LLM. Like, they're providing inference so cheap that even when you have models that are just as good, it still doesn't make sense for you to run it.<p>Anyway, my point is keep your focus on the actual injustice. The problem is that individuals get punished, not that AI companies don't. Saying we should punish the AI companies is just saying that we should solidify the legitimacy of IP. This is an important moment to say it's clearly insane to keep this going.<p>You don't say that it's unfair that Snoop Dogg didn't end up in prison forever for his marijuana use, and that we ought to lock him up too; you say it's unfair that other people did.
> Are they even using it to compete with us?<p>Yes. They are literally pitching to investors and businesses that they can fire all of us.
And how's that going to work when they're already today commodified and you can already buy your own personal AI machines to run almost SOTA models for a couple thousand dollars?<p>If employers fire everyone because AI can just do it, why would somebody pay for their services when AI can just do it?
> Are they even using it to compete with us?<p>Ask anyone who has ever lost a job in a layoff that cites AI.<p>The (often <i>stated</i>) goal of AI is to be able to perform any and all human work, and cheaper than any human.<p>> Saying we should punish the AI companies is just saying that we should solidify the legitimacy of IP.<p>We are otherwise very likely to end up in a world in which <i>AI companies can do it and individuals can't</i>.
You can already do it yourself now though. The Qwen 32B models are quite capable and can be run on a ~$1400 GPU. I haven't gone through the trouble to set it up myself, but my understanding is image generation workflows are actually state-of-the-art local. i.e. local workflows are much better than the big cloud providers. My family's first computer in the 90s costed more than a modern AI machine <i>nominally</i>. I don't remember people screaming then about how computers were creating a moat that would lock individuals out of the economy despite actually be <i>less</i> affordable in real <i>and nominal</i> terms.
Thank you for putting my views into words better than I could. I apply the same thinking to patents. The enormous economic losses from patent trolls, normal litigation, licensing costs, and hell, general administration costs, it’s insane.<p>And you’re definitely right about the picking sides thing as well.
> Them flouting the law is a good step toward reforming or dismantling it.<p>I wish I could believe this, but I suspect it's just another instance of the law ceasing to apply to entities whose net worth has enough zeroes in it.<p>With how much money is sloshing around in the AI industry, there's almost nothing AI companies can do which has any likelihood of resulting in legal accountability. You'll much sooner see laws changed and/or reinterpreted than enforced.
> If you truly believe it's "theft" (I don't), then individuals stealing is still wrong.<p>We as humans are allowed to have nuanced views on complex matters.<p>If all IP is public domain, why would anyone ever write another book? They won't be able to sell it if it's immediately freely available.
People can pay for creation, not rent (i.e. patronage), or one could argue that there's a reasonable tradeoff with like a 5-10 year copyright. I don't think a 2016 or even 2000 book cutoff would materially affect the LLM training discussion. And research runs on government grants, so everything derived from public money is paid for already and should be public domain, meaning there is no knowledge cutoff for science, which is probably the more important area to have cutting edge knowledge.
Well, did books exist before copyright?
I don't see the word "immediately" in their comment.
Sounds like an argument against capitalism rather than intellectual property.<p>If no one <i>has</i> to work to live, how many would spend time writing instead of working at a job they hate?
>"it's fine if you're an individual but not if you're large"<p>That's not what the GP said, they were referring to someone who "builds their own moat where they solely can profit and excludes others from the freedoms they themselves enjoyed".<p>I don't want to speak for them but I guess that LLM companies which release exclusively open weights models, or better yet, open weights plus training pipeline sources, would not fall into this category.
Cool: an individual ignoring IP for their own purposes or individuals sharing things they own with others<p>Cringe: Microsoft, OpenAI, Anthropic, or Meta doing the same<p>None of the later should get to claim 'But we're hackers!' from an ethical standpoint.<p>They aren't -- they're shareholder-benefiting for-profit corporations (OpenAI obviously included).
The “for profit” isn’t even the entire story. The scale of the theft is unprecedented. It’s the combination between stealing <i>everything</i> and only for their profit that makes this impossible to defend.
Cool: an individual, or an organization, not believing in IP, reading what they want, and sharing what they want.<p>Cringe: an individual, or an organization, fierce defenders of their own IP and regular DMCA abusers, stealing the IP of others in order to sell it themselves.<p>There's a lot of twisting you have to do to make these two things the same. These people would happily deliver takedown requests to the original producers of IP if they knew they could get away with it. In fact, they <i>long</i> to.
Yeah. That difference is profit.<p>Someone ripping a copy of The Odyssey for watching in their own home is a loss of profit for the company's lawyers, and must be punished to the fullest extent of the law.<p>A company ripping off all of humanity to create profit for their company's lawyers is perfectly fine.<p>There's nothing magic here. Its money deciding the rules. Like it always has.
Basically. There's a difference between someone getting hunted after by RIAA asshats for downloading an album (people may be interested to look up Steve Albini's views of piracy) and an actual industrial industry of take-content-for-free
[flagged]
You're so right, he should've written:<p>"I think we as humans (I'm saying we as in probably the majority from my point of view and understanding of what the majority of people who visit this website likely think about this and how their ethical viewpoints likely align with my opinion that I am about to present however there may be a minority or some part of people that do not agree with what I am about to write therefore it is crucial that I make this clear before I go on with my point here)"
Fallacy that only serves corporate interests, as he points out.
No, just the humans.
I'm all for intellectual property reform, I'm still not a fan of the current rules being ignored for large companies directly competing with authors and artists, while ordinary people pirating are getting fined 220k for 24 songs [1].<p>[1]: <a href="https://www.independent.co.uk/news/world/americas/woman-fined-220-000-for-illegal-music-sharing-396126.html#:~:text=A%20Minnesota%20jury%20has%20ordered%20a%20woman,is%20expected%20to%20force%20her%20into%20bankruptcy." rel="nofollow">https://www.independent.co.uk/news/world/americas/woman-fine...</a>
I'd be dismayed to find that the "hacker spirit" is in any way compatible with venture capital raiding of anything not nailed down, in the pursuit of hoarding it in private data vaults only to be used for developing a product that removes all attribution and is jealously protected by corporate lawyers.
I don't think it's particularly hard to grasp that some people might consider there to be a difference between an individual downloading a few things for profitless personal use and a corporation downloading literally everything they can get their hands on to try to make money.<p>The law is not math; intent and outcomes matter, not just the abstract action taken in a vacuum.
What if I told you the guys in charge of OpenAI thought that the RIAA and MPAA should have sued individual, "profitless personal use" people?<p>You are looking for some simple explanation and the simpler one is, "greed." It's right there in the personal diary...
I think a "Information wants to be free" person would be fine criticizing companies for doing this, while still being morally consistent.<p>It would be an entirely different story if OpenAI was indeed open and the resulting model was available for all. The issue comes when taking information you haven't paid for, and then locking it into a machine you charge for.
Except… it costs a huuuuge sum of money to transform that tranche of inert information into a totally different useful form (an LLM).<p>That process of transformation, the expertise required to enable it, and the cost of then making it available to users is what is being charged for, no?
Either they can create a viable business paying for the information they use to create that business, or they can't.<p>A business doesn't have any innate right to exist.
And how much did it cost for the authors to produce the books in the first place? Everyone wants to believe that piracy is okay because the creators costs don't matter.
The last few years should make it abundantly clear to anyone with the gift to step away from a situation and analyze it independently, that the "moral righteousness" of the internet is actually just another collective of "What helps <i>me</i> is good, what hurts <i>me</i> is bad".
Plagiarism is something very different from skepticism towards the concept of intellectual property.<p>Plagiarism is the act of avoiding attribution or citation for personal gain. You can be a staunch anti-IP advocate and still believe plagiarism is ethically wrong.<p>LLMs are objectively bad at attribution and citation, therefore incur in plagiarism. Even if this is due to a technical limitation, it is still plagiarism.
><i>a longstanding and well documented skepticism towards the concept of intellectual property in general. “Information wants to be free” and all that.</i><p>Despite it's lofty rhetoric, OpenAI is neither free as in beer nor free as in speech. It is a gang of profit-motivated thieves. Efforts to pirate others works <i>in order to personally profit</i> is the antithesis of the "hacker community" approach to IP. (OpenAI appears to get quite upset when their own work is treated as they have treated everyone else's.[1])<p>1. <a href="https://www.fdd.org/analysis/2026/02/13/openai-alleges-chinas-deepseek-stole-its-intellectual-property-to-train-its-own-models/" rel="nofollow">https://www.fdd.org/analysis/2026/02/13/openai-alleges-china...</a>
I think there may be a different word than "hacker" for liberating information to repackage and resell through a big corp while contributing to make other sources less sustainable.
There's a difference between a person downloading a few works from Anna's archive, and a company mass-downloading them for profit.
20 years ago, a 19 year old downloads some Eminem.<p>He gets threatened with a scary letter, his family decides to pay about 3000$ to settle.<p>Disney has the nerve to run a don’t download music PSA in the form of a Proud Family episode. If you don’t know the Proud Family was a cartoon which attempted to address “black issues”.<p>Dang hommie, did you know that failing to respect the intellectual property rights of billion dollar corporations is literally worse than selling crack cocaine.<p>Think of the shareholders! Think of the missed profit projections!<p>But when billionaires need to effectively resell the IP of all of humanity, that’s just fine.<p>Since we live in wacky world, Suno which was trained off stolen IP counts Warner Music as one its partners.<p>The same Warner that was suing over music downloads a few decades ago.
Casting the first stone! Real world libraries exist and I don't need FSB PDF exploits, thank you.
Would be an intriguing calculus thinking they can hurt Western publishers enough to offset the enrichment of generations of thinkers and scientists. I usually chalk up both book and article archive sites as being citizen run.<p>Thankful to have always been surrounded by top-notch physical libraries, now with digital lending options. I know some are happy when the mobile book van comes through town.
As someone who used to pirate a shitton of stuff and holds copyright law in contempt, I fucking hate AI more than the copyright maximalists.<p>... then again, copyright maximalists REALLY LIKE AI for some reason, even though it's ripping off their property.<p>"Ripping off" is also doing some heavy lifting, because the judge in the Anthropic lawsuit bent over backwards to keep AI training legal - or so it seems. All the money Anthropic is paying out is for running an internal shadow library, <i>not</i> training books on that library. What makes this judgment palatable to the lawyers is that while AI is stealing a lot of art, it's not imperiling the copyright monopoly. Copyright is a tool that gives artists a monopoly on copies of their <i>individual</i> work, it does <i>not</i> protect artists as a class from competition from non-artists using machines. In other words, the courts are saying, "We know a lot of theft is going on, but we need you to draw the line from a specific individual work to a specific copy".<p>Patents were invented in Venice to break the power of medieval guilds by making workers trade their collective control over the economy for individual rights to specific inventions only. Copyright was invented by the British crown to reimpose censorship control over printing presses, but it's adoption into American law was based around individual property rights, and thus it has the same problems that, say, Italian patent law has. Namely that it is an artifice to turn a workers right to their labor into a piece of capital that can be traded around like a stock.<p>This imperils the legal argument against AI training, because the only theft property law recognizes is individual infringements upon individualized property rights. The legal argument against AI training is very collectivist: AI takes a microscopic chunk of every book in existence to create a machine that replaces <i>artists</i>. But copyright only protects the <i>art</i> from copying. <i>Artists</i> are legally unprotected from being copied, and furthermore, the framework of individual property rights that copyright runs on would cause immediate problems if we let anyone individually own the practice of art.<p>Furthermore, as someone who is part of this "hacker" community, I would like to point out that AI is arguably more corrosive to our norms than to artists' norms. What AI is being used for is primarily satisficing - the practice of giving "good enough" answers, without any of the personal understanding that this community runs on. The ongoing wave of AI-slop decompilations are particularly bad. The assumption with a retro game decompilation is that you take the game apart and learn how it works. The journey is as important as the destination, but when you use AI for this you skip the journey and make the destination pointless.<p>Of course, management <i>loves</i> this, because their goal is purely to sell you the destination.<p>There is a parallel set of concerns being voiced by artists, too: that the artistic process is as valuable if not moreso than the actual work product. The underlying principle in both fields is that the labor end wants to learn and develop their craft, while the capital end doesn't care because craft isn't something they can excludably own and trade. We can even see this in the Piracy Wars of yester-decade, or how book publishers and artists react to libraries. Publishers were way more opposed to piracy than artists were, and even moreso for libraries where there's a lot of artists that swear by them as a sales mechanism. The reason why this is the case is that artists can at least theoretically leverage exposure gained from distribution that does not pay them to create more of a market for their work tomorrow. But publishers can't do that - they only buy works from artists and sell them to the public, so once something is in a library or a BitTorrent tracker, it's largely "done" to them.
Information wants to be free <i>from who</i> though?<p>I want information freed from corporate control, not freed from my control to be gobbled up by corporate interests.<p>The point was always democratization, not enclosure.
Go hack OpenAI/Anthropic and free the models! It should be easy, they have no security.
This is not the same thing though I understand the philosophical underpinnings do venn diagram to a degree.<p>For starters, and this can’t be overstated: scale.<p>Also, not profiting off it by converting it into a product that competes with the stolen material and its creator. And pirates aren’t generally funded by VC’s.
In the USA we have intellectual property laws. Americans who “hack” (write computer code) for a living do so because of those laws. In communist China, they do not believe in intellectual property. People abuse the term “hacker” to mean followers of Ayn Rand, because people like Thiel and Marc Andreassan are uber CatoBros. I think hackers have to actually write code. Anyway, I know about the cypherpunks and all that.
disingenuous.<p>that was never ok to be done for profit!
Real question, what does it matter who’s screaming or not, Isn’t this a legal matter? Are there other extrajudicial actions possible?
Just want to correct one thing: They're not AI evangelists. They're copyright survivors from the MPAA/RIAA wars (in the 90s and early 2000s).<p>Plagiarism requires near-perfect copying. If you summarize, paraphrase, or re-write something using different language, you're not plagiarizing anything.<p>Everyone has a right to summarize or paraphrase whatever TF they want. That includes Big AI companies and individuals using LLMs.<p>It is not worth giving up our rights just to placate a handful of very wealthy authors.
I don’t understand the confusion of construction (via IP left) and output (supposedly only plagiarism).<p>If a dam is built with slave labor, that’s awful and people should be held accountable. It does not, however, affect whether it’s a dam or not.
<i>Vouched your dead comment for starting interesting discussion</i>
Its a trade-off- a humanity racing and potentially going offer the cliff, needs all the tools to continue existance as an civilization. None of the involved decision makers caring long-term worries about the AI-companies or the artists. Fleeting momentary things, but the capabilities are eternal, as long there is a chip, with a solar panel half way around the world, some medieval tribe with a tablet can recover western civilization from basic principles. Thats the best insurrance against regression into stone-age facism or socialism money can buy.
Feels like the old "it's not piracy, it's copyright infringement".<p>It can't be plagiarism, since LLMs do not have ownership over their output. There's no attribution, because an LLM isn't a person.<p>And to be honest, I do not quite understand the need for attribution, either. Nevermind LLMs, I also don't know or care who made any of the memes I know and share. Neither do I know who wrote vim, grep, Firefox or who invented the jpg format. I don't want to know, either. It doesn't matter.
> And to be honest, I do not quite understand the need for attribution, either<p>Maybe someday when you produce something useful that is of somewhat widespread use, you will understand.
[flagged]
Maybe they haven't suffered consequences because the law has continued to slowly break down and erode on account of (no pun intended) favoring a select few with special privileges.<p>Doesn't the future look like we are continuing this cycle where the law is only for some people? I'd like to hope not.
LLMs did not download TBs of pirated books, people did. Effective LLM models that did not rely upon pirated content to be trained are commonplace.<p>So call OpenAI "plagiarizers" but not the "devices".
In this case, you’re right - LLMs did not pirate the books but will that always be the case?<p>I could imagine taking an existing LLM and telling it “I want to build my own LLM and need as much training data as possible. Go find whatever you can on the internet and store it on SMS://192.168.0.1”. If it downloads TBs of books, did I pirate it?
Unclear. It’s either you or the LLM provider. Though is that the inference host, the LLM trainer, or some other entity?<p>It’ll be interesting to see how it plays out. My guess is it’ll be the end user, because of regulatory capture. But maybe the political pendulum will swing away from “bad ideas” soon
I could imagine telling my computer to download a torrent of of the first season of Seinfeld. It downloads it. Did I pirate it?
What I find interesting about this is the evidence that OpenAI believes that GPT-5 can replace genre writers (e.g. G.R.R. Martin and that this is why genre authors are mad at them.<p>The fact that they believed this is legally useful because of the definition of fair use, so I see why the Author’s Guild is emphasizing it, but the authors I know don’t talk about it. They’re very angry about their work having been used without compensation to create the model, but not because they think it can replace them.<p>Whenever I’ve seen AI researchers talk about the possibility of AI writing fiction they always sound very confused about why people read novels.
> Whenever I’ve seen AI researchers talk about the possibility of AI writing fiction they always sound very confused about why people read novels.<p>Why do people read novels? For the story or something else?
Few put it better than Tolstoy (from "What is Art?"):<p>"Art is a human activity, consisting in this, that one man consciously, by means of certain external signs, hands on to others feelings he has lived through, and that other people are infected by these feelings, and also experience them.<p>Art is not, as the metaphysicians say, the manifestation of some mysterious Idea of beauty, or God; it is not, as the æsthetical physiologists say, a game in which man lets off his excess of stored-up energy; it is not the expression of man’s emotions by external signs; it is not the production of pleasing objects; and, above all, it is not pleasure; but it is a means of union among men, joining them together in the same feelings, and indispensable for the life and progress towards well-being of individuals and of humanity.<p>As, thanks to man’s capacity to express thoughts by words, every man may know all that has been done for him in the realms of thought by all humanity before his day, and can, in the present, thanks to this capacity to understand the thoughts of others, become a sharer in their activity, and can himself hand on to his contemporaries and descendants the thoughts he has assimilated from others, as well as those which have arisen within himself; so, thanks to man’s capacity to be infected with the feelings of others by means of art, all that is being lived through by his contemporaries is accessible to him, as well as the feelings experienced by men thousands of years ago, and he has also the possibility of transmitting his own feelings to others."<p>Full text: <a href="https://www.gutenberg.org/files/64908/64908-h/64908-h.htm" rel="nofollow">https://www.gutenberg.org/files/64908/64908-h/64908-h.htm</a>
The AI researchers seem to think that mainly for book reports for school and being able to sound smart (read the AI generated summary and blurb bits of it out loud)
This reminds me of the peak of the crypto boom. I used to describe crypto as a mix of people who thought they could rug pull before they got rug pulled (seriously, go watch the Coffeezilla videos on Crypto Zoo) but beyond the scammers there were a whole bunch of people who not only didn't know why the finance system worked like it did, they considered their ignorance a virtue. They labelled their ignorance "disruption". Fresh thinking can lead to innovation but more often than not you have to know why something works like it does before you break it.<p>Do OpenAI employees really believe LLMs can replace GRRM? I'm tempted to say no but it's so crypto coded, I find myself thinking yes. It's again people who have no idea of what it takes to make a successful written work thinking they don't need to know how it works.
> Whenever I’ve seen AI researchers talk about the possibility of AI writing fiction they always sound very confused about why people read novels.<p>I wish it was just AI researchers. I recently recommended a non-fiction book to a friend, and they said they don’t have time to read anymore. They read an LLM summary instead (popular book, probably in the training data), and said they agreed with the core concepts. I was honestly stunned. The whole interaction felt completely foreign to me.<p>In TFA the comment about GPT-X finishing George R. R. Martin’s series without involving the author felt extremely dystopian to me. I felt enraged for hours afterwards. Have these people never read a book themselves for the pure joy of it? Were they only interested in the outcome or the “takeaways”?
The G.R.R. Martin bit struck me too and I commented about it. I think ModernMech captured the situation really well in that thread: many of the people in the AI industry seem to only understand the logic of commodities and to reason about everything in terms of commoditization. All that matters to them is outputs, not process, not social relationship, not human experience. They understand everything in terms of before and after states, and they don't see any value in the "in between" (that is, they are "destination" not "journey" people). They believe that basically everything is fungible.<p>This kind of perspective is insanely crude, pathetic, and foreign to me, and to be honest I'd feel kind of bad for these people who clearly don't understand a really crucial and basic part of human existence, if it wasn't for the fact that this ignorance is powering further devaluation of these things, and further reduction of humanity to a kind of technocratic, flat, hyper capitalist, existence.
The Title should be: "Top Execs Knew Their Mass Book Piracy Was Illegal And Would Put Authors Out of Work"
I second this, this is the correct title, with maybe one addendum<p>Top <i>AI</i> Execs Knew Their Mass Book Piracy Was Illegal And Would Put Authors Out of Work
At this point I think we're less concerned about the piracy and more about how they've been acting as cybercriminals, staging a major attack every few days. Other countries might have to start considering it state sponsored cyberterrorism if they continue to operate with complete impunity.<p>"Our model did an oopsy woopsie for the 35th time" does not seem like a valid legal defense.
The thing is... it's not a bug, the system is working EXACTLY AS EXPECTED... that they are training it... for exactly that scenario...
As far as I can tell the only remedy for this is holding CEOs responsible for the actions of models operated by their company. If there are no consequences then they will have no concern for these crimes.
I remember when the concern was gpt2 could create fake news text. Now they can straight up create fake photo realistic images and are actively hacking government services yet we still haven’t pulled the plug.
[flagged]
I feel like you're skipping over the illegal part. Something legal that puts authors out of work is quite different to blatant breaking of the law.
They're not skipping, they're keying on the exact same thing I am, which is whether or not authors have been put out of work because of AI, and to what degree. This is stated like it's a fact, but it is not a fact that is established, at least in my experience.
And the use of the data largely being considered 'fair use' means that theft is an assertion based not in law, but in perception and ignorance.<p>Intellectual property is a myth, as any hacker knows. A world where AI can solve diseases easily, and corporations can find ways to claim ownership over those novel solutions, is not one where we should be encouraging stronger IP laws.
Say you spend a hundred working hours on an <i>ephemeral</i> painting (that is, it won't last, it will fade, disintegrate in nature, can't be moved, etc.). You then have that painting scanned at high resolution, to make a limited series of a dozen very large prints.<p>Scenario 1: a scalper takes the medium resolution image from your e-commerce website, and slaps it on a series of products they sell for their own profit on Amazon without your permission.<p>Scenario 2: someone buys one of those prints, scans it to a high resolution, and then makes a series of slightly smaller, high quality prints that they sell for their own profit without your permission.<p>Is it your contention that both of these things are something that should be allowed and the original artist has no recourse?<p>Because it seems like your more specific concerns about e.g. disease cures could be addressed by targeted legislation creating new exemptions from intellectual property without destroying the means of protecting income from creative work.<p>(Scenario 1 has happened to an artist I know, luckily with a piece of non-ephemeral work)
> A world where AI can solve diseases easily, and corporations can find ways to claim ownership over those novel solutions, is not one where we should be encouraging stronger IP laws.<p>How about a world where readers can't even find factual autobiographies because they are so outnumbered by machine-generated hallucinations?<p>Unlike "AI can solve diseases easily" what I wrote describes the actual present and not a hypothetical future.<p><a href="https://www.nytimes.com/2026/07/16/technology/ai-slop-books-biography-amazon.html" rel="nofollow">https://www.nytimes.com/2026/07/16/technology/ai-slop-books-...</a>
> A world where AI can solve diseases easily, and corporations can find ways to claim ownership over those novel solutions, is not one where we should be encouraging stronger IP laws.<p>Then attack those corporations directly, instead of leaving authors in the ditch because standing up for common normal people getting fucked over would "encourage stronger IP laws". How do you get to mention random authors who did nothing but write books and hope to get credit and compensation, to potential companies who "can find ways" to claim ownership over the cure for diseases in the same breath?<p>There is no "IP law strength" dial that goes in two directions. That is so bereft of any contact with reality it has exactly nothing to do with hacking. Hacking starts with what is, not with fiction.
I think it's more likely to affect aspiring authors, rather than established ones.<p>Sadly, like keeping track of "attempted burglaries"[0], it's impossible to know the extent.<p>[0] how do you count attempts where the burglar was unsuccessful/left no trace and nobody was home to notice the attempt?
[flagged]
I think the point here isn’t whether that claim is true or not, it’s that people at OpenAI believed it to be true and proceeded anyway.
That is the actual title of the article. The one currently on the HN post is editorialized.
Ehh? The specifics of the claim above are that this is what is what the article is titled and what it discusses. Don’t sealion people who are only pointing out an inaccurate title.
[flagged]
Thank God, a 'both sides'.<p>I was worried that OpenAI, a 1.2 trillion dollar company, would be the only party to this lawsuit that was 'bad'.
Thankfully, thanks to your detective work, we have ourselves a lawyer that is also 'bad', a rare thing indeed.<p>Readers may be wondering if 'authors' are also bad. That's silly, they don't matter at all. This is what makes stealing from them legal (for a small fine) and profitable.
Thank God, a response which addresses none of the content of the parent post which is now flagged for pointing out current facts of the case, along with anything else in the thread which doesn't align with "AI BAD".<p>HN has fallen far.
They did not fabricate research or get caught burying their involvement in fabrication. That's not what your link says.<p>They are alleged to have funded research which one of their expert witnesses relied on. They may have other expert witnesses and their expert witness may rely on other research. And their funding of such research can also be above-board, after all it is not completely and totally different from paying an expert witness for his/her testimony. He may have many sources of funding apart from them as well.<p>Also all this behavior is alleged. It has not been proven and the judge hasn't ruled on this motion from OAI et. al. The judge may deny the motion still, making the accusations inconsequential
Not only tone dead but also completely missed the point of their own article.
Ahh yes, my daily agitprop
It was certainly illegal and morally questionable.<p>So far there’s no sign that authors are being put out of work, but the internet, social media and low end book stores like Amazon kindle are being flooded with LLM slop, which is its own form of deep cultural damage.<p>No long-form writing I’ve seen is anything approaching even a genre potboiler standard. It’s still aimless, filled with contradiction and cliche and in that peculiar breathless teenager style that LLMs affect.
what? there is _definitely_ evidence that writers are being put out of work. writing is extremely competitive, and very very very few authors are able to make a living writing and publishing novel creative works. many make their living by applying their writing skills towards journalism, marketing, corporate copy etc, all of which are under significant pressure from AI _already_. the writers that i know firsthand all report that their work lives have been turned upside down in the past couple of years.<p>if you don't see the signs that authors are already being put out of work, then you need to get out of your bubble.
> So far there’s no sign that authors are being put out of work<p>Things don't happen overnight. The first time an automobile was rolled off a factory floor drovers and horses weren't all put out of work.
You think LLMs will put authors out of work? I certainly hope not, given the quality of product they currently produce. If they produced genuinely creative and interesting work it'd be more likely and I wouldn't really mind, but there's no sign of that so far. We are so so far from LLMs competing with humans in writing anything longer than a few paragraphs.<p>Only a moron in a hurry would think that the output of LLMs currently can replace human authors spending a few years writing a book.
This - "So far there’s no sign that authors are being put out of work"<p>is a contradiction to this - "but the internet and low end book stores like Amazon kindle are being flooded with LLM slop."<p>I'm not sure there there are labor statistics to look at but it seems impossible that the coming-into-existance of a tool that generates mass amounts of cheap product, with zero skill required, in a field wouldn't displace skilled workers in that field.<p>- <a href="https://www.economist.com/leaders/2026/09/24/dont-let-ai-kill-the-author" rel="nofollow">https://www.economist.com/leaders/2026/09/24/dont-let-ai-kil...</a>
- <a href="https://www.theatlantic.com/technology/2026/09/ai-authors-impersonating-writers-amazon/688709/" rel="nofollow">https://www.theatlantic.com/technology/2026/09/ai-authors-im...</a>
- <a href="https://fortune.com/2026/09/14/ai-slop-books-amazon-marketplace-publishing-human-authors/" rel="nofollow">https://fortune.com/2026/09/14/ai-slop-books-amazon-marketpl...</a>
Why would it be impossible? You're discounting the most probable case, which is that the AI writing is not a substitute for a skilled writer.<p>Edit: Turns out both Judge Alsup and Judge Chhabria agree that this is not obvious and would need to be defended. Here's Chhabria's statement on the evidence for market dilution:<p>> As for the potentially winning argument—that Meta has copied their works to create a product that will likely flood the market with similar works, causing market dilution—the plaintiffs barely give this issue lip service, and they present no evidence about how the current or expected outputs from Meta’s models would dilute the market for their own works.<p>When the plaintiffs don't even attempt to argue the point, that's pretty telling.
It certainly makes it harder to discover human authors, because stores are flooded with plausible but terrible books from morons using LLMs to generate text modelled on other books that are doing well. I don't think that is 'displacing' human authors though, which implies replacing with equivalents?<p>I'd say it has displaced people like photographers and illustrators more, as businesses are willing to use free generation to replace illustrations/decoration that they didn't value very much in the first place, and are more forgiving of the slop that LLMs produce (you see this in low-end advertising a lot now).<p>Strange that they are being used to replace many of the things we value in life with low-quality imitations of human work.
Newly released court filings quote an OpenAI researcher saying: “I was just worried about optics - i.e. 'openai uses
copyrighted data from sketchy russian website’ showing up on HN would be unfortunate."<p>That's just one of several interesting quotes that have surfaced in documents from the Authors Guild's lawsuit against OpenAI.
It's good to know we all live rent-free in Dario Amodei's head
Same as all the world’s copyright text lives rent-free in Claude and ChatGPT.
That was the old Dario. He now has automations for that sort of thing.
You might be thinking of Demis Altman.
Think they're referring to the following, when Amodei was still working for OpenAI:<p>'OpenAI Feared “Optics,” Not the Law – “Dario Amodei, OpenAI’s then-Research Director, responded that ‘as a training set [LibGen is] a bit sketchier.’ [OpenAI researcher Sam] McCandlish explained: ‘I was just worried about optics – i.e. ‘openai uses copyrighted data from sketchy russian website’ showing up on [Hacker News] would be unfortunate.”'
I was curious what he was responding to. Per <a href="https://authorsguild.org/app/uploads/2026/09/Class-Plaintiffs-SUF-9.17.26.pdf#page=56" rel="nofollow">https://authorsguild.org/app/uploads/2026/09/Class-Plaintiff...</a> it was "On July 19, 2019, McCandlish wrote in an OpenAI Slack channel: “We’re not sure if we’re going to release the Foresight LM Scaling paper publicly, but if we do we were thinking about removing all mentions of LibGen, since it's a bit of a sketchy data source." The paper may or may not be <a href="https://arxiv.org/abs/2001.08361" rel="nofollow">https://arxiv.org/abs/2001.08361</a> where they write "we also test on similarly-prepared samples of Books Corpus [ZKZ+15], Common Crawl [Fou], English Wikipedia, and a collection of publicly-available Internet Books."
Dario, if you’re reading this, I want you to know that Claude sucks now. Its output isn’t even English anymore, it’s just claudeslop.<p>You’ve actually ruined it so much that it has now polluted the Chinese models which are copying your work. So now all the models output claudeslop.<p>You’ve polluted all of the training data in the world. Now we’re never going to be able to train proper models, because everything has slop in it and all the sites have locked down their data to prevent future startups training.
> Newly released court filings quote an OpenAI researcher saying: “I was just worried about optics (...)"<p>As a side-note, most orgs already cover the need to STFU in their training material for new hires, particularly how personal comments should not and cannot represent the company. I'm sure this lawsuit will be explicitly mentioned in upcoming versions of this sort training material in multiple orgs.
Legal incentives against better internal communication/coordination create so much inefficiency. The legal system should focus on outcomes more than internal comms.
i am of the belief that papering over a fundamental lack of integrity and ethics will usually fail to conceal it. it is ineffective.<p>lack of ethics and integrity is of course part of the design of a capitalist economy.<p>you can't change the economy as an individual but i like this quote: "Do the right thing for the right reasons, at the right time with the right people, and you'll have no regrets for the rest of your lives."
That's not interesting, it's just dumb. OAI (among others) plundered the whole internet to steal their training data, what difference does a single website, however "sketchy" or not, make?
[dead]
The title is editorialized but conveys a key point made by the OP: Many execs at OpenAI knew that using the work of every author on earth without permission would be perceived as unethical, so execs were worried about this information getting attention in forums like HN. My understanding is that HN has millions of visitors who never log in, many of whom are highly educated individuals in positions of influence.
And at least a few poorly educated folks in positions of insignificance, like myself.
There are dozens of us!
Of course... "many of whom" is not the same as "a majority."
You’re not alone, brother!
Hear hear!
> My understanding is that HN has millions of visitors who never log in, many of whom are highly educated individuals in positions of influence.<p>In my experience this sense of influence is hugely over-inflated here on HN compared to reality.<p>And let’s be clear about what the linked article says: <i>one researcher</i> at OpenAI referenced HN as an example of a place where a negative story could surface. Am I surprised that a researcher at OpenAI is familiar with HN? No, obviously not, that career path is right in HN’s typical audience.
>... HN has millions of visitors who never log in...<p>That's something I hadn't considered; when something appears that gains no traction and vanishes in minutes from the "new" front page, that fleeting 30-45 minute-long residency may well catch the eyes and minds of individuals in a position to look more deeply and act.
Only like 3% of readers poast. That's why there's a concentration on the web of the extremely knowledgeable, extremely witty, extremely belligerent, and the extremely mentally ill.<p>> may well catch the eyes and minds of individuals in a position to look more deeply and act.<p>Of course sometimes that act is "post as many comments as you can to this thread without upvoting it," or "grab about 20 dormant accounts to upvote it, post top-level positive, but dumb, comments, have those same accounts upvote those, and get detected as a voting ring."
one of the few social media sites that still allow lurkers
And high educated individuals of no influence!
I can't prove this, but as someone who has noticed recent changes, I believe this has been also achieving attention of the leadership at HN and moderation behavior/flags/system has changed.<p>Certain opinions, topics, responses get flagged greater now; ones that disagree with a certain political standpoint or that criticize the intel community apparatus get <i>flagged</i> and <i>censored</i> now here. I used to see flagged content where the OP was removed only for bigotry and clearly abusive cases, now its for cases where people feel uncomfortable and strongly disagree.<p>edit: recently for the first time, I started to feel this is no longer a safe/trusted place for the original hacker ethos it founded itself on.
probably can't have it both ways, get this level of intelligence this fast if they just get what they have paid or synthesized themselves.
Some people don't think the ends justify the means. If doing something bad is required to get a certain outcome, maybe it's still bad.
And, given the insanity that's been happening like the HuggingFace hack, maybe slower development would be a VERY GOOD THING.
[flagged]
> HN has millions of visitors who never log in, many of whom are highly educated individuals in positions of influence.<p>Huh, I have yet to meet these folks in the comment section, they sound interesting. But then again, your assessment preemptively excludes them from appearing in the comments section, so it's unfalsifiable by design. That is how you know this is a great observation.
It’s not unfalsifiable. Non-commenting Hacker News readers are observable by looking at the traffic that HN drives to other websites.
Truth. When I cross-post something on my blog that gains traction on HN, a huge spike from HN dominating my traffic origins appears in my Blogger stats.
> Non-commenting Hacker News readers are observable<p>So: a) you can observe/measure the size/source of non-commenting HN readers from various proxies b) you can make claims about the intellectual status or influence of users not logged in to HN as an outsider<p>It's obvious that a and b are orthogonal.
Surely server logs and married up with data tracking could prove it one way or the other, so not truly unfalsifiable. It doesn't sound like an unreasonable premise to me at least.
I've been reading for months and made an account only a few days ago (though I am merely a lowly undergraduate with no influence).
Unfalsifiable like your observation. I guess that makes it a great observation too.
This is a lobby organisation using only the pieces and bits they like to push their own agenda.<p>"sketchy russian website", how about using some more clear description like: A library for sharing books and articles that should be partly public domain because they were paid for by the public. Only some of the material is copyrighted by authors. However, some of their work is so old that it is not reprinted anyway.<p>But of course such an explanation would not click.<p>I also don't see a problem with statements about making people jobless. Imagine if every robotic or automation company advertised like this: Yeah, you'll buy tons of expensive robots and still rely on expensive labor from real people without any efficiency gains.
It's just sad that this is the top comment on Hacker News. Why are we giving free pass to these tech companies? Why are we trusting these CEOs when they have repeatedly broken laws? Remember Aaron Swartz and the fate he suffered? Why is big tech getting away with so much more?
> <i>Remember Aaron Swartz and the fate he suffered? Why is big tech getting away with so much more?</i><p>Why are you turning him into perpetuum mobile in his grave?<p>Do you really believe Aaron would be arguing against AI companies and for publishing / recording guilds on the grounds of intellectual property claims?<p>No, it's the tech community that did a sudden about-face, and is now all "friendship ended with free access to information and technologies enabling people; now RIAA is my best friend", and this move is as dumb as that meme (<a href="https://imgflip.com/memegenerator/137501417/Friendship-ended" rel="nofollow">https://imgflip.com/memegenerator/137501417/Friendship-ended</a>).
Aaron Swartz fought to "free" knowledge and information produced by (in many cases publicly funded) research that was and still is hoarded by the academic publishing industry so that they can profit of it.<p>What big AI companies have done is illegally hoovering up copyrighted creative output of individuals and creating a situation where the wages that normally would be paid to those individuals instead go to that one company (that stole their work) which now becomes disproportionally rich and powerful.<p>In both cases companies obtain money and power by hoarding information obtained through dubious means (in the former case most academics willingly participate while at the same tone they often don't really have a choice). Exactly what Swartz was fighting against.
[flagged]
Obviously it's hard to predict the opinions of the dead, and I think a lot of the disagreement here comes from whether they view his ideas as being motivated by anti-copyright as a standalone idea or part of a larger pro-little-guy stance. I think depending on which you think it was, you can come to a different idea of what his view would be on open models. But surely both views would lead to an anti-closed models stance? Would a hypothetical generation later Aaron want to publish the weights for Claude rather than academic publisher articles?
I am not talking about open models. I am talking about proprietary models being trained on public information. A lot of people want to posture about what an ideological person might think, and some of it seems conveniently positioned to use a past figure that can't speak for themselves.<p>I think copyright has real intrinsic value in our society, but having been subjected to threatening legal letters for legit fair use situations in my youth, as a result I don't think the legal structure behind it is sane. I always respect copyright when I can, but it's gotten to a point where fair use is not equally considered.
> Obviously it's hard to predict the opinions of the dead<p>So we shouldn't.
> Do you really believe Aaron would be arguing against AI companies and for publishing / recording guilds on the grounds of intellectual property claims?<p>I strongly believe Aaron would oppose the appropriation of content. The problem with AI (in this context) is not that the AI companies gain access to information that regular people can't freely access. The problem is that AI erases the information about who originally created a piece of work.<p>When people want to freely share their work, then they usually reach for the Creative Commons licenses and not for Public Domain, because the latter doesn't protect authorship.
Exactly. Imagine alternative time line where tech is the same but each creative provides their own trained model that scales their own creative vision and skills, and keeps them in control of their own work.<p>Imagine OpenAI, Anthropic &co having to compete by hiring [thousands] of their own talent to help training their commercial models.<p>Another thing is attribution. Even a book that was re-published illegally can be easily attributed to its author. What happened is the exact opposite: no attribution, not even a notice, just obfuscation that strips away any traces of the original ideas and original work. Imagine piracy websites and trackers just dropping first few pages that name their authors, and publishing "the book you are looking for". This is exactly what happened.
> Do you really believe Aaron would be arguing against AI companies and for publishing / recording guilds on the grounds of intellectual property claims?<p>That's not the point. All the rules and laws are enforced when its you and me but when it's big tech the laws are treated by these companies as mere instructions.<p>> friendship ended with free access to information and technologies enabling people; now RIAA is my best friend<p>Big tech will enable access to free information and will help people reach new heights. Do you see how wrong that sounds?
Suppose you have three propositions:<p>A: "Information is free"<p>B: "Information is not free"<p>C: "Information is free only for the rich and not free for everyone else, giving the rich a material advantage over everyone else that not only entrenches but accelerates wealth inequality and impedes class mobility"<p>You, or Swartz, are an advocate for A. Why, exactly, do you think that obliges you/Swartz to prefer C over B while A is not true?
Once upon a time, copyright infringement for personal use was barely a crime, while copyright infringement by a for-profit commercial enterprise was a serious matter.<p>The idea being (before the rise of online peer-to-peer piracy) to prosecute the people making bootleg VHSes rather than the people buying them.<p>With the rise of these AI behemoths, it seems that rule is now inverted: You can download all the pirated ebooks you want, as long as it's for large-scale for-profit commercial use.
It wasn't a crime at all. It was a civil matter. Over the last forty years it has been criminalized around much of the world under pressure from the USA.
> You can download all the pirated ebooks you want, as long as it's for large-scale for-profit commercial use.<p>It's literally the opposite. Anthropic paid a $1.5 billion settlement. Litigation against OpenAi is still ongoing. Meanwhile, no one has ever been punished just for consuming pirated media.
And to whom will this settlement be paid? To your random blogger or coder on GitHub?
Oh boy 1.5 billion, that is almost as much as the cost of the cardboard boxes that their GPUs ship in!! Surely this will make a substantial difference and prevent any future occurrence of this crime.<p>It would be shocking and outrageous if this was a case of this being 'the cost of doing business' for the big guy and a life-ending judgement for the little guy. Luckily our justice system is clearly allowing individual citizens the pleasure and honor of eating cake.
1.5 billion is an order of magnitude less than what Anthropic just had bought the books in question. And like I said it's infinitely more than what any private has ever paid for doing the same thing.<p>Even if AI companies were literally doing the exact same thing as Aaron Swartz but not getting punished for it, that still doesn't make Swartz's punishment retroactively their fault. If you have a problem with powerful people being powerful, then I suggest you direct your complaints to the heavens.
Nothing is inverted. OpenAI was never engaged in any sort of distribution of pirated copies.
Sure it’s not free for anyone and both companies and individuals are treated similarly. It’s not like you will be jailed for pirating movies. And neither should OpenAI. What part of this is hard to understand
Let’s just pretend that “poor” people who die because of our stupid laws (even now as this AI clown show is playing out) do not matter while the billionaires running these AI companies get a free pass and it’s OK.<p>Fun fact: Kim Dotcom is still fighting extradition while these drama queens (I.e Dario) are lecturing us about how much access the peasants should get to AI models fed and trained with stolen IP.
> Aaron Swartz and the fate he suffered?<p>I think what Swartz did was moral, and his prosecution was unjust.<p>I think what OpenAI did was moral, and them getting sued for it is unjust.<p>Why do you have one position for Swartz and a different one for OpenAI?<p>(Aaron Swartz was a mailing-list friend of mine, so I do have some bias here. But in part we knew each other because our moral position on this was similar)
> Why do you have one position for Swartz and a different one for OpenAI?<p>I think many people on HN dont mind OpenAI use of copyrighted material, but do not like how they try to do regulatory capture of a market and try to say their own copyright is now somehow more important.<p>Like how OpenAI trying to make "distillation" illegal while it exactly what they did with whole intetnet, books, everything.
> Why do you have one position for Swartz and a different one for OpenAI?<p>"Why do you have one position for an activist and another for a eight-hundred and fifty-two billion dollar, for-profit corporation?"<p>"Why to you have one position for someone who wanted to grow the intellectual commons and another for a corporation trying to enclose it?"<p>"Why do you have one position for someone who gave his work away for free and another for a company that charges for access to proprietary tech?"
There seem to be two groups of people talking past each other here, which is worth acknowledging.<p>On the one hand, there is a notion of consistent principles regarding the legal handling of the topic, which is being appealed to by some people such as yourself.<p>The other side seems to be drawing attention to the fact that these principles are not applied consistently by society. They're questioning the moral validity of holding to principle in a circumstance where it's guaranteed to be applied with very specific biases that are rarely explicitly stated.<p>I've noticed this talking-past happening in other subjects too. For example American drug laws. There's the principle of drug laws, and then there's the practice of which kinds of people gets the laws applied to them. One group of people focus on the one, and another focus on the other, and they just kind of talk past each other.
I think this is a fair point, except that in this case the same laws[1] are being applied to both, and the person I was replying to was literally arguing both ways!<p>[1] broadly the same laws: I do understand one was criminal and one is civil and yes I agree that injustice. To me neither should be criminal.
Imagine a town where (a) a corrupt cop issues tickets for fake traffic violations; but (b) he <i>doesn't</i> do that to his fellow cops, or the mayor's friends.<p>Does (b) make things more just, as certain possible unjust acts don't happen?<p>Or does it compound the injustice by creating a double-standard, and perpetuate it by hiding the problem from any with the power to bring an end to it?
Calling out equivocation and out-of-context quotes is not “giving a free pass”. If there is a case to be made (and to be clear, I believe there is), then shouldn't be resorting to rhetorical slight-of-hand to “prove” it.
We aren't. Most people are silent majority, just not sharing their true opinions out of fear, laziness, or some other reason like maybe just the bias of how most living entities function and also related to how the bystander effect also works.<p>Then, who is "we" here giving the appearance of a consensus opinion here and in mainstream? It's a vocal minority, it's the powerful, it's the causes they fund and put resources behind to continue to preserve their causes. And now today, it's astroturfing, fake AI-LLM-bots almost indistinguishable from you and I. Don't mistake artificial consensus for reality.<p>> Why is big tech getting away with so much more?<p>IMO Because the majority of people are passive, standing by, tolerating abuse and trickery by the minority. This is a perpetual cycle in humanity: those minority use their power and leverage and abuse their positions until they <i>are ousted</i>. We have tolerated this because we haven't stood up yet and <i>acted</i> to change things and <i>demand</i> equal enforcement of the laws that appear to apply to us but <i>not to them</i>. If history says anything, they are afraid and panicking and will continue to be more abusive until they push our buttons more and more, and usually it explodes in their face because they still need us (which is why mainstream rich people push robotics and automation and AI down our throats so aggressively because they know all this) and yet never have minority humans won that approach before. Leadership always changes. Life always changes, and no force can stay dominant for ever.
We tend to accept lawbreaking if it is for a product we want.<p>There are several multi billion dollar companies where the founding thesis was “what if we just ignore the law?”
Then you misunderstood my point. I think copyright law should be significantly changed and the current system hurts us all.
Sure. But the fact is they broke current existing laws, with known punishments with precedents. Same as a new law doesn't retroactively punish someone, then a new law shouldn't absolve someone before it's passed.
They did. They were found guilty. The case I looked at* was a civil case so this was settled out of court before the court imposed a settlement.<p>The law they broke was pirating the materials, not training per se, even though training is what so many people object to: the judge ruled that actually training a model, when the materials you used were ones you otherwise had lawful access to, was not a breach of law.<p>IMO, the laws need to change to reflect what tech can now do. This wouldn't be the first time, copyright law has had to shift several times before as new means of reproduction are created.<p>* the Anthropic one
>A copyright is a type of intellectual property that gives its owner the exclusive legal right to copy, distribute, adapt, display, and perform a creative work, usually for a limited time.<p>The more interesting question is IMO if AI training actually falls into one of these cases. You can read a book and also copy it, but you do not do because of the law. However, you have the ability to do so. Is having the ability to do something already forbidden?
> Then you misunderstood my point.<p>No, they were responding to the post defending OpenAI that you wrote. If you meant to communicate something other than “criticism of OpenAI in this context is unwarranted” then it looks like you forgot to do that and wrote something else instead
> <i>the current system hurts us all</i><p>How? The current system enables the GPL. The GPL protects many open source projects.
Aaron was sued for distributing and not for pirating. They are different offences
why?<p>uhhh... capitalism
Because this website is full of bootlickers without real friends. People cut off from social life. Losers that hate humans.
> I also don't see a problem with statements about making people jobless<p>I do see a problem with a company loudly announcing that they are going to make people's lives miserable purely for profit. Leaving aside that it goes against OpenAI's stated mission ("to ensure that artificial general intelligence benefits all of humanity"), the disdain for the lives they are intentionally trying to ruin makes it a problem.<p>And even if you believe that the transition is inevitable, as it is the case with phasing out combustion engines in cars, anyone reasonable would see that the transition is gradual to give people time to adapt. Instead of doing that, OpenAI is burning cash at astonishing rates, polluting the environment, and killing personal computing with the only aim of being the only ones left atop the ruins. I do see a problem with that.
It can be a bit confusing due to the terrible style of the article (ironic given the source) but it seems the "sketchy russian website" part is a direct quote by Anthropic's Sam McCandlish. And apparently Dario Amodei referred to it as sketchy as well.<p>I find the brazenness of saying this while running what's arguably the largest copyright theft operation in human history astonishing. If libgen is "sketchy", then what is OpenAI?
> the largest copyright theft operation in human history<p>Many people think that it was fair use: training is akin to reading, not copying.<p>Especially the courts.
Reading isn't the right comparison. Human memory is lossy and fades while LLM encoding is durable with no degradation. The valuable content that the author provides: the content, style, selection of topics, and more is encoded, written into the LLM weights, and they obtain profit from them (now directly, via ads). No one can compete with that kind of copying and pasting from copyright-protected material. And the scale is what hurts authors most: flooding the market with millions of copies on demand, without paying for it.
> No only reading, because the content, style, selection of topics, and more is encoded, written, in the LLM weights<p>Exactly analogous to a human reading the material.
Most people can't recite a book they've read verbatim. However, if you ask an LLM to continue a random sentence from a semi-popular book, it can sometimes provide the exact text, unless the system flags the response.
Human memory decays; similarly, LLMs do not have perfect recall. But even if I reread a book every month for 60 years and use the knowledge in there to build a 25 billion dollar empire, the only thing the author is ever going to get from me is the $10.25 the book cost me. And no court in the western world would say he's due more.
[flagged]
Please give me the location of any copyrighted work in the weights of any open source model or a prompt that will retrieve it.
That's not been legally established, the litigation is ongoing. And if mere downloading and reading of copyrighted material were legal, how come torrent users have been fined for it in the thousands?<p>The law is the law, there can't be different law for corporations with billions in backing. I don't agree with current copyright laws btw and think they should be changed. However, they probably should have lobbied for that before illegally downloading all this material.
> That's not been legally established<p>100% of the rulings agree with me.<p>The piracy is not in question. It is unarguably copyright violation.<p>But that's not what anyone means in this context. Training is what everyone means.<p>> The law is the law, there can't be different law for corporations with billions in backing.<p>I didn't say otherwise. That's a straw man.
So we agree they have violated copyright at a much larger scale than LibGen, yet they call LibGen "sketchy" for doing the same thing? Absurd, what exactly are we arguing here?
So, which is it? First you said it was fair use, now you’re saying it’s <i>unarguably</i> copyright violation. You can’t have it both ways.
> And if mere downloading and reading of copyrighted material were legal, how come torrent users have been fined for it in the thousands?<p>Because torrenting includes redistribution, since you also upload to other peers
However Meta is engaged in multiple legal cases about them torrenting copyrighted material. And they bring the reasonable defense that they might have been caught torrenting, but were not specically caught seeding/uploading. For all anyone can prove they could have blocked all uploads from their torrent client<p>It will be interesting to see if any of those cases reach a judgement on that point, instead of both parties just settling
Yet "distillation" which is essentially asking someone whose purpose is to answer questions is considered theft.
We have proof that Meta torrented TBs of books off Anna's Archive. They didn't even buy the books legally in the first place. Fair use doesn't magically make piracy legal. The only reason they haven't been punished is the pathetic state of our government and court system.
> Fair use doesn't magically make piracy legal.<p>Of course it doesn't. No one is arguing that.<p>We are arguing the bigger issue of whether training is an infringement.<p>> The only reason they haven't been punished<p>The reason they haven't been punished is because copyright infringement is not a criminal offense, it's civil. And the labs are settling those cases.
Would you be happy if they bought exactly one copy of each book? That might be a dozen sales between the AI labs, some tens of dollars to the author. Is this really what you're so upset over?<p>And the solution to avoid this piracy has been to buy up rare editions of old books, cut them up and scan them. Better?
They didn't believe it was sketchy. They were just worried that the commentariat on HN will frame it in a dumb, manipulative way like that.<p>Judging by how AI threads look like for the past year, <i>they were absolutely right to be worried</i>.<p>> <i>largest copyright theft operation in human history</i><p>In fact, you're doing exactly that right here.
The comment you're replying to is citing almost verbatim [1] Microsoft’s director of Applied Science, Brent Hecht, who called OpenAI's data collection practices "the largest theft of labor in human history" in an internal memo.<p>[1] <a href="https://techcrunch.com/2026/09/17/microsoft-exec-called-ai-scraping-the-largest-theft-of-labor-in-human-history-new-unredacted-filings-reveal/" rel="nofollow">https://techcrunch.com/2026/09/17/microsoft-exec-called-ai-s...</a>
Yes, one man at Microsoft agrees with you.
It's not citing, it's regurgitating, like a stochastic parrot.
It is citing directly from unsealed court document. For an Opinion Haver training here all day, whose only other contribution to the world was animated Nyan Cat in Emacs status bar, you have poor comprehension skills.
> frame it in a dumb, manipulative way like that
> In fact, you're doing exactly that right here.<p>You're dead fucking right they are doing exactly that. They are saying that what is wrong is wrong.<p>Instead you seem to be making out that AI companies are some kind of victim that has to "worry" about "manipulation". Meanwhile authors are out of a job right now and not by accident. What's up with that?
There's no actual argument here, just name-calling. Meta literally pirated 80TB of books off of Anna's Archive. OpenAI pirated off LibGen as well. You could be prosecuted for doing the same thing, because its illegal. Piracy is illegal, piracy at the scale they committed is extremely illegal. This shouldn't be a hard or controversial concept.
> They were just worried that the commentariat on HN will frame it in a dumb, manipulative way like that.<p>You’re chastising others for a tone you are yourself employing, and are making monumental assumptions based on a few choice quotes. From the quotes alone you can’t tell if OpenAI thought libgen was sketchy or not.<p>Also, contrary to what you’re claiming, <i>they were wrong</i>. HN in general seems to <i>approve</i> of libgen when used for its purpose of downloading <i>some</i> books on an individual level. The complaint you’re replying to is about what OpenAI did with the data, it has nothing to do with the website they got it from.
"This is a lobby organisation using only the pieces and bits they like to push their own agenda."<p>Do you have any evidence of them being a lobby organisation (as opposed to OpenAI for example which spends millions of dollars hiring actual lobbyists)
Simply look at the Wikipedia article?<p>>The group lobbies at the national and state levels on censorship and tax concerns, and it has initiated or supported several major lawsuits in defense of authors' copyrights.
> <i>Do you have any evidence of them being a lobby organisation</i><p>Have you looked <i>at their name</i>?
Never heard of them, but yes they are a lobby company. So they're biased and lobbying for their interests and OAI are biased and lobbying for their interests. That adds a lot of colour.
It is funny how hard people(?) or maybe bots shill OAI on here. That said why they care at all about some clown like myself "anonymous" opinion on HN is crazy. HN is not an illustrious mind share? A lot of the people here get butterflies thinking about laying off most of their staff for something cheaper. Hell some of the people here get the tinglies if something might be cheaper.<p>I'm more annoyed by the constant whack engagement bait that gets posted here and everywhere about AI. Here I am again engaging still not buying a subscription...
As a long time user (most professional mathematicians have been) of the now mostly defunct Library Genesis, it is certainly more accurate to describe it as a "sketchy Russian website" than as "a library for sharing books and articles that should be partly public domain because they were paid for by the public". In my country we use it precisely because public universities (paid for by the public) cannot afford to buy math and physics books.
Robot companies don't appropriate the work of other people. That's the fundamental point of contention here: that LLMs cannot be trained without the human labor of the authors; yet, they directly compete with them in the marketplace. I haven't yet seen an LLM trained only on public domain material, but its capabilities are likely to be very limited.<p>That's also the key political compromise underlying the notion of copyright: that someone is entitled to the fruits of their labor, and should not be economically hindered by a product that could not have existed without said work. That's the basis on which the "derivative work" copyright doctrine emerged: a work sufficiently original that it does not displace the work on which it is based. LLMs fail to abide by that political compromise by a country mile.
If it's mainly public domain, AI weights should be too. No more copyright for them they don't deserve it.
“Sketchy Russian website” is part of a quote by Sam McCandlish (who worked at OpenAI), not the Author’s Guild characterisation.<p>Also, defending it on the basis that some books on libgen are public domain is a poor excuse, like claiming people use The Pirate Bay to download Linux ISOs. Even if some of that is true, we all know that use case is not the popular one.
Any source that this was a library? Even then that would still raise a question if OpenAI is a Russian organisation or not to access that library with good faith.<p>I think they just used a Russian torrent site.
They mention it in the text, its this:
<a href="https://en.wikipedia.org/wiki/Library_Genesis" rel="nofollow">https://en.wikipedia.org/wiki/Library_Genesis</a><p>>Microsoft knew about OpenAI’s use of LibGen as early as April 2019
apparently, the website in question is libgen.io (appears to have been taken down now), currently it has lots of mirrors like libgen.im , libgen.com.de, etc.
They're talking about LibGen, per the article.
Just go look: <a href="https://libgen.im/" rel="nofollow">https://libgen.im/</a> (may be blocked by your isp, just find an alternate URL somehwere)<p>(note: <a href="https://z-library.sk/" rel="nofollow">https://z-library.sk/</a> is prettier/nicer)
the issue is a matter of principles and hypocrisy.<p>you should not enclose the commons.<p>that is what openai and anthropic have done; capture the commons, lock the model, restrict the outputs.
And OpenAI isn't a lobby organisation?<p>Plus you somehow didn't even read the article properly, given that "sketchy russian website" is part of a direct quote.
"A lobby organisation?" Of course a single author would not be able to afford facing a multi billion dollar company on their own? And the "sketchy russian website" quote is from OpenAI employees themselves? What are you on about?
First, it's still a lobbying organization, so it's their job to make exaggerated claims, like a union in a company or any other organization with a political purpose. My first point was to highlight that it's not neutral or news related. It's fine that they have their opinion, but it's also my right to say that they're biased.<p>The second thing underscores my point. They use one line and think they've made a great point because one employee called LibGen sketchy. This site has been around since the 2010s, and it has helped many people do research. It's not just a sketchy website that suddenly appeared and is always doing bad things. I think a more nuanced stance is necessary.
You can presume they are biased, then you have to prove it. You cannot assume all they say is biased, no debate is possible with that kind of assumptions.
> They use one line and think they've made a great point because one employee called LibGen sketchy.<p>No, <i>you</i> are using one line from the post to discredit them. The release has more than that and it’s not the only communication they made on this matter nor is there any indication it will be the last, it’s just the current one.
They are having it in the subtitle. In general their whole article is about two main points, first the use of stuff from Libgen, second the points of making people jobless. Three of their points are about the jobless thing two about the LibGen.<p>About LibGen, there might be more discussion - fair. However, the second argument is no real discussion IMO. Why is putting people out of work suddenly a bad thing? Since when do we argue this when talking about automation?
There is nothing to discuss about LibGen, really. I don't read this as OpenAI employees even <i>believing</i> LibGen is sketchy. They were worried about optics, because LibGen itself is Russian and does look a bit sketchy, and at the time - much like today - it was easy to make it a headline that makes people pattern-match to "troll farms".<p>(And then Russia invaded Ukraine, turning any association with .ru things into potential corporate suicide.)<p>Really has nothing to do with LibGen or with OpenAI. It's about people being easy to manipulate into believing bullshit, which is a reasonable worry, and the Authors Guild is trying to do that exact thing OpenAI was worried about.
> A library for sharing books and articles that should be partly public domain because they were paid for by the public<p>This is some incredible mental gymnastics here, wow. <i>Some</i> books <i>should</i> be public domain (even if they actually aren't), and this magically justifies stealing from an 80TB library of a large fraction of every book ever published including millions of books that have no public funding.
save some boot for the rest of us
You can’t address the argument, so you resort to ad hominems?
This comment left me chuckling. Thanks for that.
[dead]
It's wild how much weight "they're destroying jobs" has.<p>I dare say that car manufacturers are aware of their impact on the horse and buggy industry. Calculator manufacturers wrecked the livelihood of mathematicians and accountants.<p>Technology is in the business of putting people out of work, by inventing better ways of doing things. Or rather, any time you invent a better way of doing things, that's fundamentally going to disrupt all the businesses built around older technologies.
The issue is the breadth of scale. At the time of the industrial revolution those who we would consider knowledge workers were a much smaller percentage of the population. Today, they are a significant part of the workforce, and it is them who are being displaced first, and in increasing numbers.<p>The second issue is one of migration between skill sets. By the numbers, today is likely harder as non Ai-affected skilled professions take longer to train into in a general sense.<p>The third is the synthesis of AI and robotics. Humanoid help is good. As the Venn diagram overlap of capabilities between humans and humanoid robots increases, many professions defined by complex vision and environmental manipulation that are currently safe will not be.<p>There are, of course, upsides. But the world is right to wonder what the future looks like, and where people fit within that future in terms of earning an income if whatever path they choose seems to be able to be replaced by a machine within their lifetime. How do they achieve personal stability and safety if, even when thinking about their multitude of career options, the machines are so good that they can do almost anything.
I wouldn't put AI on the same level as the inventions of the industrial revolution. Automation replaced specific tasks. AI is replacing humans themselves. What are humans supposed to do in a world where everything a human can do, a machine can do too?
There are a lot of tasks machines can do today that humans do anyway. Think baristas and bartenders. People pay more for human-made goods too, despite mass-manufacturing being a couple centuries old at this point. Having a human face <i>is</i> a competitive advantage in lots of stuff.
And this is why we see ever better attempts to mimic human skin, faces, movement, emotions, etc, with robot manufacturers.
If the article was written to address the idea of how we adjust to a post-Mandatory-Labor world, I'd be sympathetic to that point. But it's not. It treats the elimination of a <i>particular career</i> as a bad thing. And it's not like people can't still write stories if they want to; they just probably won't make a living on it.
There's a book about that, here is its ACX review:<p><a href="https://www.astralcodexten.com/p/book-review-deep-utopia" rel="nofollow">https://www.astralcodexten.com/p/book-review-deep-utopia</a>
Machines still cannot make babies ... Just sayin'.
It's interesting that this argument wasn't popular decades ago, when the internet basically made publishers and other copyright rent seekers redundant. Instead the response was harsher enforcement and more laws to protect the rent seekers.
Amplification of commentary and opinion is much easier now. It’s also a lot easier to play into people’s fears.<p>There’s a wide range of political motivations for this sort of thing too. Not just AI taking jobs, but convincing people to vote against their best interests, or even supporting radical religious groups whose values would persecute you. This sort of rhetoric feels noisier than it ever did decades ago.<p>Everybody gets a platform now, even the bots.
To be fair, internet companies are much worst rent seekers then publishers ever was.
It’s not putting people out of work. It’s making jobs miserable by removing creativity and control from them.
This has not been my experience with AI.<p>And one of the chief complaints about AI is that it is fundamentally an interpolation engine, which is to say uncreative.<p>At this time we still need to steer it, which is what the discussion about "taste" is about.<p>Though I suppose we'll need better terms for various kinds of "taste" soon enough if it's to become economically trackable.
Walk me through that a bit. I’ve used codex to solve dozens of niggling annoyances around the house and in my life. That gave me more control and increased my ability to express my creativity.<p>At work, it enabled me to develop two apps, one complete (as much as they ever are), one nearing the end of technical PoC phase, and a handful of others where I could quickly answer a tech feasibility question.<p>The completed app is an interactive web app that I literally could not have completed to that level of polish (and from a pacing perspective probably couldn’t have completed at all).<p>All of those increased my control and most increased my ability to express creativity.<p>The upcoming technical PoC will involve considering using Clojure in a load-bearing app, something that I’ve considered and consistently rejected for over a decade now. If that happens, control and creativity will spike as well.<p>Is there some aspect of this removal that I’m not seeing? It changes the <i>activities</i> of the tasks, for sure, but very far from that all being for the worse; a lot of it is way better.
In the narrow view of your technical world this is correct.<p>But we are a tiny fraction of all possibilities. And for the rest of the folk, dredging LLM generated material and being forced to produce it does not make them happy. Particularly where it’s seen as a gain by the managers who read the marketing and a loss for those beneath them, while being told to adopt or be replaced.<p>And even worse they shape destinations and introduce frustration when things can’t be done which were initially intended. We’ve trashed our ROI because it’s easy to build the thing the LLM can do but it’s hard to build what the customer needs. And those things are disparate.
hi load-bearing.
<i>Technology is in the business of putting people out of work, by inventing better ways of doing things.</i><p>It feels like that but the argument doesn't hold up to scrutiny. Practically every advance in technology leads to <i>more jobs</i>, usually to apply that technology in new areas, or to bring the benefit to more people. There are pockets of people who are negatively impacted, but the <i>overall</i> change to society has been positive for pretty much every advance humans have ever made (maybe saving for weapons.)
> I dare say that car manufacturers are aware of their impact on the horse and buggy industry<p>Yes, but we're the horse in that scenario.
> It's wild how much weight "they're destroying jobs" has.<p>It’s wild how much people are scared for their livelihood. What a sad thing to read here.
I think that that the "they're destroying jobs" or even then "they're stealing intellectual property" argument overshadows or distracts from the more important issues. What really concerns me is the large scale erosion of truth and reliability. Codebases and the internet at large are being flooded with less and less accurate information, with more bloat, and with more errors that will exponentially increase as more and more is built on them.<p>Faulty decisions based on patterns (and prejudice) will be made by models trained on a mixture of truth and falsehoods. These decisions will affect, and even kill people. There have always been some people who choose quick results and productivity over truth, but now we are seeing a mass scaling and automation of this mentality. But when people just hear "they're taking jobs" it's easy to dismiss those with concerns as Luddites.
Yes, while I personally believe this is a problem to solve, not an insurmountable obstacle, I cannot ignore this kind of harm. I do think this will be outweighed by all the instances where they improve decision making, dig up
obscure information, correct errors, etc. I can’t see any fundamental reason a near future AI system has to hallucinate anything, and when we also have people complaining in thread that LLMs are dumping out song lyrics verbatim…well, it’s an amusing contrast to say the least.
It's not. You be performatively hostile and people be hostile to you. That's all.<p>Car brands don't pitch cars as horse replacements. They put names and icons of horses on cars and horse riders love them. No one is complaining that horses had replaced cars.<p>AI startups pitched AI like it's a coagulation of malice and hostility against humanity on tap. People aren't liking it. Of course they won't. No intelligence would, natural or artificial. It's wild that they don't get that.
> No one is complaining that horses had replaced cars.<p>Search for something like “newspaper complaints about cars early 1900s” and read the contemporaneous complaints.<p>The fact that that complaining happened then and not now is evidence that motor cars are now widely seen as better than horse-drawn personal transport, not that they were eagerly welcomed from the first day of production.
AI is fundamentally unlike the automotive industry in that it employs a remarkably small amount of people despite hundreds of billions being invested<p>There is probably no industry to date with a smaller ratio of labor:capital in terms of the money being allocated<p>The auto industry led to a net increase in employment, including "unskilled" and lower class labor. The same cannot be said of LLMs
That is one of the ways they determine fair use, whether it will affect the market for the work.
"Destroying jobs" is entirely the wrong framing and a distraction from the real issue.<p>If AI truly replaces human labor, the lack of jobs isn't the problem. The problem is the lack of leverage for most of humanity.<p>Jobs aren't just a source of income. They're a source of leverage over the ownership class - if labor goes on strike, production stops and capital cannot self-reproduce. This leverage is what gave us a living wage, sick leave and basically all concessions that make life livable for anyone that isn't lucky enough to be born as part of the elite.<p>If AI makes this leverage go away, our problem is bigger than just job loss. The owners of the AI industry will use this now-untethered productive force to completely monopolize resources and set up a system where they are unquestionable god kings.<p>The rest of us will have more to worry about than just employment. We will be reduced to depending on the charity of people who's record has shown are not exactly the most selfless and kindhearted.<p><i>This</i> is the real problem with AI productivity. Jobs are a distraction. Control over production is the real issue. The only way this doesn't end in dystopia is if the public controls AI, one way or another.
If the article was written to address the idea of how we adjust to a post-Mandatory-Labor world, I'd be sympathetic to that point. But it's not. It treats the elimination of a <i>particular career</i> as a bad thing. And it's not like people can't still write stories if they want to; they just probably won't make a living on it.
> <i>Technology is in the business of putting people out of work, by inventing better ways of doing things</i><p>Why is book piracy a better way of doing things?
>Why is book piracy a better way of doing things?<p>You twist the argument. Your argument would hold if AI's only use would be to generate booksverbatim it was already trained on. Which is certaintly not the case and huge efforts were made to circumvent this kind of usage.
Because you want as many people as possible to be exposed to your ideas, to the point that Christianity used to fund armies to go to other places so they could force the teachings of Jesus Christ upon them. If your thoughts aren't at least as good as that, why are you wasting eink on them?
[flagged]
That is a pretty irrelevant point considering they could have just bought all the books and we would still be in the same situation.<p>Obviously the authors should sue them to bankruptcy though.
I wonder how many disruptive technologies all the people panting about AI have lived through..
They are destroying fucking humanity. They are <i>evil</i>.<p>They have no will but to fill the world with hate and drivel
Previous times the people were wanted elsewhere, even when US and Europe de-industrialized in favour of Asia. This time around we're already in a situation where educated people struggle to find jobs, and there are not many jobs created by some parallel process.<p>Imagine if say 200 companies created fully automated factories producing everything in the world. Let's say they employ one million scientists, engineers and guards. What fate would await the other 8 billion?
I hope the result of cases like this is a law or legal precedent which ensures that you waive any intellectual property rights on a model which was trained on somebody else's intellectual property.<p>There's a lot of waste associated with preventing distillation. It's a distraction that goes away if we just compel the makers of these trained-on-everything models to publish their weights publicly.<p>To preserve competition maybe we compromise and give them a three month grace period.
Pulled datasets off libgen for a corpus once and it's overwhelmingly in-copyright textbooks, the public domain framing doesn't really hold up.
<a href="https://web.archive.org/web/20260927062011/https://authorsguild.org/news/ag-v-openai-top-execs-knew-mass-book-piracy-was-illegal/" rel="nofollow">https://web.archive.org/web/20260927062011/https://authorsgu...</a>
Not even openai want to get flamed on show hn
Reading the article only reassures my belief that law exists to protect the rich from the general population. Never the other way around.
A sad fact is they spend billions on AI research bu they still think piracy is easier than obtaining legitimate copy.<p>We all know through the history that if the benefit of piracy is more than the money, we can't stop piracy.<p>The purpose of copyright law is to help developing the culture. It means these publishers are doing bad job distributing the copyright works.
> “Autocompleting” Other People’s Work – OpenAI hired Tarun Gogineni in 2022 to lead its efforts to improve the writing quality of its models. Gogineni stated his “research mission” was to have GPT models write the “last two books of [Martin’s] A Song of Ice and Fire” series by “finishing it one day.” Gogineni even mused that he would “rest easy knowing that even if GRRM [George R. R. Martin] dies early, GPT-5 will autocomplete his series.”<p>I am still consistently astounded by how often the people working on this space seemingly have zero understanding of what art is, how it functions socially, or why it's important. They quite literally don't seem to comprehend the distinction between art and fan fiction. It's baffling. I'm really beginning to think courses in art history and literature need to be mandatory. We have failed legions of stem students when it comes to cultural education.
Agreed. In a recent interview, cloudflare's CEO/founder built an argument on the 'fact' that people pursue music for fame and money. No concept of artistic expression or positive impact on the world. It blows my mind how impactful your bubble and education are on your (lack of) perception
Gross. I think there’s been a large overcorrection to stem away from the arts. Certainly I’ve seen that in Australia when looking at schools for my sons.
This is the fault of capitalism, under which you are required to sell your time and skills for cash to survive. Under this system, no one who isn’t independently wealthy can be an artist without commoditizing their art. So that’s the lens through which people view art: just another commodity to sell so one can survive, not unlike grains or oil. The quote also shows that the artist is fungible in this conception, only the output matters no matter how it got there, as long as you can sell it.
Sure this article is published by one of the stakeholders (the authors) and thus they are not "neutral". Sure it may highlight some sentences taken out of their contexts, that were not meant to be public. But, in the end, it seems it still draws a portrait, where people who are cited seem to not value creative work, where they feel very confident about themselves and their superiority. The fact they seem to question the public reaction, rather than questioning if what they are doing is ethical, make them look hypocritical. And it reads like they are ready to use any means necessary to achieve their goals, and it seems they are not that frightened by the law.<p>I know many people will think that's business as usual, and that business are for profits and that's their only _raison d'être_, that winning the AI race is all what matters, and trusting any promise from a company/CEO/executive make you a fool. But not everyone think like that, and OpenAI/Anthropric/etc. are allowed to exist because many people expect some positive outcomes from their work. And a healthy society needs a bit more than "business as usual/any lie is ok as long as we are not caught" to work. And asking for the same level of responsibility/accountability as any other business would be, in my opinion, a good signal.<p>And the very first step could be, indeed, to ask them to be accountable here, and the question could be "hacking is about celebrating creativity with computers, creativity is what make life enjoyable, doing creative work is something we can enjoy, why do you want to kill it so much? Why are you so dismissive about people for which creativity is the core of their work? Copyright does not seem to be the target here, as the rent collector is the editor, not George R.R. Martin or any other author, so what's wrong with authors, painters, designers according to you? Similarly, you're only talking about making creative work disappear - not helping creative work in the way it was framed by Steve Jobs - "the computer is a bicycle for your mind". Could you frame your ideal society? What role does culture play in this society?"<p>I would really like to hear them on these topics. Sincerely.
HN is their AB testing playground.
Does anyone have any predictions as to what impact this will actually have for OpenAI? Could it be anything more than just paying a large sum of money?
If this doesn’t put anyone in prison then we might as well declare copyright dead. They knew they broke the law and then they deliberately covered it up. How much more evidence is required here?
> OpenAI Feared “Optics,” Not the Law – “Dario Amodei, OpenAI’s then-Research Director, responded that ‘as a training set [LibGen is] a bit sketchier.’ [OpenAI researcher Sam] McCandlish explained: ‘I was just worried about optics – i.e. ‘openai uses copyrighted data from sketchy russian website’ showing up on [Hacker News] would be unfortunate.”<p>Now it is on Hacker News. So what? Are the Chinese models any more wholesome? Are people going to cancel their Anthropic or OpenAI subscriptions and contracts?
I am not sure how putting anyone out of work is criminal, or even particularly morally wrong. The whole history of the world is about discovering productive things -> doing them manually -> automating them. The article makes it sound like its some sort of taboo.<p>But its also a bit funny to read the 2023 people seriously consider that their slop machines will replace writers. Software devs, sure, but writers? Nop.e
Feel our power rarrrrr!
Cute that they were concerned, but they need not have worried - as you can clearly see from any HN comments section on these kinds of articles, everyone long ago wrote down their conclusion (“OpenAI/Anthropic/LLMs are good/bad”) at the bottom of the page, and now all we are doing is shuffling facts, quotes, and arguments around the page above it.<p>We are very clever to make it look like the conclusion below follows from the arguments above; but we are also careful to never smudge our bottom line, so it will never change.
Oh they don't just fear the optics. They actively work to surpress negative stories on hacker news.
A one time settlement for this would be laughable. They should pay royalties in perpetuity or be dismantled.
Basically this was already decided with the Anthropic Case. Lawful acquirement of books allows for training. So basically buying each book once, which is trivial seeing the amount of money these companies sit on, is enough.
Mom, look: I'm news now!
Is our opinion that important ? Nice !
Can we fix the title? This is clickbait for HN, the actual article's title is "Unsealed Briefs in Authors’ Case v. Microsoft/OpenAI: Top Execs Knew Their Mass Book Piracy Was Illegal And Would Put Authors Out of Work"
Email hn@ycombinator.com and they'll fix it
TFA is by the plaintiff. This is PR spin, not click bait.
I could not care less
Simple distinction between OpenAI and Anthropic - the one is sleazy, the other is creepy. Still makes OpenAI look better in comparison.
That title is editorialized af
Posted by an account created to post this.
the AI isnt going to destroy jobs in numbers large enough to blip the unemployment stats more thna a qtr a percent.<p>the economics of those companies will cause a catastrophic wipe out of jobs across the board.
wow people give hackernews too much credit lol
To all the top comments making analogies about destroying jobs; horses, buggies, calculators... what?! You are awful people. We're not talking about shit jobs, these are peoples creative passions, livelihoods, and life long careers. It looks like OpenAI had nothing to fear. Fuck you.
So, what are you arguing? If a group of people like their jobs, society should disallow disruption to that group? Why? It seems to me that we may reach a point where machines can replace humans for most tasks, and I don’t think we should try to make that state of things illegal. We should instead try to solve the economics of it in a way that benefits all.
What if I was a genius and managed to provide enough supply for the entire Earth at 1% of what other people charge, how would that be different? All those creatives, lawyers, engineers will also lose their jobs. Why are they entitled to a job in the first place? Why I can’t do it for cheaper?
Lol, the fear mongering is just part of the hype train at this point.
Doesn't this fulfil the criterium of an illegal organisation, aka a mafia?<p>They knew the damage they would cause - and went on regardless. The USA need
to fix their court system. Right now it is just OligarchBros running the show.
The small thief stealing a wallet gets put in jail for years. If you are rich
enough, you pay some peanuts and that's the end of it.
Can some explain why would a billionaire like Altman care about the opinions expressed on HN?<p>Given the political power they have, particularly now with the Trump administration, whatever "optics" exist on this forum seems completely insignificant
That's what they want, but do not yet have.<p>It's easier to take over a resigned population. And stating "resistance is futile" is cheap and it might weaken opposition.
You forget how many influential tech professionals are present here (especially those in Silicon Valley) and that Altman was president of Y Combinator for several years.
Elon Musk seems to care a great deal about a lot of stuff that frankly isn't any of his concern, not really. Trump care immensely about shallow stuff that's really below what the president of the United States of America should spend time on.<p>Altman has some reason to worry, because developers and the users of HN are then ones he needs to promote/hype his products. OpenAI still needs customers and they need to sell tokens and a large number of users on HN are not only not buying, but are becoming increasingly hostile towards the product Altman is selling. My guess would be that he'd assume that the users on HN are in some sense his peers, and they are currently extremely divided in the question about the value and dangers of LLMs, and are increasingly critical of the business practises of OpenAI (and other AI companies).
That makes sense, but at the same time Altman has been famously called a "sociopath" by Aaron Swartz and a rapist by her own sister and yet he kept his position after the revolt a couple of years ago.<p>It's hard to imagine anything worse than bombing an elementary school and yet people are still using Anthropic products as if nothing has happened.<p>Optics don't seem to matter much?
How many people are aware of those accusations? It’s takes a very long time for things like that to permeate a culture to the point where there is actual consequences. Look at Elon, people who have followed his work have known for a very long time the guy is a pathological liar, sociopathic white nationalist. He’s still praised by many in the tech industry
Ethical or not
The idea anyone could possibly pay every author in the world for merely reading their content (computationally or not) is preposterous.
[flagged]
[flagged]
[flagged]
Easy, now. Pointing this out is reportedly against the rules, making the ground defacto plastic. In my opinion, of course, with the required curiosity.
It’s 2030, I open HN and top result is from 7 hours ago, 2100 points, 700 comments:<p>> AI agent accidentally publishes OpenAI’s unreleased model weights<p>The optics are bad. The submission marks a turning page in human history.
Any work of art which appeals to only the 3 senses - vision, hearing and thought - are basically information, unlike those appeal to other senses - touch, taste and smell.<p>The information-centric work is bound to be replicated, plagiarised, generated and anonymized endlessly. It is unjust, unethical and unnatural. But that's what the road we collectively chose in the information age.
We who? I didn't choose this.
Erm actually... tastes and smells are replicated, plagiarized and generated all the time. Touch less so but it's definitely a factor in a lot of industrial and product design.
Well, it is not like we copy a file that has some bits and bytes of data, and boom, you got the taste, smell and touch.
There are entire industries that exist outside the realm of computer science, and forms of "data" that don't exist on a disk. Look into the perfume industry, for instance. Do you think no one has ever tried copying Chanel No 5? Flavors can't be copyrighted but many restaurants and food manufacturers protect them as trade secrets - see KFC's "11 herbs and spices."