Humanity hasn't caught up to the reality of the times; we should adjust copyright laws so anything over say 10 years old can be hosted for free on the Internet. We should have colossal databases of music, books, movies, etc. that can be freely traded and downloaded without breaking any laws. Maybe in the past the commerical protection of art and IP was necessary but with the Internet and AI I think we need to, as a species, get over it and focus on information archival and dissemination over protecting income streams.
Related:
<a href="https://news.ycombinator.com/item?id=49068738">https://news.ycombinator.com/item?id=49068738</a><p><a href="https://news.ycombinator.com/item?id=49127284">https://news.ycombinator.com/item?id=49127284</a>
On the surface it seems bad, although many of these books were probably rotting in place rather than being read. If the companies are willing to make the digital version available this may actually prreserve the books.<p>More troubling to me is the focus on pre-2022 data. This indicates that the companies are seriously worried about model collapse, where the future models become worse becuase they are being fed by generated data instead of real human data.<p>This is actually my worse case scenario for AI: the models become good enough to disrupt entire industries in the next 10 years, but then stay frozen at that level. Model drift then kicks in because the world will keep changing but the models do not, and in a few decades we actually regress instead of progress becuase they won't be enough skilled humans to drive advances.
As long as economics incentivizes next quarter thinking nothing will be done. We can't even handle human portion of global warming which is arguably more important than just some AI model collapse scenario.
Yeah, I really don't know where this goes from here. Like certain sources, particular Reddit and Twitter, have to largely be considered spoiled at this point. It's hard to quanitfy the amount of bots but my intuition tells me it's really high.<p>But changing how non-AI people write, that's an interesting angle. Because where do we go from here? In 100 years will we still be overvaluing pre-AI sources? That doesn't make sense. Of course a lot can (and will) change in 100 years.<p>But i'm reminded of pre-atomic steel, which is steel made before the first atomic bombs were detonated and thus have really low background radiation. This is necessary for making MRIs and such. People will go and find it from shipwrecks and such (ironically, many of which are from WW2). It's also a finite resource. What happens when we run out?<p>Now pre-2022 texts aren't consumed (other than destructively scanning books of course) but it is also finite. We can't make more of it.
There is good discussion on the previous thread from 6 days ago: <a href="https://news.ycombinator.com/item?id=49068738">https://news.ycombinator.com/item?id=49068738</a>
Yes, since we complained about them violating copyright so much, it does not violate copyright if they destroy the original, so they are now destroying the original. We got our demands met.
Why don't they just put the pdf versions on Internet Archives? Then they would at least be preserved for the future
Forever preserving the content of a book by consuming one (1) copy of a rare book is the kind of thing we should be doing more of.<p>Roughly nobody would have had access to that copy. Roughly everyone can now benefit from its content.
One (1) company destroys and (illegally?) copies a rare book. This knowledge is now of that company, not of humanity. Other companies will feel the need to do the same and before you know it knowledge is walled off in the gardens of the AI companies, and the rare books are gone and destroyed.
> Roughly everyone can now benefit from its content.<p>You mean, roughly everyone who uses that particular model provider’s products. For truly rare books, this has an anticompetitive flavor, since it ensures others can’t train models from the same knowledge.
I have a deep rooted doubt that "roughly everyone can now benefit from its content" If there's any value in the content it will immediately be monetized with an aggressive pay structure and had it sat in a library, that same value would be free.
Where can I read any of these destroyed copies from anthropic?
It's going to be hard for you to do more of that, considering that AI forms are destroying the books. You know, the whole point of the article...
Roughly Anthropic owners benefit from the content, unless the scans are made available somehow.
> Roughly everyone can now benefit from its content.<p>You mean every Anthropic?
Who knew Ray Bradbury could be so wrong.
Yay capitalism! Imagine how much value is being created for shareholders by this destruction! Another example of "great work" being done!<p>Jokes aside, I don't know what I would do in this situation either. Maybe find a way to mark "rare" or "out of print" books and sell/auction the "pages"? Or make the scanned book available via some mechanism?
Nowhere are the supposed rare books listed.<p>Confuses "(not all that) old books" with "rare books".<p>Ragebait, nothing more.
They are academic titles from the last few decades for the most part, and are rare because the print runs were typically short, and the prices typically eye-watering. “Semantic analysis of cotton processing terminology in Des Moines, 1993-1996, 8th edition”, that sort of fun.