10 comments

  • quacked8 minutes ago
    Humanity hasn't caught up to the reality of the times; we should adjust copyright laws so anything over say 10 years old can be hosted for free on the Internet. We should have colossal databases of music, books, movies, etc. that can be freely traded and downloaded without breaking any laws. Maybe in the past the commerical protection of art and IP was necessary but with the Internet and AI I think we need to, as a species, get over it and focus on information archival and dissemination over protecting income streams.
  • themaninthedark0 minutes ago
    Related: <a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49068738">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49068738</a><p><a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49127284">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49127284</a>
  • cowanon7720 minutes ago
    On the surface it seems bad, although many of these books were probably rotting in place rather than being read. If the companies are willing to make the digital version available this may actually prreserve the books.<p>More troubling to me is the focus on pre-2022 data. This indicates that the companies are seriously worried about model collapse, where the future models become worse becuase they are being fed by generated data instead of real human data.<p>This is actually my worse case scenario for AI: the models become good enough to disrupt entire industries in the next 10 years, but then stay frozen at that level. Model drift then kicks in because the world will keep changing but the models do not, and in a few decades we actually regress instead of progress becuase they won&#x27;t be enough skilled humans to drive advances.
    • storus9 minutes ago
      As long as economics incentivizes next quarter thinking nothing will be done. We can&#x27;t even handle human portion of global warming which is arguably more important than just some AI model collapse scenario.
    • jmyeet13 minutes ago
      Yeah, I really don&#x27;t know where this goes from here. Like certain sources, particular Reddit and Twitter, have to largely be considered spoiled at this point. It&#x27;s hard to quanitfy the amount of bots but my intuition tells me it&#x27;s really high.<p>But changing how non-AI people write, that&#x27;s an interesting angle. Because where do we go from here? In 100 years will we still be overvaluing pre-AI sources? That doesn&#x27;t make sense. Of course a lot can (and will) change in 100 years.<p>But i&#x27;m reminded of pre-atomic steel, which is steel made before the first atomic bombs were detonated and thus have really low background radiation. This is necessary for making MRIs and such. People will go and find it from shipwrecks and such (ironically, many of which are from WW2). It&#x27;s also a finite resource. What happens when we run out?<p>Now pre-2022 texts aren&#x27;t consumed (other than destructively scanning books of course) but it is also finite. We can&#x27;t make more of it.
  • eleventen26 minutes ago
    There is good discussion on the previous thread from 6 days ago: <a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49068738">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49068738</a>
  • inigyou31 minutes ago
    Yes, since we complained about them violating copyright so much, it does not violate copyright if they destroy the original, so they are now destroying the original. We got our demands met.
    • lookingdesk28 minutes ago
      That&#x27;s not how copyright works, at all.
      • inigyou25 minutes ago
        Yes, actually it is
        • pfdietz10 minutes ago
          Destroying your copy of a copyrighted work does not violate copyright.
        • netsharc19 minutes ago
          I&#x27;m going to buy a Taylor Swift CD, rip it as MP3s, destroy the CD, and put the MP3s online...<p>Will I get sued?
          • pfdietz14 minutes ago
            The &quot;putting it online&quot; part is where you are f-ing up.
      • madaxe_again23 minutes ago
        It absolutely is. By destroying the original they retain the one copy they were licensed. It explicitly discusses that this was judged fair use in a court of law, because of the destruction of the original, as this becomes a transformation of the original work rather than a duplication.<p>Don’t hate the player, hate the game.
  • DrDeese14 minutes ago
    Why don&#x27;t they just put the pdf versions on Internet Archives? Then they would at least be preserved for the future
    • pfdietz12 minutes ago
      That would be illegal.
  • jstummbillig27 minutes ago
    Forever preserving the content of a book by consuming one (1) copy of a rare book is the kind of thing we should be doing more of.<p>Roughly nobody would have had access to that copy. Roughly everyone can now benefit from its content.
    • 2830428340923418 minutes ago
      One (1) company destroys and (illegally?) copies a rare book. This knowledge is now of that company, not of humanity. Other companies will feel the need to do the same and before you know it knowledge is walled off in the gardens of the AI companies, and the rare books are gone and destroyed.
      • pfdietz13 minutes ago
        The scan doesn&#x27;t go away, and can be released when the copyright expires.<p>Oh no, a lump of cellulose is gone. Will no one think of the fibers?
    • anon37383924 minutes ago
      &gt; Roughly everyone can now benefit from its content.<p>You mean, roughly everyone who uses that particular model provider’s products. For truly rare books, this has an anticompetitive flavor, since it ensures others can’t train models from the same knowledge.
    • polarbearballs21 minutes ago
      I have a deep rooted doubt that &quot;roughly everyone can now benefit from its content&quot; If there&#x27;s any value in the content it will immediately be monetized with an aggressive pay structure and had it sat in a library, that same value would be free.
    • maccard26 minutes ago
      Where can I read any of these destroyed copies from anthropic?
    • breezybottom14 minutes ago
      It&#x27;s going to be hard for you to do more of that, considering that AI forms are destroying the books. You know, the whole point of the article...
    • Certhas25 minutes ago
      Roughly Anthropic owners benefit from the content, unless the scans are made available somehow.
    • Sha1rholder13 minutes ago
      &gt; Roughly everyone can now benefit from its content.<p>You mean every Anthropic?
  • polarbearballs26 minutes ago
    Who knew Ray Bradbury could be so wrong.
  • callmeal27 minutes ago
    Yay capitalism! Imagine how much value is being created for shareholders by this destruction! Another example of &quot;great work&quot; being done!<p>Jokes aside, I don&#x27;t know what I would do in this situation either. Maybe find a way to mark &quot;rare&quot; or &quot;out of print&quot; books and sell&#x2F;auction the &quot;pages&quot;? Or make the scanned book available via some mechanism?
    • madaxe_again18 minutes ago
      You could “ingest” them all into a “model” and sell access to the information within them via “token usage”
  • pfdietz26 minutes ago
    Nowhere are the supposed rare books listed.<p>Confuses &quot;(not all that) old books&quot; with &quot;rare books&quot;.<p>Ragebait, nothing more.
    • madaxe_again20 minutes ago
      They are academic titles from the last few decades for the most part, and are rare because the print runs were typically short, and the prices typically eye-watering. “Semantic analysis of cotton processing terminology in Des Moines, 1993-1996, 8th edition”, that sort of fun.
      • pfdietz15 minutes ago
        I&#x27;m reminded of a quote that&#x27;s been (mis)attributed to Dorothy Parker.<p>&quot;This is not a novel to be tossed aside lightly. It should be thrown with great force.&quot;