21 comments

  • jacquesm1 hour ago
    It&#x27;s the robbery of all of our culture to sell it back to us at a mark-up. Crimes this large are crimes against humanity. So many people whose life&#x27;s work got appropriated without consideration, compensation or consent it is baffling.<p>It is said that at the heart of every great fortune there is a great crime, so it should be no surprise that the most valuable companies on the planet will most likely result from this crime. And given that justice can be bought by those with the most money you can forget about anything coming of this.
    • pingou52 minutes ago
      &gt;It&#x27;s the robbery of all of our culture to sell it back to us at a mark-up<p>Would regulation help with that? Right now you can download free models that have been trained on that &quot;stolen&quot; data.<p>With regulation and compensation, only rich companies would be able to do that, and they would definitely not give it back for free. I put &quot;stolen&quot; in quotation marks because it&#x27;s still unclear if we can call that stealing. Nobody would say a human reading a book and learning from it is stealing. I&#x27;m not saying that a machine doing the same is equivalent, but the only think I am sure of is that I am not sure we can call it &quot;stealing&quot;.
      • CJefferson36 minutes ago
        We don’t have to treat people reading books and companies stealing all human knowledge the same.<p>Also, companies spent a long time telling us downloading single songs via Napster was the worst thing ever, before torrenting every book in existence themselves. I don’t believe any of these companies have paid for all the books they have trained on.
        • hlynurd34 minutes ago
          &gt;companies spent a long time telling us downloading single songs via Napster was the worst thing ever, before torrenting every book in existence themselves<p>really not the same entities here
          • CJefferson28 minutes ago
            No, but why should we accept they get away with it? Also Microsoft has definitely sued people for pirating windows and now collaborates with OpenAI and uses their ai trained on stolen materials.
            • hlynurd22 minutes ago
              Yeah that&#x27;s more fair
          • ipython11 minutes ago
            Well, it&#x27;s kinda converging, because Napster and Microsoft have teamed up to build a multimodal interactive video agent through a simple proxy API (this is a direct quote from Napster&#x27;s blog post)<p><a href="https:&#x2F;&#x2F;www.napster.com&#x2F;blog&#x2F;napster-heads-to-microsoft-build-with-omniagent-api" rel="nofollow">https:&#x2F;&#x2F;www.napster.com&#x2F;blog&#x2F;napster-heads-to-microsoft-buil...</a>
          • cgio26 minutes ago
            Not at leaf level, but if you trace the trunk, pretty sure you end up on the same one.
        • philipallstar4 minutes ago
          &gt; companies spent a long time telling us downloading single songs via Napster was the worst thing ever<p>One&#x27;s world cannot be so drawn in crayon that &quot;companies&quot; is a useful level of detail with something like that. There&#x27;s no irony in two totally different companies (one of which was actually an industry body, the RIAA) doing two totally different things.
        • hkt21 minutes ago
          &gt; We don’t have to treat people reading books and companies stealing all human knowledge the same.<p>We don&#x27;t. People engaging in piracy have their lives ruined, companies engaging in piracy pay a tiny fraction of their revenues out to authors who can&#x27;t legally outgun them.<p>(Sorry, I just wanted to air the juxtaposition as clearly as possible, I sense we are actually in agreement)
          • stego-tech5 minutes ago
            Came here to post this, got beaten by someone putting it far more succinctly than I would have.
          • ImHereToVote9 minutes ago
            I mean if you cite a copyrighted book verbatim. You are held liable. So should a company producing copyrighted work.<p>For instance a image&#x2F;video generating model.
      • barnabee8 minutes ago
        Regulation that said something like “we own 50% of your profit or 20% of your revenue, whichever is the larger” would.<p>If Apple can charge 30% to gate-keep mobile payments, we can surely charge that for the total information output of humanity.
      • Luker8815 minutes ago
        &gt; and they would definitely not give it back for free.<p>...not like they are doing it for free now either.<p>open-weight is an economic war strategy of trying to undermine your competitors and prevent it from rising prices, thus preventing profit, driving them out of business.<p>&gt; I put &quot;stolen&quot; in quotation marks because it&#x27;s still unclear if we can call that stealing<p>It never was stealing: you can&#x27;t steal a book by copying it. You <i>can</i> however commit copyright infringement.<p>This blatant disregard of licenses and copyright is clearly infringing on the authors ability to make a profit from their work, which was the whole point of copyright.<p>They knew it too, which is why they said nothing about the pirating and infringing until they got too big to fail.<p>So now we are left discussing and wasting time on what <i>technically</i> counts as infringing, pirating, stealing and whatnot.<p>All the while the small authors who can&#x27;t possibly lawyer up against the literal biggest corporations on earth will just have to shut up.<p>Yet, somehow they had deals with Disney and other big names, proving that they did actually feel they need approval.<p>Their actions are two-faced, thus proving malice. Now we can go back to pointless technicalities.
      • steveBK12320 minutes ago
        Well theres at least two different buckets of this.<p>First is the scraping of the open internet.<p>The second is the paywall bypassing, YouTube audio recording, and pirated content training that the labs have basically admitted to in one form or another.<p>Content from both gets served back to us, in exchange for watching ads&#x2F;paying a subscription&#x2F;paying tokens.<p>The second is more immediately hypocritical because they are license&#x2F;copyright&#x2F;DMCA violations that the little guy could get sued for while the labs get $2T valuations for. The automation of crime at scale, which is a common VC pattern.
        • robinsonb52 minutes ago
          And yet a third bucket is the license-laundering of GPL code when the entire github corpus was vacuumed up.
      • fzeroracer19 minutes ago
        &gt; Would regulation help with that? Right now you can download free models that have been trained on that &quot;stolen&quot; data<p>We do have regulation against these issues. Companies spent years railing against piracy and IP theft enshrining it into law but now that it&#x27;s being done by them en masse it&#x27;s considered acceptable. The reality is that no regulation would help because we don&#x27;t have regulators willing to enforce it nor do we have a legal system designed to help individuals against mass theft by corporations.
      • jappgar3 minutes ago
        Regulation can mean all sorts of things, including declaring the models themselves illegal.<p>Tech bros have a hard time understanding this, but a state can and will enforce its laws, even seemingly absurd one, if it wants to.
      • mitxela35 minutes ago
        Culture robbery is not limited to AI. Any big concert for example is capitalismed to hell. So are neighborhoods. Where you used to have people just living, now you have an intentionally designed facade for people to live <i>within</i>. There was a comment on the 40C3 thread saying it&#x27;s got too capitalist because of the ticket cost, and idk about that because it&#x27;s always been hosted in commercial venues to my knowledge, but the vibe of the conference and the club itself are much less rebel than they used to be. Stuff like Burning Man now exists for people like Elon to go there and say &quot;I was at Burning Man&quot; and for people to get T-shirts saying &quot;I was at Burning Man&quot; and photos of themselves being at Burning Man more than for whatever the first few ones were about.
        • loloquwowndueo19 minutes ago
          The what thread, now? Got a link instead?
          • n_plus_1_acc3 minutes ago
            <a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49737787">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49737787</a>
      • embedding-shape37 minutes ago
        &gt; With regulation and compensation, only rich companies would be able to do that<p>Well, with some imagination, you can have regulation that forces companies to open up, not just close down.<p>Imagine a law that stipulates that if you want to offer &quot;LLM-inference-as-a-service&quot;, you need to also publish exact details about how it was trained, what datasets were used and also offer those exact weights for download.<p>Sure, this would never happen, but just offering another perspective on how laws and regulation can be used if it was wanted, locking stuff down and pulling up the ladder behind you isn&#x27;t the only way to use laws, although that is a very popular reason and approach.
        • ben_w12 minutes ago
          I am unclear how this would help anyone?<p>Any argument that writers and artists lose from these existing, would remain unchanged.
    • dataviz10001 minute ago
      The first time a saw a documentary about Tetris it really hit me what communism is -- nobody owns anything they invent or create. [0] It was a long time ago and I remember feeling sad watching the story.<p>It is this one line, Article 1 Section 8 Clause 8, that separated the United States from the disaster that was the Soviet Union:<p>&gt; To promote the Progress of Science and useful Arts, by securing for limited Times to Authors and Inventors the exclusive Right to their respective Writings and Discoveries;<p>I don&#x27;t think it is far fetched to call ignoring and disregarding the Copyright Right clause a communist revolution.<p>[0] <a href="https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Tetris#Spread_beyond_the_Soviet_Union_(1985%E2%80%931988)" rel="nofollow">https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Tetris#Spread_beyond_the_Sovie...</a>
    • CrimsonRain26 minutes ago
      Every time you&#x27;re writing software or building machines&#x2F;factories (which is automating things), you are committing a crime. Every time you learn from your superiors or colleagues, get better than them, get promotion or they get fired, you are committing a crime. Provide justice there first.
    • cyber_kinetist53 minutes ago
      At least the Chinese AI companies are doing good service open-sourcing their models back to the public.
      • mitxela34 minutes ago
        There are no open-source LLMs.
    • sneak9 minutes ago
      Copying data isn&#x27;t a crime.
      • nullbio2 minutes ago
        Agreed, but now they&#x27;re trying to stop other people from copying data so that they can be the sole gatekeepers of humanities collective knowledge.
    • pluc52 minutes ago
      Has it affected DD as deeply as it as affected software engineering? Guessing clients feel a lot more &quot;empowered&quot; or &quot;independent&quot; and knowledgeable these days? I liked doing DD, just as much as I enjoyed developing, but it must be dying a slow death too. What&#x27;s changed in how DD reports are produced?
      • dgellow29 minutes ago
        What is DD supposed to mean?<p>Nit: Please don’t use obscure acronyms when writing things to an international audience without defining them first… DD can mean so many different things
        • turzmo20 minutes ago
          Due diligence — pretty standard acronym in this community.
          • oblio15 minutes ago
            No it&#x27;s not, and I&#x27;ve been here more than a decade.
    • TacticalCoder13 minutes ago
      [flagged]
      • api9 minutes ago
        IMO if they didn’t have proper licensing to train on the data the model should not be copyrightable.<p>In the long term though I think models have no moat, so the cost will fall to the cost of compute and storage. Which is why they’re pushing AI safety panics: regulatory capture to outlaw open models and outlaw competition.<p>And yeah, EA is neither effective nor altruistic. It’s a cult, part of the “Rationalist” and adjacent cluster of tech cults. They’re to tech what Scientology is to Hollywood I guess.
    • Razengan26 minutes ago
      &gt; <i>sell it back to us at a mark-up.</i><p>What if it was for free, like Wikipedia?<p>&gt; <i>Crimes this large are crimes against humanity.</i><p>jfc no, sit down.<p>Try doing something about the actual evil shit like arms manufacturers and the politicians ordering the deaths and misery of millions from the comfort of their couch.<p>At this point in our civilization, all human knowledge NEEDS to be collated in one place and easily queryable. Otherwise it&#x27;s just too damn difficult to make any further progress at the edge of our understand; there&#x27;s just too much shit to learn &quot;manually&quot; (wait I&#x27;m not advocating for low-effort slop, chill)<p>It&#x27;s helping common folk who wanted to do something but didn&#x27;t know where to start, while legacy search engines increasingly lead to spam, shallow knowledge or outright predatory shit.<p><pre><code> Example: </code></pre> Not long ago I had the misfortune of becoming interested in some WarHammer 40K lore. Most of the links led to Fandom (the enshittification of Wikia) and that place is a cesspool of obnoxious ads.<p>That content was written by unpaid volunteers. Should Fandom keep profiting from their work for perpetuity? Should I not be able to get the gist of what the heck a Qoiazrjirnowerx@# is without wasting my mortal lifespan on a horrible website?<p>Or, if I need to ask something peculiar, should I post on Reddit or StackOverflow or HN and wait for someone to see it and deem to give a sufficient answer, only to have a pricky mod decide that the question doesn&#x27;t &quot;fit&quot; the community?<p>God hell no, if you don&#x27;t know how much bullshit AI could eliminate for the silent majority then you were either not doing much to begin with or you were probably part of that bullshit.
      • steveBK12318 minutes ago
        Defending the same entities getting large DoD contracts to use AI for killing?
        • Razengan3 minutes ago
          They&#x27;ve been using computers for killing for decades, who&#x27;s taking up pitchforks against computers?
    • noosphr9 minutes ago
      &gt;It&#x27;s the robbery of all of our culture to sell it back to us at a mark-up. Crimes this large are crimes against humanity.<p>Yeah the introduction of copyright was truly criminal.<p>&gt; So many people whose life&#x27;s work got appropriated without consideration, compensation or consent it is baffling.<p>Oh wait ...
    • erulastiel6 minutes ago
      “It is said that at the heart of every great fortune there is a great crime”<p>lol at this edgy 5th grade statement. So ridiculous.
  • sedan_baklazhan0 minutes ago
    AI overall is the ultimate piracy crime.<p>I wonder what a token cost would be if AI companies were to pay royalties to every author who made their business even possible.
  • leonidasrup35 minutes ago
    In case of programming.<p>How much do the current LLMs invent solutions for user tasks, how much they just copy and adopt existing open-source solutions from from Github and other code repositories?<p>This not a problem for open-source code under permissive software license, but works derived from open-source code with copyleft software license should be also under copyleft license.<p>Could the biggest commercial benefit of LLMs be just working around limitations of copyleft licenses?<p>What is the monetary value of human work put into copyleft software and later used to train LLMs? It&#x27;s hard to estimate, but the study &quot;Estimating the Total Development Cost of a Linux Distribution&quot;, estimated that it would cost $1.4 billion to develop the Linux kernel alone.<p><a href="https:&#x2F;&#x2F;consortiuminfo.org&#x2F;metalibrary&#x2F;estimating-the-total-development-cost-of-a-linux-distribution&#x2F;" rel="nofollow">https:&#x2F;&#x2F;consortiuminfo.org&#x2F;metalibrary&#x2F;estimating-the-total-...</a>
    • menaerus23 minutes ago
      They do invent code solution for the problem that exists in your codebase. Latter implies that the code solution LLM synthesizes is usually unique of a kind, so, it&#x27;s not a copy-paste neither it is a simple extract from &quot;another codebase&quot; and adopted.<p>IMO they operate pretty similarly to humans - we synthesize our solutions, and therefore build-up our knowledge, by collecting knowledge from multiple other sources, including technical books and blogs, open-source code repositories, and our past experiences.
  • TutleCpt20 minutes ago
    The most shocking point is that they have a Microsoft exec who knows what he&#x27;s talking about.
  • 472828477 minutes ago
    “Information wants to be free“.<p>It’s not “theft of labor”; the work was already done. If anything it is theft of “intellectual property” (aka “copyright infringement”), if you believe that is a thing, but not of the “labor” that went into it.
  • sajithdilshan40 minutes ago
    If someone asked what is &#x27;the largest theft of labor in human history&#x27; I would have thought slavery.
    • y-curious18 minutes ago
      Yeah gulags and other forced work camps also come to mind. But I guess this is a larger scale in terms of man hours
      • bcjdjsndon12 minutes ago
        But it&#x27;s copying...how is it theft? Your labour WASNT stolen was it?
    • mitxela34 minutes ago
      Never ended, just changed in form.
    • TacticalCoder8 minutes ago
      &gt; If someone asked what is &#x27;the largest theft of labor in human history&#x27; I would have thought slavery.<p>Then I take it you&#x27;re interested in factual information as to whom the biggest slavers were, which country was the last to abolish slavery (an african one, in the 1980s) and in which countries, today, there are still people selling slaves.
    • bcjdjsndon14 minutes ago
      No actually it&#x27;s when someone copies that blog post you did about react.js and puts it into a dataset, I&#x27;m not sure how they sleep with themselves the absolute monsters
  • rich_sasha3 minutes ago
    It’s not that different to the US helping itself to indigenous peoples’ lands in North America, decimating them with smallpox and alcohol, then generously offering reservations.<p>At least it’s consistent, is what I’m saying.
  • nullbio5 minutes ago
    It&#x27;s humanities collective knowledge and work. That&#x27;s why nobody should ever buy the narrative of distillation being a crime or theft. It should be a human right to distill these models. Distillation should be being provided as a service.
    • rich_sasha2 minutes ago
      Yeah, distilled, hosted by OpenAI and charged for. And don’t you try reverse engineer what they did!<p>If this was all open, I’d maybe half agree.
  • Weryj1 hour ago
    I think it’s more like ‘The absolute maximum possible degree of theft’ there can’t be larger, it’s everything current and past.
  • American8739 minutes ago
    I remember techchrunch.com making the argument that IP Infringment != Theft in the music piracy era.. how quickly the tide turns :)
    • mitxela33 minutes ago
      they did say theft of labor, not theft of the things being trained on
  • Neil4444 minutes ago
    I understand the sentiment and partly agree. But also, the original has not gone anywhere. You&#x27;re free to accumulate knowledge in the old way just as before. So maybe it&#x27;s not theft of knowledge that we should be angry about, it&#x27;s something else harder to define.
    • tdb78938 minutes ago
      I&#x27;ve heard people say &quot;theft&quot; of intellectual property a lot. Also stealing an idea is common parlance. Maybe it&#x27;s regional or something but I hear &quot;theft&quot; or similar used all the time for things other than physical goods that you lose access to.
    • pluc41 minutes ago
      Lots of the original content is no longer available. Bots kill sites, AI kills monetization - both results in the original material disappearing.
      • bcjdjsndon9 minutes ago
        Copying means we can both share in the knowledge, surely everyone on HN wants that right? Share the open source code for the good of everyone?<p>Hackers used to say &quot;information yearns to be free&quot; now they&#x27;re saying &quot;that&#x27;s my information and I don&#x27;t want you using it&quot;<p>Probably indicative of America&#x27;s wider downfall that they&#x27;ve all become so self interested
        • sneak6 minutes ago
          Hackers are irrationally anti-corporation. This is where the nonsensical AGPL came from, too.
  • proc043 minutes ago
    If corporations weren&#x27;t already owning the consumer, with AI it does this by many orders of magnitude. If something isn&#x27;t done to prevent AI from being used to farm the masses for data, we will be living in a sci-fi dystopia without a doubt.
  • wj13 minutes ago
    How is this different from Microsoft scraping to build Bing?<p>Honest question. There is a line in the sand somewhere apparently.
    • sethops12 minutes ago
      Bing sent traffic to the original source. AI answers don&#x27;t. That&#x27;s the line. It&#x27;s not complicated.
    • haxiomic8 minutes ago
      Linking to work, where ownership and attribution is clear and the owner has the ability to commercialise is a very different thing to “laundering” content through the model, quoting the midjourney developers here<p>&gt; &quot;We just need to launder it through a fine-tuned codex.&quot; [0]<p>[0] <a href="https:&#x2F;&#x2F;cybernews.com&#x2F;news&#x2F;midjourney-ai-images-art-lawsuit-copyright&#x2F;" rel="nofollow">https:&#x2F;&#x2F;cybernews.com&#x2F;news&#x2F;midjourney-ai-images-art-lawsuit-...</a>
    • oreoftw12 minutes ago
      Huge difference between building AI and a search index.
      • sneak7 minutes ago
        Why? In both cases the SaaS downloaded the whole web and derives 100% of revenue from content they didn&#x27;t make.
  • noosphr10 minutes ago
    I&#x27;d call the introduction of copyright the largest theft of human labor in history.<p>No one was compensated for all the free labor they did before the introduction of copyright which copyright holders then privatized. For example the Disney corporation would have had to pay the Brother&#x27;s Grimm estate for the use of Snow white under the copyright regime they instilled in 1998 with the Mickey Mouse Protection Act.<p>That we are finally having a sane pendulum swing towards no copyright is a breath of fresh air.<p>The only way the AI bubble could improve the world more is if we end up becoming a Type I Kardashev civilization to feed the data centers. Then when the bubble pops we suck up all the extra CO2 with all the now idle nuclear power plants we can&#x27;t shut down.<p>At the same time it&#x27;s truly baffling going on a site called _hacker_ news and seeing corpo talking points from the 90s&#x2F;00s regurgitated wholesale. Information wants to be free.
  • gyosko52 minutes ago
    And here we are, just watching and doing nothing..
  • bcjdjsndon13 minutes ago
    Copying isn&#x27;t stealing you babies
  • redsocksfan451 hour ago
    [dead]
  • aaron69552 minutes ago
    [dead]