10 comments

  • nacs25 minutes ago
    Claude&#x2F;OpenAI etc have taken and continue to take literally all data from the internet, printed books, image, audio, and video humans have created in all of existence without permission to train their models.<p>But when same AI company gets &quot;distilled&quot; or it&#x27;s own AI-generated content used to train other models, it&#x27;s suddenly immoral or illegal?
    • mrngld7 minutes ago
      People keep saying this without attribution. Meta and others got caught, and brought into court, over using torrented files. But those models trained on that data have long since been retired and replaced with new models based on new from-scratch training runs. OpenAI, Anthropic, etc all pay studios, newspapers, Reddit and others for access to data for training. They scrape the open web, but if that&#x27;s illegal a court hasn&#x27;t said so. The open web is open, after all. And they don&#x27;t seem to be stealing books, they seem to be buying physical copies and scanning. Seems legit, that&#x27;s what a human would do to learn from a book. They also pay big bucks for commercially curated data and training sets.<p>Just feels like there&#x27;s enormous CCP effort to put their labs on equal moral footing with everyone else when it&#x27;s not demonstrably the case. They want the West to hate themselves so we&#x27;re happy to squander our technological lead.
      • rsstack2 minutes ago
        &gt; they seem to be buying physical copies and scanning. Seems legit, that&#x27;s what a human would do to learn from a book<p>There’s still the open question on learn vs copy&#x2F;mimic&#x2F;repeat.<p>As a human, I can read a book I bought. I’m definitely not allowed to scan it and post its pages online and upload them to an archive of scanned PDFs without the authors’ and publishers’ permission.<p>IIRC the Meta legal case wasn’t even about LLMs, they just torrented and shared pirated files, whether with strangers or among employees. Those may or may not have been later used for training, but it was already illegal to just share among employees.
  • AlanYx2 minutes ago
    The open question here is how is Anthropic retaining these exchanges?<p>If Moonshot is using the API, normally Anthropic would not retain the exchanges, at least that&#x27;s the promise. If Anthropic is consistently retaining all exchanges from a class of customers because they&#x27;re &quot;flagged&quot; but not notifying those customers, how can an average customer trust it won&#x27;t happen to them?<p>If Moonshot is buying accounts and using those rather than the API, wouldn&#x27;t they set the &quot;no training on my data&quot; flag in the settings so as to go undetected for longer? If so, we get back to the question of why would Anthropic be retaining the exchanges?
  • tancop5 minutes ago
    So we&#x27;re randomly getting free Claude <i>and</i> helping China beat America at the same time?<p>The only problem with it is they might have worse security than Anthropic and your personal info gets leaked, but I don&#x27;t think it can happen that easy.
  • segmondy17 minutes ago
    So how come super genius Fable and Mythos didn&#x27;t stop them?
  • ahsillyme30 minutes ago
    This is weird. I can&#x27;t think of a better compliment that is simultaneously safer. They&#x27;ll always trail on data generation then, no? So I can&#x27;t see the point. Best case for moonshot they get to lie about benchmarks that are gamed regardless, is how I see it.
  • nojito16 minutes ago
    We also get a glimpse into what those models are being used for.<p>&gt;PLA-affiliated surveillance activity. One user that we assess was likely affiliated with the PLA used what they thought was Moonshot’s Kimi model to load surveillance data from a CCTV archive about a single targeted individual. The user asked Kimi to analyze the CCTV data to understand whether the tracked person was behaving abnormally. The CCTV data included video surveillance from hundreds of cameras in Chengdu, including cameras outside PLA facilities, institutes affiliated with the China Electronics Technology Group Corporation, and a major state-owned enterprise (SOE).
  • okokwhatever25 minutes ago
    Expected
  • OutOfHere27 minutes ago
    There is absolutely nothing ethically wrong with other models using GPT and Claude for training. Heck, GPT-Sol even helped train GPT-Luna.<p>There also is nothing wrong with using customer data for training, since millions if not billions of other users benefit from it. Again, even the big players recognize its relevance and do it.
    • mrngld14 minutes ago
      Except it violates terms of use, thus illegal, and then they lie when they claim they weren&#x27;t doing it when they know full well not only did they do it in the past but they were actively still doing it even as they denied it.<p>That&#x27;s pretty significant. Paid Chinese bots can deny it all they want, but every human not on their payroll can see how bad it is. If they lie about that then they&#x27;re capable of lying about absolutely anything, and it tells you everything you need to know about how they view you as a user.
    • nojito17 minutes ago
      They are using stolen api keys and credentials.
  • Laurel12341 hour ago
    So?