3 comments

  • croemer1 hour ago
    In principle interesting, but I can&#x27;t stand the Claude writing.<p>&gt; Keys with real blast radius<p>&gt; Here is what they unlock.<p>&gt; This is a floor, not an estimate of actual balances or unauthorized usage. The keys were verified but never used.<p>&gt; We cloned the public dataset hub end to end: every repository, every branch, every large-file object<p>&gt; The size is only half the story. These are the training sets behind models people actually use. The worst-hit ones are named, card-documented pretraining corpora that open models were built on. We verified every credential we cite against its provider, so they were live when we looked.<p>The whole post looks like a Claude artifact with random little cards.<p>It&#x27;s also just too long, which is a side effect of using LLMs, it&#x27;s just too easy to create walls of text.
    • fractorial24 minutes ago
      I similarly find it interesting; I do not understand why it is not the default to just generate the thing as a draft, research anything you aren’t clear on, and re-write it in your own voice.
  • lorreyfum1 hour ago
    Wouldn’t it just be easier to crawl the net? Not sure what huggingface has to do with anything here.
    • croemer1 hour ago
      I guess Huggingface datasets are easy to crawl - those datasets are hosted to be crawled. In contrast to the web at large which will be behind Cloudflare etc