10 comments

  • ladyanita224 hours ago
    This is something I&#x27;ve been fantasizing about for long.<p>Let&#x27;s say we took Rust, a language that makes parallelization easier than others (as it helps you avoid some common footguns). How difficult would it be to have a massively parallel computer system made out of many tiny, simple microcontroller-like chips? Let&#x27;s say we picked many little Risc-V&#x27;s. Surely this would be an interesting experiment (though I&#x27;m not sure whether it&#x27;d make economic sense or not...)
    • sigmoid104 hours ago
      It would certainly not make any economic sense, and I guess that&#x27;s also why noone is seriously looking into stuff like volunteer&#x2F;enthusiast clusters of home computers to do inference in the same way that e.g. LHC@home works. The main bottleneck for LLMs is still memory bandwidth. Any memory bus not directly soldered on your GPU is terribly slow. That&#x27;s why one big GPU with twice the VRAM will always perform significantly better than two GPUs with half the VRAM each. And it&#x27;s also not like you can just solder more memory onto a chip. At modern speeds, the speed of light is a hard limit. For current GDDR7, signals may only travel like 10mm per cycle.<p>If you spread such a system out over dozens or hundreds of tiny chips, you&#x27;ll be wasting most of its resources and lose hard to anyone who built a single chip setup.
    • akavel2 hours ago
      See GreenArrays&#x27; 144-core Forth chips by Chuck Moore.
    • ur-whale3 hours ago
      &gt; How difficult would it be to have a massively parallel computer system made out of many tiny, simple microcontroller-like chips?<p>It&#x27;s scaling the communication that becomes hard.<p>In this project they daisy-chain SPI. I don&#x27;t believe that would scale very far.
  • tdhz778 hours ago
    Soon ai in every lightbulb running Kubernetes
    • oneZergArmy5 hours ago
      Praise the Omnissiah.
      • Tade04 hours ago
        With the proliferation of Abominable Intelligence? Quite the contrary!
    • tombert6 hours ago
      You know, I&#x27;ve always liked Futurama but I always kind of thought it was silly that literally <i>everything</i> has an AI and a personality.<p>But, you know, I actually think that there might be a logic to it. Economies of scale might mean that almost-literally every computer you buy in the year 3000 has some kind of AI-assistance chip in there, and sure maybe it will have full AI with a personality spitting out one-liners.
      • abroadwin5 hours ago
        Kind of like how disposable vape pens often have a 24 MHz Cortex-M0+ with 3 kB SRAM and 24 kB flash, which would have seemed ludicrous a while back.
      • KeplerBoy5 hours ago
        Change the year 3000 to the 2030s and it might be just as accurate.
  • NDlurker8 hours ago
    I&#x27;m curious how this would handle grammar checking on a basic word processor. Or maybe generate worlds for small text based games. I have no idea what the capabilities are of a cluster like this.
  • librasteve4 hours ago
    haha … this is precisely the kind of project that <a href="https:&#x2F;&#x2F;bil-lang.org" rel="nofollow">https:&#x2F;&#x2F;bil-lang.org</a> is aimed at: Go for parallel (ie in this case pipeline processing).<p>don’t get too excited until we get the TinyGo backend built though ;-)
  • cameron_b10 hours ago
    It is a bit of a bummer to see that the degree of &#x27;compression&#x27; makes it a fancy llm noise-maker. It is still charming.
  • matthewfcarlson8 hours ago
    I’m actually working on a small project that’s exactly this! Less quant so it’s only 150M parameters but this is amazing.
  • sjakati989 hours ago
    Gemma 4 when?
    • nkozyra6 hours ago
      We&#x27;re gonna need a bigger ESP.
      • nkko15 minutes ago
        looking forward to try this one <a href="https:&#x2F;&#x2F;www.espressif.com&#x2F;en&#x2F;products&#x2F;socs&#x2F;esp32-s31" rel="nofollow">https:&#x2F;&#x2F;www.espressif.com&#x2F;en&#x2F;products&#x2F;socs&#x2F;esp32-s31</a>
      • pantalaimon1 hour ago
        <a href="https:&#x2F;&#x2F;www.espressif.com&#x2F;en&#x2F;products&#x2F;socs&#x2F;esp32-p4" rel="nofollow">https:&#x2F;&#x2F;www.espressif.com&#x2F;en&#x2F;products&#x2F;socs&#x2F;esp32-p4</a>
  • sneak4 hours ago
    Seriously though, what are the low cost chips that <i>can</i> usefully run LLMs? Is a Mac Mini the lowest we can go? Are there iGPUs on mini-itx that can do it, or are there dedicated AI chips that one could turn into a pi HAT?
    • Risse2 hours ago
      Depends on what you consider to &quot;usefully run LLMs&quot;.<p>Earlier this year, I bought a mini pc from Aliexpress, specs are roughly Ryzen H255, 24GB LPDDR5, 1TB SSD. This was around 350€ including VAT, customs, shipping etc. I would personally consider this somewhat of a lowest class of useful LLM box. It can run 8B models well, up to somewhere around 24B. I currently run Gemma 4 26B A4B Q5 on it, with MTP, and it is quite slow, but smaller models would run okay on it.
  • aidiscoverywire5 hours ago
    [flagged]
  • nonasking_6 hours ago
    Thanks for sharing. It&#x27;s fascinating to see a 0.5B LLM being split across seven ESP32s like this.