2 comments

  • librasteve10 minutes ago
    As the video points out, there is a hardware&#x2F;software feedback loop at play here. Since the hardware is deeply pipelined SIMD FPU datapath, the software is hand tuned machine coded transformations. In addition to limiting flexibility (every variation has to be ground out in CUDA), this prevents sparse matrix type optimisations. I predict that a set of general purpose CPUs - non shared memory at this scale - would be a much better use of transistor&#x2F;power. And you can code that at high level give a CSP style approach such as <a href="https:&#x2F;&#x2F;bil-lang.org" rel="nofollow">https:&#x2F;&#x2F;bil-lang.org</a>
  • tolugenius38 minutes ago
    I wonder what would &quot;compute&quot; mean if cpus were more efficient at matrix multiplication say 15 years ago. And on the flipside, what it would take to say train a frontier model entirely on cpus in the future.
    • actionfromafar6 minutes ago
      I wonder what &quot;compute&quot; would mean if CPUs were more efficient at matrix multiplication and vendors had the balls to pair each core to its own dedicated DDR and a star interconnect between.