2 comments
As the video points out, there is a hardware/software feedback loop at play here. Since the hardware is deeply pipelined SIMD FPU datapath, the software is hand tuned machine coded transformations. In addition to limiting flexibility (every variation has to be ground out in CUDA), this prevents sparse matrix type optimisations. I predict that a set of general purpose CPUs - non shared memory at this scale - would be a much better use of transistor/power. And you can code that at high level give a CSP style approach such as <a href="https://bil-lang.org" rel="nofollow">https://bil-lang.org</a>
I wonder what would "compute" mean if cpus were more efficient at matrix multiplication say 15 years ago. And on the flipside, what it would take to say train a frontier model entirely on cpus in the future.