Really dumb question from a software guy. Why aren't the labs burning their frontier models into chips already? Seems like the performance gains and cost per request would be worth it. That said, I understand neither the economics nor the physical challenges to doing this.
Model SOTA moves faster than chips can be designed or produced. You'd need to commit to a particular model for years to get payoff while still burning buckets of money producing new SOTA models to keep up with the competition.<p>It's why everyone and their dog runs these things on GPUs. When a new model supercedes the previous one, so long as you've got the memory for it your chips aren't obsolete.<p>I'm looking forward to someone picking a model to be "good enough" (say, qwen 4.0 or something) and selling them as peripheral hardware
At this point, LLM's are "good enough" for all kinds of tasks. Instead of making them more capable, now the efforts are making them smaller and cheaper.<p>All aboard! We're racing to the bottom now.
<p><pre><code> > Model SOTA moves faster than chips can be designed or produced.
</code></pre>
From what I remember working in that area the hardest part is getting masks for a design. Masks were developed in the span of half an year. Masks also reusable, they can be mixed and matched and this is why fabless companies work with fabs to produce specialized masks for them, it saves time for consumer to have masks for some macroblocks prebuilt.<p>Here's my analysis of how to etch relatively big LM into silicon: <a href="https://news.ycombinator.com/item?id=47109252">https://news.ycombinator.com/item?id=47109252</a><p>Given some amount of work with the fab before main pipeline set (I think a year long process), one can then spew LM-on-a-chip in six months or less and much more than 2 per year, because there can be several LMs in pipeline.
FPGA speed is a common misconception. They often have run at quite modest frequencies, and come with limited memory capacity compared to a GPU at same price range.<p>Speed (clocks speed) , is dependent on your design layout, and for any resonably advanced layout, it takes lots of knowledge to push the clockspeed beyond 200Mhz on 'consumer'/prosumer models. (Compare to a few GHz for GPU/CPU).<p>FPGA speed shine where they can pipeline massively parallel calculations through pipelines with minimal lookups.
I don’t disagree with most of what you’re saying, except for one point: I must have gotten a dumb dog, I’m a little jealous…
They have trillions..<p>Lots of people would have happily taken GPT-4o as good enough for a lot of use cases a year ago and not lived to regret it.
They have trillions worth of obligations to their partners in terms of compute purchase agreements and equity. Burning models to chips isn't really something they can blow money on just for funsies, they need to be able to justify it. The above comments are hypotheses as to why they haven't yet.
I know FPGAs are more expensive than GPUs, but are they fast enough to justify the extra cost?
It isn't just a matter of speed, it's also a matter of model quality. If they take 6 months to burn Fable to chips, and it takes 2 years to break even between design, custom fab, energy savings, etc, are those chips even worth running when the new models that are running on GPUs at that point are producing 10x better quality results?<p>Sure, your 2.5 year old models are running faster, but you can't drop prices on them without pushing the break even point further out.<p>If the cost difference isn't incredibly significant, will people even want to pay for the 2.5 year old model, or will they get more value for their money paying more to get better results from the newer model?<p>There's a lot of open ended questions that I don't have the insiders knowledge for to suggest whether or not such a capital outlay would be a worthy investment.<p>My guess is that state of the art stuff will stay on GPUs and models burned into chips will be for "good enough" applications that people are still teasing out. Probably highly specialized models in automated sensor units and such.
FPGAs are FPGAs by virtue of putting on the chips vast, vast arrays of wiring that can be controlled by software. Any given utilization of the FPGA will leave large fractions of the chip resources unused. If you've got a highly stereotypical use case FPGAs will have a "highly stereotypical" set of components being unused, where it would be better to use that space instead to do real work. A lot of people only see the "pro" side of the FPGA proposition without realizing they come with some very substantial "cons" that are intrinsic to the way they work.
I guess that's why they work in particular niche spaces like a synthesizer where you have a max of 8 voices and every voice goes through the same pipeline (osc / filter / env / amp) and everything is necessarily running all the time. In that sense I suppose they're very good for modelling any kind of analog circuitry?<p>Even then, while there are some amazing FPGA-based synths available, companies like Korg just put their code on a raspberry pi and call it a day. The same is true for emulators (SNES Mini etc. are also just raspberry pis under the hood iirc)
You don't really get an FPGA for capability. CPUs are much more capable, and they're general-purpose so they can do absolutely anything with about the same efficiency and just a little more code.<p>You get an FPGA for timing. They're less capable, but (in many common design architectures), they output their results once per clock, every clock, on time, every time. If you can hit a fabric clock of say 100MHz, clocking all the weird logic you can stuff in there, it gives 100 million outputs per second, <i>never skipping a single one</i> for any reason (short of total failure). The penalty is that making a small change to your desired "program" can be very expensive, and many things won't be realistically possible at all. Or at least won't fit into a part that <i>you</i> can buy. But things like audio, video, and high-frequency trading <i>love</i> being able to guarantee timing.<p>(Of course there are other ways to write your FPGA HDL, but that's one of the more common ones. And you do see DDR-style clocking, and similar, every now and then.)
> with about the same efficiency and just a little more code.<p>It depends. For some things, CPUs don't even come close. An XCVU13P FPGA can handle 1.2Tbps of full-duplex Ethernet @ 1 billion pps. And that part costs less than a grand at moderate qty, and with significantly less power consumption than a CPU that'd be capable of operating a dataplane at these speeds.
Sorry, I spoke pretty sloppily there.<p>The point I was trying to make is that the CPU is a general-purpose creature and doesn't really care what you want it to do. If you <i>had</i> a CPU that could handle 1.2Tbps of Ethernet packets at 1Gpps, it could do a whole lot of <i>other</i> things involving 1.2Tbps of data flow too, very easily, if someone wrote the software. And more. (But you're probably not getting 2.4Tbps out of it, no matter what you do.)<p>An FPGA can not. There's plenty of things that those XCVU13Ps just can't do, or would do worse than a $1 microcontroller. (Setting aside for a moment implementing a CPU inside the FPGA... which does actually happen in just about every large-enough FPGA design, which is its own discussion....)
> In that sense I suppose they're very good for modelling any kind of analog circuitry?<p>That would be better suited to FPAAs (field programmable analog arrays). FPGAs can usually only work with clocked digital signals.
They're not magical go faster juice. I don't know of a microarch where they're faster than modern GPUs at ML training or inference.
They are more like a way to proving the architecture of the accelerator before committing 100's of millions into a custom ASIC with TSMC
1. No<p>2. They don't have enough capacity either<p>The current largest FPGA, the AMD Versal Premium VP1902 has 18.5 million logic cells. That's not even enough for the smallest whisper.cpp model (75M).<p>You'd have to order hundreds of thousands of them (or millions) to serve even a single copy of a frontier model, and at that scale inference quickly becomes starved by the speed of light.
Well, you'd use BRAM to store model weights, not fabric. But still, you only get a couple hundred MB for probably close to US $100k per chip.<p>It's likely that the major FPGA vendors will soon announce parts specifically architected to support LLMs and similar models. But the current generation isn't suitable for that at all.
I think one limiting factor here is that Big AI cannot focus on building solid, stable products that do a job well. It's okay if it happens incidentally, but doing it on purpose would be self-defeating.<p>Their stock price, be it public or estimated, is heavily pricing the notion that they are first and foremost Growth companies. Therefore their focus must remain on ever better and greater things. If they lose focus and get distracted by lesser endeavors, their valuations crumble, their ability to raise capital vanishes, and their runways collapse before they ever have a chance to reach their end goal, whatever that may be.<p>That means the boring job of productizing AI models into reliable systems that won't vanish in six months is left for a smaller company willing to pick up the crumbs. Unless they get acquired by Big AI before getting it done.
They [1] are [2].<p>[1] <a href="https://taalas.com/" rel="nofollow">https://taalas.com/</a><p>[2] <a href="https://chatjimmy.ai/" rel="nofollow">https://chatjimmy.ai/</a>
8 months ago, Taalas claimed <a href="https://taalas.com/the-path-to-ubiquitous-ai/#:~:text=Upcoming%20models" rel="nofollow">https://taalas.com/the-path-to-ubiquitous-ai/#:~:text=Upcomi...</a> that "Our second model, still based on Taalas’ first-generation silicon platform (HC1), will be a mid-sized reasoning LLM. It is expected in our labs this spring and will be integrated into our inference service shortly thereafter. Following this, a frontier LLM will be fabricated using our second-generation silicon platform (HC2). HC2 offers considerably higher density and even faster execution. Deployment is planned for winter."<p>Nothing was released in spring, and 2 months ago AMD announced their acquisition of Taalas. That doesn't exactly inspire confidence that their frontier LLM will arrive as promised.
> AMD announced their acquisition of Taalas. That doesn't exactly inspire confidence that their frontier LLM will arrive as promised.<p>Why not? AMDs chip design experience and production capacity are magnitudes larger than a small startup. Assuming that they acquired Taalas for their technology, I don't see a reason why they couldn't.
I think this company was recently acquired by AMD, so hopefully they'll start getting some this into production. I know OpenAI was working on model-on-a-chip too.
Lead times are so long that there is a lot of risk the chips would be obsolete by the time they come out.<p>Also, it's hard to get fab capacity for any project. Let alone something so experimental.
Most responses here are along the lines of "model capabilites move too fast to build hardware for".<p>I think the fact that there are plenty of 1yr+ old models on openrouter serving hundreds of billions of tokens a month shows that there's plenty of use case for models that are "good enough. Cerebras' entire business is serving older models at high speed. I would happily use an opus 4.7 at 15k tokens per second. The intelligence per second of an ASIC still makes sense even with rapidly evolving models.
"Intelligence per second" is a striking phrase! I'll be turning it around my head at a moderate IPS until I hopefully make something of it.<p>But you're right in sense: moderate intelligence at superhuman rates (and presuming moderate energy usage) is very compelling compared to an intelligence that takes 1000 years to return "42"
I look at it this way. A year ago is a long time in AI terms, but <i>not that long</i>. Those models were already decent for the tasks we're using SOTA models for today.<p>So imagine taking a year-old SOTA model and running it at 100 tokens per second on an edge device. That's enough to feed a screen's worth of content through it and power <i>decent</i> multilingual message suggestions on IM.<p>Imagine running it at 1000 tps. That's enough to reparse that screen mid-keystroke, and give you semantic autocomplete in text. Or fully general "the phone has a good idea of what you're attempting to do" context at all times.<p>There's many, many new classes of features that will open up if decent enough models can be run on edge devices at 100+ "intelligence per second".
Also "faster" indirectly means "more intelligent" with reasoning models.<p>For example, Opus 4.6 Max was somewhere between Opus 4.7 Medium and High in some benchmarks, but it was slower. If there was a way to run it 10x faster, the economics would be different.
Totally. But it's worth noting that this is a pretty new thing! I wouldn't have bet on that a year ago, but now I would.
Yes, and startups do exactly that. Check out Etched <a href="https://www.etched.com/" rel="nofollow">https://www.etched.com/</a> where they made a Transformer specific GPU (basically a form of ASIC) where they bet that transformers would be the dominant GPU architecture for running AI / LLM workload.
From what I've read, not only are some labs doing it (other commenters already mentioned).<p>But it's complicated for other reasons, one being that the number of parameters for frontier models (especially with MoE models) are so high, and not always utilized (once again, thanks to MoE) that it would actually be incredibly cost prohibitive, if not impossible, to attempt to make giga-chips that would allow running it.<p>I definitely do believe that we will see more and more specialized chips over time, but putting the entire model on a chip is still a ways away.<p>I believe Taalas has a heavily handicapped llama 8-billion parameter model. And it still pulls >200W to run.<p>I can't imagine how anthropic or open ai would be able to burn a multi-trillion parameter model on a chip, we just aren't there yet.
I take 200W for 17,000 tokens/sec any day, even just for Llama 3.1 8B. For that size I want to see this one on silicon.
<a href="https://huggingface.co/xunlinkx/MiMo-V2.6-Distill-Qwen-9B-Text-oQ4e-mtp" rel="nofollow">https://huggingface.co/xunlinkx/MiMo-V2.6-Distill-Qwen-9B-Te...</a>
Etched is a startup doing exactly this.<p><a href="https://www.etched.com/progress/frontier-inference-clusters" rel="nofollow">https://www.etched.com/progress/frontier-inference-clusters</a>
The bottleneck, for inference at least, is memory bandwidth. And that you can't make any faster by making it specific to your model.<p>So companies try to maximize the memory bandwidth they can get, balancing tradeoffs of power/area/programability of their chip. Right now they feel like the economy on power/area is not worth the decrease in programability/flexibility.
Presumably though the kernel has a pretty specific set of operations done against the weights in memory. Burning the weights into the memory with local memory cores capable of the kernel operations would be a lot more efficient than round tripping busses.<p>The primary constraint isn’t likely what’s possible to do, but that the kernel and weights are too variable right now and the patterns too poorly established to bake into hardware accelerators yet. Margin pressure is also not there yet.<p>I suspect as the marginal utility of the frontier improvement settles into diminishing returns (I suspect we are there already tbh) baking hardware models with ROM, working set, and kernel cores collocated will be the frontier space as the goal will become reducing capital spend to utility levels rather than research levels.<p>Once someone has a model that is sufficient for almost any practical use, making marginal inference cost effectively zero will be the competition frontier. I do shed a tear for all those lonely data centers as compute densities will almost certainly make most of them a terrible investment.<p>But such is the cycle
You're starting to hint at compute-in-memory as a general replacement for CPU/DIMM layouts. That could be useful for far more than LLMs, world models, or any sort of AI. It takes a bit of a different software development stack than a standard architecture though.
> Burning the weights into the memory with local memory cores capable of the kernel operations would be a lot more efficient than round tripping busses.<p>Sure, but now we're not talking about just burning the weights into the chip, but also designing a new architecture that has memory local to each core. A new architecture would then require a new programming model, which means new inference stack, which may mean new training stack.
I don’t think it would require a new training stack, and I’d imagine it makes more sense to distribute cores with memory. The cores can be simplified to the functions of the kernel since the inference kernel can be expressed as a reduced set of optimized functions in the pipeline rather than a general CUDA core. If the model is burned into ROM, the compute pipeline can be baked into the core.
Burning the weights into static mask ROM is pretty trivial. Taalas is doing it.
Would you pay to crystalize one of today's models in silicon so you can use it in 2028, or would you wait for another 6 months to see how models improve before pulling the trigger on that kind of commitment?
Naive question: would flash or some kind of write once memory be dense enough to replace a mask ROM? I'm wondering whether it is feasible to manufacture "blank" chips at the semiconductor factory that have a fixed architecture, but are model-agnostic. These chips would get the then-current weights burned in on first use. It's probably quite wasteful, but it could keep the same chip design alive for a longer time.
If I controlled a budget like this, I think I would put some portion of it toward paying to crystallize one of today's models in silicon, yes. Not 100%, but I do think this makes sense to invest in at this point. I would not have said so a year ago.
This question would be better answered if people were careful about distinguishing between "models" and "transformer architecture".<p>If you bake a given transformer architecture into silicon and then, a year later, changes in transformer architecture give a large inference performance boost, you may have to throw away all that now nearly-useless silicon that gets outperformed by humble GPUs.
We are still in the middle of the AI race. Commodity hardware is easy to use, can do everything and is fast enough.<p>Your optimized hardware chip might be obsolete before its back from the fab.<p>SOTA Frontiermodelhardwarechip is a benchmark point of a potential model slow down.<p>Google is doing it right now under project Frozen v2 which should be ready by 2028? which is either just a small experiment or flexible enough and thats why it takes so long for it to happen.
Cuz once they're in chips they can be put into robots, and once they're in robots it won't be so easy to reach the off switch, and once we can't easily reach the off switch, we're doomed.
Models fully deprecate in a few months. Why would you burn an algorithm that fully depreciates in value faster than a bag of potato chips. The 'inefficient' general purpose hardware is constantly renewed with every released model. Even 6 year old Ampere GPUs are still usable.
Because the iteration speed on models is so fast that by the time they have an ASIC ready for one model version, they are already significantly ahead in capability. Think of how big the jump between Opus 4.8 and 5.5 has been. They were released four months apart.
I'd presume because it take too long to go from design to tapeout to production. Their whole business is predicated on having better models.<p>Also can't keep them closed source if you do that.
Extropic?
taalas did: <a href="https://taalas.com/" rel="nofollow">https://taalas.com/</a>
Yeah, and what about FPGA? Which was the same interim state when Bitcoin went GPU -> FPGA -> custom chip fab?
Not sure that's actually practical at the scale of SOTA models.
They do. It takes time to deploy those chips though. Check out OpenAI and Broadcom deal.
Even dumber question: What is new or novel about this openTPU?
In addition to some of the other replies you got, here is one more:<p>Much of a model are weights, and high-density ROMs are very very very hard.
[flagged]