Cerebras CS-4

(cerebras.ai)

410 points by sunils3416 hours ago

27 comments

  • KronisLV8 hours ago
    God I wish they&#x27;d back up all of those claims by offering a subscription of Kimi K3 and GLM 5.3, not some outdated GLM 4.7 instance that they then proceed to call a preview model and say that they&#x27;ll remove it, leaving users only with GPT-OSS 120B which is nigh useless nowadays: <a href="https:&#x2F;&#x2F;support.cerebras.net&#x2F;articles&#x2F;9996007307-cerebras-code-faq" rel="nofollow">https:&#x2F;&#x2F;support.cerebras.net&#x2F;articles&#x2F;9996007307-cerebras-co...</a> and <a href="https:&#x2F;&#x2F;www.cerebras.ai&#x2F;pricing" rel="nofollow">https:&#x2F;&#x2F;www.cerebras.ai&#x2F;pricing</a><p>Guess they don&#x27;t care about regular devs atm and are focused only on hardware sales.
    • dannyw4 hours ago
      Why would they offer a coding subscription and start competing with some of their biggest customers; when they are capacity-bound and companies like OpenAI will take however many wafers that Cerebras sells to them?<p>OpenAI&#x27;s Sol ultrafast (powered by Cerebras) is still in preview, presumably because they&#x27;re overall capacity bound.
      • KronisLV2 hours ago
        &gt; Why would they offer a coding subscription and start competing with some of their biggest customers; when they are capacity-bound and companies like OpenAI will take however many wafers that Cerebras sells to them?<p>Because they already have &#x2F; had an okay coding subscription product for a bit and it gives them visibility and mindshare (in regards to their hardware, even if they don&#x27;t compete with other providers that much). They could do what Kimi did - make a good subscription with good models, once you get enough customers to get some good PR and such, pause the signups so you don&#x27;t have to spend more on running the service than you want&#x2F;can. Do enough of that and people will talk about your offerings organically, make yourselves known to even devs as &quot;That one company with their own hardware and the super fast subscription.&quot; experiencing which would do more than any marketing.
        • dannyw2 hours ago
          Coding subs are good when they promote usage and adoption of <i>your</i> models in enterprises at API rates.<p>Cerebras is a B2B hardware company. It feels like a distraction: think of the opportunity cost, and resources&#x2F;headcount not working on other things that would drive more impact.<p>Should NVIDIA do a coding subscription too? I&#x27;m sure they can make money off it, but I think it would be -EV.
      • giancarlostoro1 hour ago
        They do have Cerebras Code<p><a href="https:&#x2F;&#x2F;www.cerebras.ai&#x2F;code" rel="nofollow">https:&#x2F;&#x2F;www.cerebras.ai&#x2F;code</a><p>But it&#x27;s not fully open to just anyone, I wasted time signing up to find out that I couldn&#x27;t even sign up for it to test it out.
    • GenerWork3 hours ago
      &gt;Guess they don&#x27;t care about regular devs atm and are focused only on hardware sales.<p>Why would they want to target regular devs right now? If they sold to regular devs instead of enterprises, the complaint wouldn&#x27;t be about model choice, it&#x27;d be about how expensive it.
      • tedivm2 hours ago
        Seriously, even as well paid as many devs are these are not machines that are affordable for personal use. Their market is people slapping down millions on frontier model training.
    • scosman7 hours ago
      I don’t think they will until they change the architecture.<p>They don’t have a prefix cache like other providers, or at least don’t have a discount in their billing structure. Each message charges for the whole context window. It’s wildly more expensive for long multi turn scenarios with lots of tool calls (coding). It’s better for short few turn tasks.<p>Edit: I don’t know if they actually have a proper cache. This could just be a billing artifact.
      • dannyw4 hours ago
        Cerebras supports prompt caching and has a doc about it. A fairly standard automatic prefix-based implementation with 5min expiry.<p>They do not seem to discount cached input for the self-serve Developer tier. Maybe they do for enterprise rate cards?<p><a href="https:&#x2F;&#x2F;inference-docs.cerebras.ai&#x2F;capabilities&#x2F;prompt-caching" rel="nofollow">https:&#x2F;&#x2F;inference-docs.cerebras.ai&#x2F;capabilities&#x2F;prompt-cachi...</a>
        • scosman1 hour ago
          ah, that makes it feasible! Okay, glad it&#x27;s not technical limit. They should fix the pricing...
    • dilap2 hours ago
      GLM 4.7 is gone (at least for us), with no suitable replacement from Cerebras. I think all they care about now is hardware and OpenAI hosting.
    • preommr8 hours ago
      &gt; GPT-OSS 120B which is nigh useless nowadays:<p>I still think that was a really great model that got overlooked. It was really great in terms of latency&#x2F;throughput while still being fairly intelligent.<p>I was planning on using it for a design tool, but moved over to luna since it&#x27;s comparable speeds and cost for a lot more intelligence.
      • KronisLV7 hours ago
        &gt; I was planning on using it for a design tool, but moved over to luna since it&#x27;s comparable speeds and cost for a lot more intelligence.<p>Everyone should occasionally go back to the old models to see how much worse they were, like even a year ago you could generate results but they were typically full of bugs and you have to fix a non-insignificant amount of it all manually: <a href="https:&#x2F;&#x2F;blog.kronis.dev&#x2F;blog&#x2F;i-blew-through-24-million-tokens-in-a-day&#x2F;" rel="nofollow">https:&#x2F;&#x2F;blog.kronis.dev&#x2F;blog&#x2F;i-blew-through-24-million-token...</a><p>Admittedly that post was before agentic development truly took off and that 3k EUR figure when paying per API tokens would nowadays be closer to like 6k EUR for the volume of work I do, but still.<p>It&#x27;s the same how Qwen 2.5 was pretty problematic for anything remotely serious, same with Qwen 3 Coder Next (80B), and at least the most recent versions are getting better but still not <i>quite</i> good enough in real world use cases outside of benchmarks. They&#x27;ve come a long way, regardless!
        • 5555watch6 minutes ago
          &gt;Everyone should occasionally go back to the old models to see how much worse they were, like even a year ago you could generate results but they were typically full of bugs<p>Oh yeah, I&#x27;m still amazed how good the current iteration of models are for coding (I have a fear it&#x27;s too good to be true - so will get taken away..). Exactly a year ago I switched from GPT 5 to Gemini just because the coding with R language was terrible; and even with Python it kept forgetting and mixing basic stuff. Gemini at the time had much longer context window and was miles ahead on R syntax.<p>Current experience of just leaving a Codex Agent chug until a stable solution is completed is still mind blowing to me.
      • wongarsu7 hours ago
        As MoE with 5B active parameters it&#x27;s pretty fast. But you still need a lot of vRAM, or have to run small quantitations. Qwen models just gave you more bang for your buck, and the gap became worse with every qwen release
      • LoganDark7 hours ago
        gpt-oss-120b is absolutely unusable over Cerebras. It fails to call tools half the time and just continues to think about what tool it&#x27;ll call repeatedly. Like it says it&#x27;ll call a tool and then it doesn&#x27;t, and then it says it&#x27;ll call the tool again and then it doesn&#x27;t, and it just does that in a loop forever. It&#x27;s awful. Also forgets to end the thinking block too. Even if the model itself was just-okay for its time, even at 1000t&#x2F;s+ it&#x27;s not worth it. And it&#x27;s EXPENSIVE, like $5 per minute expensive
    • maxdo3 hours ago
      Why they should go with Chinese models if they have a line up of gpt models and a very good partnership with someone who lives in the same jurisdiction and not in the country that convinces their citizen that it’s a good idea to go on war with western world ? Just curious ?
      • KronisLV2 hours ago
        &gt; Why they should go with Chinese models<p>Because they generated some buzz and are near-SOTA and would be a great benchmark for a PoC subscription that doesn&#x27;t necessarily aim to compete with other vendors at a similar scale (since their main business is the hardware). Mistral is conceptually cool but is lagging behind. I guess Muse Spark and Laguna would also be okay, just not as recognizable. Meanwhile both Kimi K3 and GLM 5.3 are near-SOTA in performance and considerable in size, a great choice for proving the platform!<p>As for the 2nd part of your question - that wasn&#x27;t a relevant concern or consideration here, unless the models would be tainted to a degree to prevent them from having a good coding subscription that gets more developer mindshare towards what their chips can achieve and generate some good PR.
    • 0xbadcafebee1 hour ago
      Cerebras is the fastest provider <i>by far</i> on OpenRouter, and gpt-oss-120b is still very useful. They have backed up their claims very well.<p>&gt; Guess they don&#x27;t care about regular devs atm and are focused only on hardware sales<p>They aren&#x27;t trying to make a few bucks off tokenmaxxers. They&#x27;re trying to be the underpinning of compute for all AI. They&#x27;re going to beat Nvidia.
  • bearjaws3 hours ago
    I was hoping to see Cerebras launch something other than GPT-OSS-120b in production this week, especially with GLM4.7 going away.<p>If they could launch Qwen 27b or Deepseek Flash that would be amazing.
  • syntaxing15 hours ago
    I think the fun takeaway from this is that GPT 5.4 is probably 45B active parameters and GPT 5.6 Sol is closer to 50B.
    • yorwba9 hours ago
      You cannot infer this because they only show the tokens per second per user. One way to get a higher number is to have fewer users per chip.<p>I&#x27;m pretty sure Cerebras has a confidentiality agreement with OpenAI, and this press release was carefully constructed to avoid leaking details about the model weights. For example, the graph of tokens per second vs. tokens per second per user doesn&#x27;t have any numbers that would allow you to translate between the two. (And in any case the relationship depends on the model.)
      • petu2 hours ago
        They show that CS-4 can&#x27;t really do batching (or rather it can&#x27;t properly benefit from it), total throughput barely changes (25%?): <a href="https:&#x2F;&#x2F;cdn.sanity.io&#x2F;images&#x2F;e4qjo92p&#x2F;production&#x2F;6a132331880d41c8ea20b584f4fd37c270741692-1920x1080.png" rel="nofollow">https:&#x2F;&#x2F;cdn.sanity.io&#x2F;images&#x2F;e4qjo92p&#x2F;production&#x2F;6a132331880...</a><p>Which I think makes it feasible to approximate activation from CS-4 tokens per second per user.
    • logicallee14 hours ago
      (Where did you see that?)<p>This was also interesting: &quot;CS-4 delivers more than 1,000 tokens per second on models exceeding 10 trillion parameters.&quot; Was it known that there were 10 trillion parameter models in use?<p>I think the frontier providers keep the size of their models carefully hidden.
      • redox9911 hours ago
        You don&#x27;t really need to train a 10T model to test cerebras against a 10T model. You can feed it an untrained (randomly initialized) model and benchmark it. Result will be gibberish but performance the same.
      • nl12 hours ago
        Mythos&#x2F;Fable are around 10T:<p>&gt; According to FT, industry estimates say Anthropic&#x27;s most advanced Mythos 5 has about 8 trillion parameters and Fable 5 about 5 trillion<p><a href="https:&#x2F;&#x2F;www.reuters.com&#x2F;technology&#x2F;bytedance-targets-mega-ai-model-nearing-anthropics-mythos-ft-reports-2026-08-07&#x2F;" rel="nofollow">https:&#x2F;&#x2F;www.reuters.com&#x2F;technology&#x2F;bytedance-targets-mega-ai...</a><p>I believe this report has confused Opus (which is known to be around 5T) and Fable.<p>Other reports say 10T. See for example <a href="https:&#x2F;&#x2F;eu.36kr.com&#x2F;en&#x2F;p&#x2F;3760679047267075?ref=explainx" rel="nofollow">https:&#x2F;&#x2F;eu.36kr.com&#x2F;en&#x2F;p&#x2F;3760679047267075?ref=explainx</a> where Musk talks about the models being trained on Colossus2
        • YmiYugy10 hours ago
          I&#x27;m confused. I thought Mythos 5 and Fable 5 were exactly the same model just with a different security layer in front of it. Could they mean the Mythos 5 Preview?
          • nl8 hours ago
            Yes. One reason why I think that report has confused Fable and Opus.
        • zozbot23411 hours ago
          &gt; I believe this report has confused Opus (which is known to be around 5T) and Fable.<p>5T for Opus feels quite high though. DeepSeek V4 Pro is a mere 1.6T and often described as a match with Opus in overall quality. Even the largest open models in common use are around 2.8T.
          • Implicated10 hours ago
            &gt; and often described as a match with Opus in overall quality<p>It&#x27;s not. Idk about who has more T&#x27;s but, unfortunately, DS4 pro is not a match to Opus, at least not Opus 4.8.
          • nl8 hours ago
            It&#x27;s not an Opus match.<p>The difference is very visible in long tail applications. Exactly where you&#x27;d expect parameter count to matter.
      • alightsoul12 hours ago
        pretty sure 10 trillion parameters is now the norm among closed ai labs, given that nvidia also references the same 10 trillion number for their nvl72 racks
        • KronisLV8 hours ago
          Pretty bad efficiency then unless that only applies to Fable class but even then - Kimi K3 is around 3T and does similarly well in most benchmarks.
          • esafak3 hours ago
            As others have noted, most benchmarks stress the torso, not the tail.
      • ewild14 hours ago
        It&#x27;s rumored fable is around that 10T number
        • walrus0114 hours ago
          If this is true, it&#x27;s even more impressive that some of the open weight models that are &lt;3.5T in size, approx 33% of its size, are within a few points of it in the artificial analysis leaderboard.
          • manquer14 hours ago
            Not necessarily, there could be diminishing returns on mere parameters count .<p>There is nothing to say for example a 1 Quadrillion parameter model will be vastly more intelligent than current SOTA especially since new training data is largely synthetic today
            • HDBaseT13 hours ago
              That&#x27;s precisely what he is saying, there is diminishing returns (or optimization left on the table).
              • manquer13 hours ago
                I read it as it is impressive because smaller models 2.5T are squeezing similar returns as 10T models despite being 1&#x2F;4th size not that there beyond 2T today the number or parameters do not have much meaning
                • verdverm12 hours ago
                  or the latest qwen3.8 27B doing so well at ~1&#x2F;100 the size of K3
                  • rnewme12 hours ago
                    What about general knowledge you can get out of it before hallucinations start?
                    • stymaar10 hours ago
                      Storing general knowledge in VRAM has always been a dumb idea in the first place.
                    • walrus0111 hours ago
                      It did OK on schlongbench v1.0 (test of a specific niche word that doesn&#x27;t make it into smaller LLMs) but it sure does love to count words<p><a href="https:&#x2F;&#x2F;pastes.io&#x2F;r8F1AY8h" rel="nofollow">https:&#x2F;&#x2F;pastes.io&#x2F;r8F1AY8h</a>
                    • Balinares8 hours ago
                      Qwen 3.8 27B beats Opus, Fable and GPT 5.6 by a comfortable margin on the AA-Omniscience Hallucination Rate benchmark.
                    • verdverm11 hours ago
                      I do not rely on any LLM of any size for general knowledge baked into the weights, they all hallucinate and that is the wrong way to hold them imo<p>I think there is some merit in that smaller models cannot memorize so much of the training data, i.e. that they are less likely to do copyright infringement, and by analogy not having memorized SDK &#x2F; API surfaces that have since changed from the training data
                      • walrus0110 hours ago
                        &gt; I do not rely on any LLM of any size for general knowledge baked into the weights<p>You have to rely on it to a certain level for agentic&#x2F;coding work, presuming that&#x27;s the general subject we&#x27;re talking about here... For instance I recently encountered a project where it would have been a lot worse if the LLM didn&#x27;t already know &quot;what is&quot; xterm.js and a bunch of its associated npm-related&#x2F;node related software. If it was still smart but had to google and find results for everything it would have been a lot more time consuming and risked sending it down a wrong path.
          • habosa13 hours ago
            GLM 5.3 is &quot;only&quot; 753B parameters. Much much smaller.
          • mlmonkey13 hours ago
            You want to take a look at the &quot;Scaling Laws&quot; paper, so you can extrapolate from these numbers.
            • stymaar10 hours ago
              This paper, as well as the Chinchilla one, aged like milk though.
          • scosman7 hours ago
            And GLM is only 0.7T!<p>But these labs distill off the larger models. Both officially at the labs with the big ones, and unofficially. We need the giant models to get the smaller models.
        • johnnyApplePRNG12 hours ago
          Fable is most definitely nowhere near 10T.<p>The cost to train and infer that would be insane, even by today&#x27;s standards.
          • nl12 hours ago
            Fable is strongly believed to be around 10T. The most conservative estimate I&#x27;ve seen is 8T.<p>Eg: <a href="https:&#x2F;&#x2F;www.reuters.com&#x2F;technology&#x2F;bytedance-targets-mega-ai-model-nearing-anthropics-mythos-ft-reports-2026-08-07&#x2F;" rel="nofollow">https:&#x2F;&#x2F;www.reuters.com&#x2F;technology&#x2F;bytedance-targets-mega-ai...</a><p>That reports Mythos as 8T and Fable as 5T, but I think they mean Opus as 5T, which is widely known, eg: <a href="https:&#x2F;&#x2F;eu.36kr.com&#x2F;en&#x2F;p&#x2F;3760679047267075?ref=explainx" rel="nofollow">https:&#x2F;&#x2F;eu.36kr.com&#x2F;en&#x2F;p&#x2F;3760679047267075?ref=explainx</a><p>Both Grok and Bytedance are training 10T models.
            • stymaar10 hours ago
              The fact that Musk claims Opus is 5T to justify why Grok is far behind should be taken with a massive grain of salt given he&#x27;s a recidivist mythomaniac.<p>Honestly if Opus is 5T parameters while being matched by the biggest open models that are at least twice smaller, it would mean that the US is already behind China in the AI race, despite a significant edge in compute.
              • nl5 hours ago
                The open models don&#x27;t really match Opus.<p>For example I regularly do Fable+Opus agentic coding runs over 24 hours without intervention.<p>I think I&#x27;ve had GLM do a run that was a few hours. That&#x27;s the closest I&#x27;ve had an open model come on that kind of work.
                • stymaar4 hours ago
                  Even if they don&#x27;t match current-day Opus in everything, they do beat 6 month old Opus, which we have no reason to believe it was smaller than the latest version.
              • WinstonSmith849 hours ago
                Yes. And Opus goes a very long way compared to Fable, Anthropic isn&#x27;t doing any favour, it&#x27;s clearly just 2 models with a very different amount of parameters.
            • johnnyApplePRNG37 minutes ago
              If Fable is seriously around 10T and Kimi K3 sidles up to it at 2.4T<p>That would be extremely surprising and a massive blunder by Anthropic in model design architecture ... which I highly doubt to be the case.
            • andai11 hours ago
              Wasn&#x27;t Opus ~1.5T and Fable is about twice that?
          • riknos31411 hours ago
            Kimi K3 is a 2.8T model that&#x27;s available at about 1&#x2F;4-1&#x2F;3 the cost of Fable from multiple providers on openrouter. The math doesn&#x27;t seem wildly off.
            • zozbot23411 hours ago
              The raw margins on proprietary model inference are rumored to be quite high though (they have to successfully defray the entire investment into model training and datacenter capacity for inference, which is massive enough). The API cost you&#x27;re paying for the model includes that raw margin.
              • alightsoul2 hours ago
                The Chinese have similarly high profit margins on Inference via their first party api
          • x-complexity10 hours ago
            &gt; The cost to train and infer that would be insane, even by today&#x27;s standards.<p>This assumption is likely what has led to the erroneous failure.<p>Enterprise compute per rack has scaled multiple fold in the last 3-5 years. Alongside the training efficiency gains &amp; datacenter scale increases, even 50T+ is well within reach at the top end.
          • kube-system1 hour ago
            Yeah, that&#x27;s insane, you&#x27;d need to have many billions of dollars and buy up a huge chunk of the worlds memory supply to do that &#x2F;s
    • Copenjin8 hours ago
      Didn&#x27;t they say that they can support bigger models now?
      • rbanffy4 hours ago
        Memory capacity on the WSE is the same as before, but access to off-wafer memory is much slower, so the sweet spot is a given fixed balance of memory and compute. They have announced a partnership with AMD in which CPU&#x2F;GPU hardware is used for part of the workload and the WSE-3 machines are used for inference for specialized smaller models, but I&#x27;m not really sure of the details on that.<p>And there is, of course, the educated guesses about what WSE-4 will be, one being adding a LOT of stacked SRAM or DRAM to tip the balance towards memory (which could also be done by having a few different tile designs with various configurations of compute and memory capacity). I am curious about which way they&#x27;ll go.
  • sreekanth85015 hours ago
    AMD along with cerebras may probably compete with NVIDIA monopoly in near future. Also, NVIDIA will have competition form multiple companies. Just my prediction.
    • eitally13 hours ago
      Maybe, but GPU is just one aspect of NVIDIA&#x27;s dominance. If you are buying Vera Rubin GPUs, you&#x27;re getting an NVL72 rack, which is only one of several racks that you&#x27;re probably buying. You&#x27;ll also need your NVIDIA racks with NVIDIA networking &amp; storage gear, too. At the end of the day, they&#x27;re &quot;vertically integrated&quot; for your accelerated computing data center (e.g. the &quot;AI Factory&quot;). This doesn&#x27;t even count the software layer, where CUDA + CUDA-X (not to mention the software for all the sysadmin pieces) has a huge first mover advantage over anyone else.
      • zarzavat12 hours ago
        Is CUDA still a moat? Are we not at the point where frontier models can reimplement software stacks, given you throw enough tokens at the problem.
        • fooker10 hours ago
          If you are making a decision to spend 50B on hardware, would you use the proven tech stack or rely on engineers taking an unspecified amount of time vibecoding your software stack while the hardware sits idle?<p>How about in a month or so when you have to run a slightly different workload?
          • epolanski9 hours ago
            Hyper scalers like Google or Microsoft, which are the big spenders, have all the incentives in the world to get more out of their gargantuan spending.<p>In fact both of them, actually Amazon too, invest in their own inference hardware and owns the stack.<p>You can&#x27;t possibly think that these companies will keep shelling 50-100B per year in hardware alone where 60%+ is <i>margin</i> for Nvidia and not invest there.
        • alightsoul12 hours ago
          if you are developing your own hardware, you provide your own stack to avoid lawsuits with nvidia. i don&#x27;t think it&#x27;s a technical problem at all, but a legal one. this is probably why zluda was scrapped by AMD and Intel. Nvidia technically bans the creation of CUDA reimplementations in their TOS if i remember correctly
          • zarzavat11 hours ago
            Doesn&#x27;t Google v Oracle provide protection here? Copying APIs is fair use.<p>If it&#x27;s patents that are the problem then presumably all these large semiconductor companies have defensive parent portfolios.
            • alightsoul3 hours ago
              I don&#x27;t know then, amd and Intel decided to get rid of Zluda before that lawsuit ruled that apis are fair use. And now, their cloud customers are ok with using their own stack such as rocm, and Intel got rid of their Garuda ai chips. Those things also happened before the lawsuit was settled<p>It would be a breach of contract not a copyright issue.
          • sreekanth85011 hours ago
            [flagged]
        • aurareturn9 hours ago
          You still need experts to know what is good and what is not.
        • someothherguyy6 hours ago
          i haven&#x27;t seen it. see the recent browser attempts.
        • vatsachak12 hours ago
          You just proved that AI cannot currently do that
          • incrudible11 hours ago
            It can definitely create a software stack for you if you hold it right, but the software stack supported by a trillion dollar company with decades of expertise, that <i>also</i> uses AI to improve its stack is <i>probably</i> gonna be better.
            • tomrod4 hours ago
              On the reverse side, there are fundamental limits to the number of ways you can perform certain actions, and agents are both diligent as well as able to swarm. If you have your tests beforehand, there is a chance.
    • anonzzzies10 hours ago
      Like NVIDIA bought Groq, AMD might do well buying Cerebras.
      • FranGro786 hours ago
        AMD did enter into an agreement to buy Taalas, which is speculated [1] will be used to augment their Helios offering.<p>1. <a href="https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=3MKRjt59hh4&amp;pp=0gcJCRMMAYcqIYzv" rel="nofollow">https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=3MKRjt59hh4&amp;pp=0gcJCRMMAYcqI...</a>
      • sreekanth85010 hours ago
        They are Ex AMD employees. Nvidia tried back, but they rejected.
        • vezycash8 hours ago
          The rejection might be overrulled by investors if the offer is enticing enough.
    • _zoltan_9 hours ago
      Press doubt. Single GPU? Maybe. MultiGPU behemoths like NVL144 and NVL576? I don&#x27;t think so.<p>NVLink is at gen9. they had a lot of teething problems and can codesign the hardware and software.<p>in the name of openness (AMD&#x27;s only &quot;&quot;&quot;weapon&quot;&quot;&quot;), the UALink spec is a hodgepodge of corporate opinions with very different implementations (looking at you, Broadcom). at spec version 1 (in hardware).<p>I wish them good luck as I really like AMD, but they compete no more on this than Lambo vs Bugatti.
    • epolanski9 hours ago
      It&#x27;s not really a far fetched prediction: high margins and huge market attract competition, that&#x27;s just the law of economics.
  • xmorse6 hours ago
    Cerebras is very fast but you can basically never use it because of its scarcity
  • reilly300014 hours ago
    &gt; CS-4 delivers more than 1,000 tokens per second on models exceeding 10 trillion parameters<p>Oops did they just out GPT-5.6 sol’s parameter count?
    • sho14 hours ago
      Sol is supposed to be 5T according to rumour. The imminent Astra is allegedly 10
      • nozzlegear13 hours ago
        Rumors and allegations aren&#x27;t worth much. Why don&#x27;t they just tell us mere mortals?
        • WinstonSmith849 hours ago
          Because that would reveal their edge to investors, or the lack thereof.<p>If Fable turns out to be a 10T or 20T model, there is little to boast vs Kimi at 3T. But the opposite is true: if Fable were to be e.g. a 500B model, that would show how far ahead they are from the open models. This isn&#x27;t likely to be the case ...
          • magicalhippo7 hours ago
            I guess there&#x27;s also the economic aspect. It would make it much easier for competitors to figure out your costs and margins if they know the model parameter sizes you operate.
        • brookst12 hours ago
          Why would they? What the upside, for them?
          • eigenspace12 hours ago
            Yeah, its not like this js some sort of <i>Open</i> AI company. That&#x27;d be ridiculous.
            • brookst4 hours ago
              You think their <i>name</i> means releasing competitive details would be good for them?<p>I’ve got sone bad news about Federal Express.
    • whatever114 hours ago
      I mean we kinda know the frontier models are multi trillion parameter models. The only open weights that are close to the frontier are that size too
      • verdverm12 hours ago
        save qwen3.8 27B which is outclassing much larger models and is in spitting distance of the top 10 in <a href="https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;models#intelligence" rel="nofollow">https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;models#intelligence</a>
        • Vax-12 hours ago
          I wonder why they removed DeepSWE from their incorporates evaluations
          • verdverm11 hours ago
            They didn&#x27;t afaict <a href="https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;agents&#x2F;coding-agents?coding-agents-performance-chart=deep-swe" rel="nofollow">https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;agents&#x2F;coding-agents?coding-ag...</a><p>It seems it takes some time to run a new model on all the benchies, not sure they run all models on all of them either
    • kanwisher10 hours ago
      cerebras model are different size then the original models
  • 9cb14c1ec015 hours ago
    Just a reminder for everyone that we are only several years and 3 or 4 iterations into hardware being optimized for LLMs. We should all expect orders of magnitude improvement in speed and&#x2F;or cost over the next 5 years. Then we can have fun conversations about &quot;unlimited&quot; &quot;intelligence&quot; and about what the price wars and profit margins of consumer AI products are when your average ChatGPT user costs the company $0.10 per month.<p>&gt; CS-4 delivers more than 1,000 tokens per second on models exceeding 10 trillion parameters<p>Wow!
    • SwellJoe15 hours ago
      And, the software side isn&#x27;t finished being optimized, either. We&#x27;ve seen with Qwen 3.8 27B and DeepSeek V4 Flash 0731 and GLM 5.3 that quite small models can pack a punch. Intelligence density will improve, efficiency of kernels will improve, efficiency of KV caching and MTP will improve, algorithms for splitting workloads across compute units will improve.<p>It&#x27;ll all be as cheap as DeepSeek was before the price hike. And, it&#x27;ll become more and more realistic to run near-frontier intelligence on personal devices.
    • dgellow8 hours ago
      &gt; Then we can have fun conversations about &quot;unlimited&quot; &quot;intelligence&quot; and about what the price wars and profit margins of consumer AI products are when your average ChatGPT user costs the company $0.10 per month.<p>We can have that discussion now: sounds like that would kill OpenAI and Anthropic
    • brausepulver2 hours ago
      What LLM-specific hardware improvements should one expect? Seems to me that LLM inference is simple architecturally (matmul et al) so most scaling in hardware should come from general improvements (memory BW, packaging, interconnect, power).
      • dorkypunk1 hour ago
        What you describe is basically Cerebras case, at the bottom it&#x27;s just a really big die (about x28 an NVIDIA GB200) with a lot of work to reduce memory latency and improve throughput. What it&#x27;s actually amazing is how can they make a chip so big and still have a decent yield to be commercially viable.
    • moralestapia14 hours ago
      Hence why taalas was one of the best strategic acquisitions of the year.<p>I&#x27;m honestly baffled they were not acquired by somebody else (sorry AMD).
      • adventured12 hours ago
        Taalas will be one of the great disaster investments of the early AI era. It&#x27;ll be a near total write-down.<p>The absolute worst market time to etch a model to a chip is right now (very rapid iteration). There is no scenario where they can keep up. The Taalas approach will be viewed as comically foolish within just a few years.<p>Cerebras will win in terms of approach.<p>It&#x27;s 1998: hey, I can drastically speed up your web service, let&#x27;s etch it right to silicon.
        • Iolaum10 hours ago
          I would still pay ~500 for a chip that runs 10kt&#x2F;s of a ~100b model on my machine even if the half life is 6months. I m sure my employer would too.<p>Qwwen3.5 122b was released 6 months ago and is still best in class overall 100-140 B param model.
          • x-complexity9 hours ago
            &gt; I would still pay ~500 for a chip that runs 10kt&#x2F;s of a ~100b model on my machine even if the half life is 6months.<p>....<p>:T<p>Considering the 8B model uses 53 billion transistors, that&#x27;s 6.625 transistors per parameter.<p><a href="https:&#x2F;&#x2F;taalas.com&#x2F;products&#x2F;" rel="nofollow">https:&#x2F;&#x2F;taalas.com&#x2F;products&#x2F;</a><p>Assuming they can get it down to 3 (somehow), that&#x27;s still 300 transistors, or 5.565 RX 9070s.<p><a href="https:&#x2F;&#x2F;www.techpowerup.com&#x2F;gpu-specs&#x2F;radeon-rx-9070.c4250" rel="nofollow">https:&#x2F;&#x2F;www.techpowerup.com&#x2F;gpu-specs&#x2F;radeon-rx-9070.c4250</a><p>You&#x27;re looking at<p>1) waiting for another 3-5 generations of transistor improvements before it can fit into a single conventional chip, or<p>2) another generation before getting a monster of a chip (1000+ mm^2), and prices for flawless etching scale quadraticly (likely $1000+ for manufacturing costs alone).<p>Could happen, but it&#x27;s a long shot for a market that could be satiated by specialized accelerators.
            • mdp20217 hours ago
              In Taalas HC2 a chip embeds 20b parameters, and the declared idea is linking the chips. A card with two of them chips and you can already have a dense Qwen at staggering speeds.
          • dgellow8 hours ago
            500 what? You’re missing the unit
          • jeffybefffy51910 hours ago
            10000%
        • walrus017 hours ago
          &gt; It&#x27;s 1998: hey, I can drastically speed up your web service, let&#x27;s etch it right to silicon.<p>I distinctly remember 32-bit&#x2F;33 MHz PCI accelerator cards for SSL being a real thing (for use on OpenBSD or FreeBSD), in an era when something like a single core 700 MHz Pentium 3 1U system was a relatively powerful individual bare metal httpd box.<p><a href="http:&#x2F;&#x2F;www.aster.si&#x2F;partnerji&#x2F;compaq&#x2F;atalla&#x2F;axl200.html" rel="nofollow">http:&#x2F;&#x2F;www.aster.si&#x2F;partnerji&#x2F;compaq&#x2F;atalla&#x2F;axl200.html</a><p>The CPU load of doing a lot of SSL purely in software was a problem in terms of scaling things up, so this was one attempt at a (very short lived) solution. Note that this predated TLS1.0.
        • NitpickLawyer11 hours ago
          &gt; The absolute worst market time to etch a model to a chip is right now<p>Slightly disagree. It really depends on the price-point at which they can do that etching. ~1k usd &#x2F; ~30B model in a hdd-sized case that fits on your desk? I&#x27;d buy one right now, even knowing that I&#x27;m &quot;stuck&quot; with whatever model of the day is.
          • riknos31411 hours ago
            Time to market also matters a ton. If they can start shipping chips &lt;1 month after the weights drop that&#x27;s much more compelling than if it&#x27;s a 6+ month development pipeline.
            • mdp20217 hours ago
              In the case of Taalas, the pipeline was said to be 2 months:<p>&gt; <i>From the moment a previously unseen model is received, it can be realized in hardware in only two months</i> ( <a href="https:&#x2F;&#x2F;taalas.com&#x2F;the-path-to-ubiquitous-ai&#x2F;" rel="nofollow">https:&#x2F;&#x2F;taalas.com&#x2F;the-path-to-ubiquitous-ai&#x2F;</a> )
        • WithinReason9 hours ago
          The 500x efficiency gain makes their approach a no brainer. Just make a new chip every 6 months, you still win.
        • alightsoul12 hours ago
          maybe AMD wants the IP to deploy it once ai model development slows down in a few years. Or, their large cloud customers do want to burn through silicon, basically paying rent to AMD for models etched on silicon.
        • pimeys11 hours ago
          I just want to but hardware so I can run a model at home that is fast. I don&#x27;t see myself installing a server that burns almost two hundred kilowatts but maybe a card which runs a 27B Qwen...
          • mdp20217 hours ago
            At 250w <i>when it&#x27;s working</i> (I understand), and it works for tiny amounts of time per query...
        • moralestapia12 hours ago
          It&#x27;s 2026: let&#x27;s etch nginx into silicon and get 10,000,000 rps at a cost of 0.1 US&#x2F;day.<p>Yes, please!
          • redox9911 hours ago
            Most people probably don&#x27;t care about nginx performance. It shouldn&#x27;t be your bottleneck unless you serve massive amounts of static data.
            • mdp20216 hours ago
              In the case of needs to process natural language, instead, massive efficiency (esp. time) can be a game changer. It&#x27;s like &quot;you have two years to complete the project&quot; vs &quot;you have two hours to complete the project&quot;: if you can squeeze that &quot;two years worth&quot; into a negligible delay, it&#x27;s a game changer.
            • dyzone9 hours ago
              Ok, how about postgres?
        • conception10 hours ago
          Etched model into a chip? A… mobile chip eventually? Seems prescient.
    • api15 hours ago
      This is part of why I think the data center build-out is a bubble. We&#x27;ve barely scratched the surface when it comes to hardware optimization. We&#x27;ll see exponential improvements in energy efficiency and speed over the next decade. Exponential, not linear.<p>GPUs really aren&#x27;t that great for AI. They just happen to be the best chips we have in mass production right now for this work load, and it takes time to field new designs. Basically every chip engineer on the planet is working on this right now.
      • mindwok15 hours ago
        Whether it&#x27;s a bubble or not depends on how much the demand for compute and the type of workload keeps growing, though.<p>If AI tends to be something used mainly in ideation and development, which is how a lot of people use it today, then once consumer hardware gets good enough you could see a bunch of the current data centre workloads move onto consumer devices.<p>But if AI starts being used more in repeatable, operational workloads I think it makes sense to have significant cloud infrastructure for it. TBH I haven&#x27;t seen much of this, and I&#x27;ve been skeptical about people using agents for much of anything when it can be done with just software. But we are starting to see more of this kind of workload, like the taggable Claude in your slack etc that people seem to really love.
      • aurareturn9 hours ago
        By the way, this is the same argument that Michael Burry used to short Nvidia.<p>He claims that GPU depreciation&#x2F;obsoletion is much faster than hyperscalers are assuming because new chips will be much better. He&#x27;s being proved wrong right now because H200 rental prices have been claiming for the last 8 month despite B200 having 10-20x better inference efficiency.[0]<p>The logic is fundamentally flawed in my opinion. Let&#x27;s use future Nvidia chips being much better optimized for LLMs for example.<p>New Nvidia chips 10x better than H200 --&gt; data centers buy a lot --&gt; Nvidia profits a lot.<p>New Nvidia chips 10x better than H200 --&gt; data centers don&#x27;t buy --&gt; no faster than expected obsoletion.<p>In other words, the very act of buying many new Nvidia GPUs would be the event that causes faster than expected obsoletion. Yet, if you don&#x27;t buy those new Nvidia GPUs, then there is no faster than expected obsoletion.<p>We also live in a world where there is competition. If Amazon doesn&#x27;t buy but Microsoft does, suddenly Microsoft can offer better $&#x2F;token prices.<p>[0]<a href="https:&#x2F;&#x2F;inferencex.semianalysis.com&#x2F;inference" rel="nofollow">https:&#x2F;&#x2F;inferencex.semianalysis.com&#x2F;inference</a>
        • haldujai7 hours ago
          1. The same isn’t necessarily true of the rest of the hardware stack which may be reused between accelerator generations.<p>2. You’re missing the “New Nvidia chips 10x B200, compute requirement grows less than 10*software improvements YoY -&gt; buy less Nvidia.” Valuations are based on forward projections (&gt;1T annual for NVDA) which can be revised down leading to a drop in valuation.<p>&gt; If Amazon doesn&#x27;t buy but Microsoft does<p>The big 3 all have their own proprietary accelerators. Meta is buying TPUs as well for now.<p>I would bet Nvidia’s major customers in 2 years are neoclouds and it seems that Jensen is making the same bet.
          • aurareturn5 hours ago
            1. So this makes Burry’s argument even less convincing since those auxiliary hardware can last longer.<p>2. Jevons Paradox. More efficiency should lead to bigger models, faster inference, and more total tokens.<p>3. By all accounts, Trainium and Maia and Meta’s internal chip are struggling to keep up with Nvidia. That’s why they order as many Nvidia chips as possible. They’re not giving up but it isn’t as easy as buying stock Arm cores and taking them to TSMC.<p>Neoclouds may very well be Nvidia’s biggest customers and this probably what Nvidia wants.
            • haldujai34 minutes ago
              1. Not really, current valuations are priced for persistent 80%+ margins based on spot. If auxiliary hardware lasts longer (I.e. next gen GPU reusing the same shell) then that reduces supply pressure and spot prices.<p>2. Jevon’s paradox is about total consumption, not margins. Valuations are about margins (and their projections). Many coal mine owners went bust despite increased total coal consumption.<p>3. Source? Gemini for example is 70% on TPU. I have yet to see data on Maia-300 beyond Microsoft PR. Remember it doesn’t have to be <i>better</i> it has to be more cost efficient. The overwhelming majority of inference spend does not care if token output is 20% slower if it is 50% cheaper.<p>&gt; Neoclouds may very well be Nvidia’s biggest customers and this probably what Nvidia wants.<p>What Nvidia <i>needs</i>. Whether neoclouds can stay competitive vs hyperscalers paying Nvidia tax is far from clear, particularly when inference margins compress.
      • petra14 hours ago
        I wonder: in world where inference is cheap, how many engineering agents that use simulation as their feedback we will use?<p>In the scenario, engineering everything becomes so easy - so why not optimize everything? every component, every product, every system?<p>And maybe llm&#x27;s could invent. So even more to simulate. And simulation is inherently compute-heavy.<p>So unless there are some other bottlenecks, we&#x27;ll use a lot of simulation servers.
      • winrid15 hours ago
        On the plus side, lots of cheap servers to swoop up :)
        • sroussey15 hours ago
          But power hungry.<p>In that 5+ year timeline, the compute per watt could change by three orders of magnitude.<p>GPUs are to LLMs what CPUs are to gaming — not a good fit.
          • amluto13 hours ago
            A cursory estimate courtesy of ChatGPT suggests that there is a grand total of one order of magnitude or less of power efficiency improvement available compared to current Blackwell if the entire system’s power consumption outside the ALUs went all the way to zero.<p>If you want three orders of magnitude improvement, you probably need to find two of those orders of magnitude somewhere else: process improvements, different ALU design, model architecture changes, etc.
        • dgellow8 hours ago
          Look at their power supply, it’s not something you can run in a home lab. Unfortunately most of that will likely go to the bin eventually :(
      • RachelF13 hours ago
        True, I have to agree with you. The AI giants might be investing a huge amount of money in generation 1 technology. There might be a much better way to do it just around the corner. They might know this and thus the hurry to IPO.<p>A rough analogy would be if the first generation of ISP&#x27;s spent billions on dial-up exchanges, when fibre could be invented next year.
      • jeffybefffy51910 hours ago
        Exactly right, and nVidia is protecting their moat through business practices rather than genuine product innovation.
      • __turbobrew__13 hours ago
        By the time these gigawatt datacenters are done being built the hardware will be so far behind state of the art they may be mostly useless.
    • rvz15 hours ago
      Congratulations! You have just realized that the AI data center build out is a total scam, built on both the insurmountable trillions of debt, and the assumption that <i>only</i> GPUs are all we need to continue scaling.<p>There exist other AI accelerators (TPUs, ASICs) that perfectly exceed the throughput that LLMs need to scale as well. But the true solution is more software optimizations. There&#x27;s a tiny handful of them but more needs to be discovered so that we can reduce building hundreds of more data centers as the alternatives mature.<p>As better software becomes more useful for the alternative AI hardware for developers with LLMs running efficiently you then would have more choices of hardware to run your LLMs on rather than just only GPUs.
      • blovescoffee14 hours ago
        TPUs and ASICs run in data centers too. Your argument only holds true if there&#x27;s some satisfied limit to demand for inference. If not, data centers will continue to spring up to host more and more agents. Even if agents were running on hardware and software as efficient as the human brain, its conceivable we want trillions of them running at any given time which would require data center scale.
        • georgeecollins14 hours ago
          Everything has some satisfied limit to demand, often depending on the price. If you assume there will never be any satisfied limit to demand for inference at any price you can justify any investment.
          • aldonius14 hours ago
            Yeah, but there&#x27;s certainly a part of the curve where price drops by X OOMs and demand increases by much more than X OOMs. (Presumably some of that is substitution and some of that is new use cases.)
          • x-complexity9 hours ago
            So far, at least by Openrouter&#x27;s weekly numbers, there doesn&#x27;t seem to be a satisfied limit.<p><a href="https:&#x2F;&#x2F;openrouter.ai&#x2F;rankings#top-models" rel="nofollow">https:&#x2F;&#x2F;openrouter.ai&#x2F;rankings#top-models</a><p>And their market share sits at around 16-20%.<p>At 75.3 trillion tokens for the week ending 10 Aug 2026, that means that up to 450 trillion tokens were plausibly demanded by the whole market for that week.<p>My take: At max saturation, each person on earth could have their demands satiated by an average of 16 agents running concurrently. Sometimes more, often times less, but the average would likely be at 16.<p>At 200 tokens&#x2F;second for each agent, that would mean 15.48288 quintillion tokens per week.<p>We&#x27;re currently at about 0.00290643601% of the calculated demand ceiling.<p>Even if the demand limit per person is just 1 agent at 50 tokens&#x2F;second, the current demand&#x27;s still 0.186011905% of the theoretical ceiling.
          • adventured12 hours ago
            Looking back nearly 80 years, what has been the limit to transistor demand so far?<p>Unlimited.<p>What has been the limit to electricity demand globally?<p>Unlimited.<p>We can&#x27;t get enough and never will. Costs have to become pretty severe to turn back the demand as well.
      • skyberrys14 hours ago
        I wonder what this looks like in 5 years... Will there be a massive push to repurpose these giant boxes into housing? Will they get turned back into the farm land from where they came? When a data center goes bust, what happens to the parts left behind?
        • gpm14 hours ago
          I&#x27;d think the infrastructure would tend towards factories, smelters, and so on. Industrial things that have reasonably high power demands, can use the building, and don&#x27;t care about the lack of windows.<p>They&#x27;re typically not built where you want housing, and the buildings are distinctly the wrong shape.<p>If you can&#x27;t use the power infrastructure profitably my next thought would be warehousing.<p>But also... we&#x27;ve seen a pretty continually increasing demand for compute. Even if AI busts a bit (or becomes a bit more efficient) I bet most data centres stay data centres, just less profitable ones.
      • jryle7014 hours ago
        Huh, why I&#x27;m not surprised that HN is full of opinions confidently stated without any numbers or resources to back up?<p>&gt; built on both the insurmountable trillions of debt, and the assumption that only GPUs are all we need to continue scaling.<p>Insurmountable according to whom? And who assume that only GPUs are all we need to continue scaling? Google, Amazon, Microsoft, Meta and OpenAI, all have or plan custom non-GPU AI chips. Do they plan to use them not for scaling?
  • rajnathani7 hours ago
    Interestingly they’re still on the WSE-3 (5nm TSMC) wafer chip and slightly bumped up the specs there (overlocking mostly it seems), for why it’s called WSE-3 Turbo now. I think people were also expecting WSE-4, as it’s been 2 years now since WSE-3 was launched.
  • ethanzhang102414 hours ago
    If cerebars is performing well, why didn&#x27;t its predecessor, server S-3, become the largest API token provider on OpenRouter, surpassing the official model releases?
    • walrus0114 hours ago
      Without having any inside information, one possible theory:<p>All or a vast majority of of the cerebras manufacturing capacity was going to a few companies that aren&#x27;t publicly available inference providers on openrouter, for their own internal use.<p>or<p>The asking price of the S-3, no matter how speedy it might be, for small&#x2F;medium size customers made it economically prohibitive to purchase and use to sell public inference vs. buying more common nvidia b200 or whatever.
    • wmf12 hours ago
      Cerebras provides high-speed inference at high cost. It&#x27;s never going to be the cheapest and thus it will probably remain niche.
      • zurfer8 hours ago
        But that&#x27;s supply and demand, not technology. Right now a lot more people want their inference than they can supply. as supply catches up in the next 5-10 years, the underlying tech at scale is probably cheaper than GPUs per token produced.
    • aurareturn9 hours ago
      Probably the same reason why there are more people who takes buses, subways, trains than drive Ferraris.
      • epolanski7 hours ago
        I love how instead of comparing a Ferrari (fast and expensive) to some average car (not fast, not expensive) to make your point..you went for public transport where your comparison cracks from multiple angles.
    • smallerize14 hours ago
      If you&#x27;re willing to pay a significant premium for latency, why use openrouter? And anyway Cerebras only supported a few specific models.
    • gampleman8 hours ago
      Cerebras capacity was pretty much entirely bought out at some point. We needed it and couldn&#x27;t get it.
    • aseipp14 hours ago
      The WSE is very expensive to build, and they have a waiting list of customers who are already willing to pay a lot of money for the available supply.
    • doctorpangloss14 hours ago
      it only takes ~445 GB300 NVL72 (about $22b) to run ALL of openrouter demand for a year. Microsoft rolled out $32b of DC 2026Q1.<p>imo the issue is that most openrouter demand is inauthentic activity (things that anthropic and openai models will refuse to do like pretend to not be bots when interacting with humans)
      • aurareturn9 hours ago
        I thought your numbers must be wrong.<p>So I plugged 288 trillion tokens&#x2F;month (OpenRouter&#x27;s current rate), 500 billion MoE model average, and the math comes out to be around 620 B200 GPUs minimum.<p>So basically, OpenRouter&#x27;s volume must be absolutely tiny compared to the volume hyperscalers are getting.
        • Mattwmaster583 hours ago
          For reference, Google serves &gt;3 quadrillion&#x2F;mo.<p>[1] <a href="https:&#x2F;&#x2F;x.com&#x2F;ren_stocks&#x2F;status&#x2F;2056946641815396718?s=20" rel="nofollow">https:&#x2F;&#x2F;x.com&#x2F;ren_stocks&#x2F;status&#x2F;2056946641815396718?s=20</a>
          • aurareturn2 hours ago
            Which is only 10x more than Open Router by the way.
      • HDBaseT13 hours ago
        It is worth mentioning, the OpenRouter demand isn&#x27;t static though. It has increased week on week since early 2026.
      • senordevnyc13 hours ago
        I was curious so I looked it up: looks like a GB300 NVL72 is about $4M. So $22B would buy you 5500 such racks, no?
        • selectodude9 hours ago
          GPU cost is about half of datacenter cost. Other half is cooling, power and networking.
    • 0xbadcafebee1 hour ago
      Because they aren&#x27;t selling inference, they&#x27;re selling hardware. The only reason they sell any tokens on OpenRouter is so they get on the benchmark that shows them as the fastest provider. It&#x27;s free advertising.
    • petesergeant7 hours ago
      I use Cerebras via OpenRouter. It’s every bit as fast and reliable for my needs as claimed. I suspect the reason is that they can either be making peanuts selling inference to plebs like me via OpenRouter, or making bank selling the more expensive models to businesses directly. In short: I would be very surprised if they have die capacity, and are at this point maximising revenue per chip.
  • rbanffy5 hours ago
    Impressive that this is an &quot;interim&quot; product, the start of a new line that ought to be continued with the WSE-4 family, where they are supposed to use a 3nm process and, maybe, 3D stacked SRAM. The modular architecture also points towards field upgrades that are badly needed for AI datacenter builders.
  • anonymous_user915 hours ago
    Conspicuously missing: power consumption figures
    • wmf15 hours ago
      162 kW
      • walrus0114 hours ago
        I guess we know why there&#x27;s a fair bit of investment money going into small modular nuclear reactor startups now.
        • fragmede14 hours ago
          And advanced geothermal. Fervo Energy let&#x27;s us get energy that&#x27;s not based on burning fossil fuels but is, instead, able to produce energy from the ground.
          • ricardobeat5 hours ago
            <i>Remove</i> energy from the ground. I wonder what the consequences may be once we are cooling the underground at several MW&#x2F;h.
            • Isaackoz1 hour ago
              Earth will be swallowed by the sun before humans could make a dent in the earths core temperature.
              • ricardobeat58 minutes ago
                The core is obviously not in question here, but at scale and long term it could have an impact on the water bed, vegetation, underground ecosystems, soil stability, etc.
        • dakolli8 hours ago
          Actually that&#x27;s mostly just the military funding those, with a few of them having data center partnerships so they can shield themselves from the criticism of what they really are: military contractors.
          • walrus018 hours ago
            I heard a great deal of noise from early 2002 to the present date that the US military had a high interest in small portable nuclear reactors for large bases in Iraq and Afghanistan. And particularly around the peak period of troops on the ground in AF and IQ. And indeed a place like Bagram or Kandahar used a shitton of diesel to run generators. But nothing ever came to fruition to actually implement it, has something changed now that they actually consider it worth doing?
      • roughly14 hours ago
        God, I was going to ask if this could be deployed in a standard existing datacenter, but I guess that answers that question.
        • KeplerBoy11 hours ago
          So the answer is yes? Putting in a few of those racks for special tasks shouldn&#x27;t break the power assumptions of a data center.
          • roughly27 minutes ago
            Distribution&#x27;s always the problem - how much power actually gets delivered to each rack. 162 is ~an order of magnitude higher than normal, which means nothing in a standard data center is going to be built to deliver that kind of power to one rack.
      • xattt15 hours ago
        I presume per rack?<p>Can you imagine something radiating that much energy into a space in your home?
        • walrus0114 hours ago
          It&#x27;s mandatory liquid cooling, so it&#x27;s meant to be attached to a specialized liquid cooling loop that gets the heat outside the building.<p>This is far beyond the practical maximums of like 10 to 15kW per 44U cabinet front to rear air cooling for &#x27;regular&#x27; rackmount server stuff.
          • ttul14 hours ago
            Indeed. You need 45 to 60 liters per second of cooling water flowing over a Cerebras wafer every minute to keep it under 90C. And that’s assuming the water leaves at 90C…<p>More realistically, you need much more cooling water.
            • 0xbadcafebee1 hour ago
              CS-3 used 100 liters of water with a cold plate (<a href="https:&#x2F;&#x2F;www.brownstoneresearch.com&#x2F;bleeding-edge&#x2F;ai-infrastructure-investment-continues-to-increase&#x2F;" rel="nofollow">https:&#x2F;&#x2F;www.brownstoneresearch.com&#x2F;bleeding-edge&#x2F;ai-infrastr...</a>). But you can also use refrigerants with a cold plate (<a href="https:&#x2F;&#x2F;eng.umd.edu&#x2F;engineering-ai-public-good&#x2F;cooling-data-centers-safer-two-phase-fluids" rel="nofollow">https:&#x2F;&#x2F;eng.umd.edu&#x2F;engineering-ai-public-good&#x2F;cooling-data-...</a>) or dielectric fluid in total immersion. Until they find a more power-efficient design, my guess is total immersion will come back in style.
        • wmf15 hours ago
          I guess because I have actually set foot in a data center I don&#x27;t imagine literally every product in my home.
    • logicallee14 hours ago
      &quot;10x more throughput per watt than CS-3&quot;
  • anarticle24 minutes ago
    Really wish they’d host more models for us normies. My guess is OAI will buy &#x2F; subsidize them with terms that will close off open models.
  • aneryu14 hours ago
    It would be even better if a version available to individual users were released soon.
    • gpm14 hours ago
      They do offer API services to individual users... though with a set of models that makes it unlikely that you want to use it. They are promising Qwen 3.8 27B any day now though*.<p>if you have the money as an &quot;individual user&quot; to purchase one of their racks... save your money and retire.<p>* Actually they sent out an email claiming they already have it, but I don&#x27;t seem to have access, they&#x27;re promising to release it to the &quot;shared tier&quot; any day now.
      • drcode12 hours ago
        not only that, but I was so happy with their GLM 4.8 that they got rid of yesterday :(
        • apitman12 hours ago
          What do they do with old ones? Their hardware physically can&#x27;t run other models right?
          • airspresso12 hours ago
            The Cerebras hardware is not locked to specific models &#x2F; model families. Taalas is the company that&#x27;s etching models into their silicon, locking it to that model forever.
            • apitman4 hours ago
              That&#x27;s pretty sweet. Thanks
      • fragmede14 hours ago
        &gt; save your money and retire.<p>Now that this hypothetical person has retired, what are they gonna do all day? Just sit on the beach and drink Mai Tais? If that&#x27;s what they wanna do, sure, but nerds gonna nerd, and if I had that kind of money to retire on, I&#x27;d totally buy some ridiculously expensive AI box for fun.
        • gpm14 hours ago
          Ah but if you have the kind of money where this is a reasonable retirement hobby purchase, you aren&#x27;t bothered by representing yourself as a &quot;enterprise&quot; :P
    • ttul14 hours ago
      I’ll get that 250kW home power service dropped in next week!
      • z213 hours ago
        For now you can rig an adapter to your nearest DC EV charging station, but make sure it&#x27;s near a body of water for the cooling.
        • ttul1 hour ago
          Hot water for the whole neighbourhood!
    • dnautics12 hours ago
      needs an sla that says power will never ever ever go out or else you will have a useless shattered plate of silicon.
      • gpm11 hours ago
        Huh, why would it shatter if the power goes out?
        • dnautics1 hour ago
          by my understanding:<p>there is an extensive and complicated cooling system that permeates the wafer. some of the cores are completely turned off because they fail qc (see the tsmc logo) if all of their neighbors have been going at full bore the thermal differential can cause stress fractures if the cooling system suddenly fails.
    • johntash10 hours ago
      I&#x27;d like to see a consumer version too, I don&#x27;t need a whole rack of them. I probably can&#x27;t even afford one gpu-sized one
  • kobe_bryant13 hours ago
    can these vibe coded sites please set a max width and overflow so their sites work fine on mobile
    • dgellow8 hours ago
      Good news, future models will have your comment in their training set, making them slightly more likely to fix that problem!
  • 4k0hz15 hours ago
    &gt; Introducing the all new Cerebras CS-4, a revolutionary rack-scale solution that delivers upto 30x faster inference compared to GPUs, enhanced economics, and a simple path todeploy [sic] hyperscale capacity.<p>Did nobody proofread this?
    • algoth115 hours ago
      If they had ask Claude it would probably look like this: Introducing the all new Cerebras CS-4, a revolutionary rack-scale solution that delivers up to 30x faster inference compared to GPUs, enhanced economics, and a simple path to load-bearing hyper scale capacity.
      • jm415 hours ago
        That&#x27;s unusually honest and the sharpest thing in this thread.
        • wren699114 hours ago
          You&#x27;re underselling it, and here&#x27;s why.
    • SoMomentary15 hours ago
      Sometimes I wonder if mistakes are now used to indicate the possibility that a human actually wrote it.
      • ceejayoz14 hours ago
        There’s been a spate of Reddit AI bots using all lower case in hopes of evading detection.<p>It’s still incredibly obvious.
        • VladVladikoff13 hours ago
          I don’t really visit Reddit much these days but would love to see an example.
          • ceejayoz3 hours ago
            <a href="https:&#x2F;&#x2F;www.reddit.com&#x2F;r&#x2F;ModSupport&#x2F;comments&#x2F;1tu67kn&#x2F;influx_of_bot_comments_with_gen_z_wording_anybody&#x2F;" rel="nofollow">https:&#x2F;&#x2F;www.reddit.com&#x2F;r&#x2F;ModSupport&#x2F;comments&#x2F;1tu67kn&#x2F;influx_...</a>
    • helloplanets10 hours ago
      Well at lest it&#x27;s written by a human.
      • dgellow8 hours ago
        « Make it look like human written »
    • geodel15 hours ago
      Maybe it is just part of their &quot;compact design&quot;.
    • dpkirchner15 hours ago
      An error no frontier LLM would make, eh
  • arthurcolle10 hours ago
    What&#x27;s the sticker price? If I have 20 million in the bank can I just like buy one or what
    • runeks8 hours ago
      Surely it would depend on your model, since they need to etch this into silicon.
      • yvdriess7 hours ago
        You&#x27;re thinking of Taalas.
  • lostmsu14 hours ago
    KV caching status?<p>What&#x27;s the point of 1000tok&#x2F;s if you have to do prefill on every agentic turn which at 100k depth would make it 1.5 min latency every turn?
    • walrus0114 hours ago
      Information about RAM type&#x2F;size and connection topology of the RAM to be used for context cache seems to be conspicuously absent from the slick looking marketing materials.
      • gpm14 hours ago
        There&#x27;s a few more details at the bottom of this page: <a href="https:&#x2F;&#x2F;www.cerebras.ai&#x2F;blog&#x2F;introducing-cerebras-cs-4" rel="nofollow">https:&#x2F;&#x2F;www.cerebras.ai&#x2F;blog&#x2F;introducing-cerebras-cs-4</a><p>44GB on-chip-sram * 3 chips. Per chip: 43.2 PB&#x2F;s memory access + 53.5 PB&#x2F;s on-chip fabric bandwidth + 2.4 Tbits&#x2F;s &quot;IO&quot; bandwidth (I think that means their RoCE v2 RDMA over Ethernet interface).<p>I suspect there might be a certain amount of customization for how much RAM they attach when you order it.
        • porridgeraisin12 hours ago
          They have managed to make the external link 300GBps&#x2F;2us. Cs3 was 150&#x2F;5.<p>This is 1&#x2F;3rd blackwells nvlink c2c bandwidth already. Not too bad. We can make KV cache offload work with that I suppose.<p>If magically KV cache was not an issue, pipeline parallelism on cerebras can be quite pleasant. As for the KV cache offload, I have hopes their CPO solution they&#x27;re trying with that canadian company ends up bearing fruit.
  • denizay13 hours ago
    The comparison seems incomplete. CS‑4 is a full rack-scale system with three wafer-scale processors, but the exact GPU models, GPU count, power consumption, price information are not disclosed. We still don&#x27;t know if buying a multi-GPU rack (or racks) is cheaper and&#x2F;or more efficient in power. The fact that they didn&#x27;t disclose these numbers makes me believe that the numbers are not in their favor. And personally, makes me see them as disingenuous.
  • sva_14 hours ago
    &gt; enabling massive clusters and models with more than 50 trillion parameters
  • charlielidbury4 hours ago
    only 44GB * 3 of VRAM per rack :O<p>I guess you&#x27;d need a DOZEN(s) of these to host a large model with long context KV caches?
  • selimonder9 hours ago
    That &quot;GPU&quot; comparison is the vaguest i seen so far
    • epolanski9 hours ago
      True, it&#x27;s also &quot;per user&quot;, somehow, but I think it&#x27;s a misleading metric. Cerebras chips take the whole wafer?<p>A single TSMC wafer contains 60 to 65 B200s, assuming 70% yields that&#x27;s 40ish wafers per die.<p>Cerebras cannot redefine wafer economics.
      • tjoff8 hours ago
        Depends on what you mean, they have more redundancy which means that the yield can be much higher.
  • avantnyc13 hours ago
    Cerebras should slowly also move to dgx&#x2F;ryzen market for a desktop version for masses at affordable price yet providing substantial tokens&#x2F;second on desktop
    • wmf12 hours ago
      Desktop SRAM isn&#x27;t really viable because it could cost $100K just to load the model.
      • aenis10 hours ago
        ...so, in the same ballapark as ddr5? :-)
  • tamimio14 hours ago
    I wonder what are the benchmarks of hashcat on different hashes.
  • OutOfHere15 hours ago
    Five years from now, I don&#x27;t know why anyone will still be using Nvidia for inference. Note that Cerebras is for inference only, not for training.<p>I understand that Cerebras has competition, but this bodes even more poorly for Nvidia for inference. Nvidia may still have a role to play for training, however.
    • dgellow8 hours ago
      NVIDIA has the best supply chain in the entire game. They are the only ones who can produce at their scale. You really shouldn’t underestimate their position
    • kcb12 hours ago
      Nvidia is at this time a pretty well run company tech wise. They are going to keep iterating on the inferencing hardware stack over the next five years too.
      • OutOfHere6 hours ago
        The only way I see in which Nvidia can catch up is by buying Cerebras.
    • wmf15 hours ago
      Cerebras is only claiming ~2x the performance of Groqvidia which usually isn&#x27;t enough for people to switch.
  • adventured12 hours ago
    OpenAI needs to immediately move to acquire Cerebras.<p>Nvidia&#x27;s extreme margin is the opportunity for OpenAI&#x27;s cost reduction. Buying Cerebras would pay for itself and they should take all of its future production (after filling required contracts).<p>Right now China&#x27;s models have no silicon moat. Cerebras as a drastic speed-up &#x2F; cost-reduction potential, can assist in building a competitive moat. And every time a Cerebras pops up, OpenAI or Anthropic should eat them if at all possible.<p>There&#x27;s no stand-alone frontier AI company of great scale in the near future that doesn&#x27;t have a large silicon advantage in-house. Apple knew it in smartphones, Google figured it out a long time ago as well.
    • thefounder12 hours ago
      With what? More debt? What will nvidia say?
    • wmf12 hours ago
      Do you know about Jalapeno?
      • mmmeff11 hours ago
        ^<p>OpenAI is partnering with Cerebras while simultaneously investing in their own silicon play. Hedged bets.<p>After sitting thru their keynote today, it makes sense. The main throughput speedups they tout are an obvious evolution of the GPU that all companies will be building in the next year. Wafer-scale interconnected memory and compute is just going to beat out mountains of network cabling any day on both cost and performance metrics.
  • gpm15 hours ago
    Is it just me or is it bizarre that they&#x27;re advertising old open-weight models.<p>GLM 4.7 (December 2025) not 5 (Feb) 5.1 (April) or 5.2 (June). 5.3 (4 days ago) is, to be fair, not open weights yet... but there&#x27;s a lot since 4.7.<p>Kimi K2.7 (April) not K2.7-code (June) or K3 (July).<p>Gemma 4 (April), Llama (April), and gpt-oss (August 2025) are up to date, but old (for models).<p>Meanwhile the closed source GPT 5.6 sol is up to date (June)...<p>Should potential purchasers take away from this that they&#x27;re not going to be able to run recent models unless they front the cost of developing software or something?
    • eli15 hours ago
      I think they run whatever models they get paid to run. But mostly from enterprise. They are clearly not interested in consumer dollars.
      • gpm15 hours ago
        I mean the product is a server rack and while there&#x27;s no advertised price I would assume it&#x27;s six figures. So yes, an enterprise product.<p>But even an enterprise is going to care about the difference between &quot;we can run the model we want with support from the manufacturer&quot; and &quot;we have to purchase the product, and then spend another 6 figure sum having developers port a recent model to the product to use it&quot;.
        • kube-system13 hours ago
          You’re at least an order or magnitude under… likely two.<p>A single AI server with a mere 8 GPUs from Nvidia is already mid 6 digits. A rack system from Nvidia is mid 7 digits.<p>There’s some info out there that suggests the CS1 had an 8 digits price tag, so it wouldn’t be surprising to see that here.
          • gpm13 hours ago
            Yeah, did some googling after WarmWash&#x27;s comment and I concur.
        • WarmWash15 hours ago
          I feel like 6-figures would be the clearance price on it...
          • oceanplexian14 hours ago
            6 figures is a single mid range Xeon or Epyc server these days.
  • jaumesnts3 hours ago
    [dead]