26 comments

  • simonw4 hours ago
    The numbers I was most interested in are tucked away in a chart towards the bottom - the speed comparison of the Mac Studios v.s. a RTX 5090:<p><pre><code> Qwen3.8 27B tokens&#x2F;sec generation speed Prompt size 8K 64K 128K 256K RTX 5090 PC 59 51 44 n&#x2F;a M5 Ultra 48 39 32 24 M3 Ultra 31 23.5 20 15 </code></pre> A whole bunch more comparison numbers in this section: <a href="https:&#x2F;&#x2F;www.macstories.net&#x2F;stories&#x2F;m5-ultra-mac-studio-review-the-dream-mac-for-local-ai-agents&#x2F;#mx-pc" rel="nofollow">https:&#x2F;&#x2F;www.macstories.net&#x2F;stories&#x2F;m5-ultra-mac-studio-revie...</a>
    • gpugreg3 hours ago
      Those RTX 5090 numbers are bad. You can get over 200 tps with ninfer using NVFP4 and MTP.
      • beastman822 hours ago
        can confirm.<p>I dont&#x27; know why people spend huge money on these and Spark. The 5090 is running qwen 3.8 at 200+ tps!! That&#x27;s 1-2 orders of magnitude faster.
        • _hugerobots_2 hours ago
          Have a 5090, and yes it&#x27;s very fast. But it&#x27;s like the worst ADHD team member and requires constant supervision and review from larger models. It&#x27;s context size on-card is good for super, suuuuuper shallow precision work. The gb10&#x2F;spark on top of it, that thing can refactor enormous monorepo architecture. The time it takes the 5090 to compact, reiterate and execute a plan is often the same time as the gb10.
        • tomega21341 hour ago
          Is a 5090 still cost efficent when it is (currently) unobtainable? Or when obtainable only at current prices (min. $6500 USD)?
        • cyanydeez4 minutes ago
          Qwen3.8-Flash-Next is pretty damn worth the extra ram you need.
        • nacs2 hours ago
          People don&#x27;t buy Sparks and M5 Ultras to run a 27B model - you buy it to run an MoE model like Qwen Next which this M5 excelled at.
          • ProllyInfamous59 minutes ago
            Exactly; when I first got my RTX 5070 Ti (16gb, to <i>game with!!!</i>, upgrading from VEGA56), I loaded then-latest Qwen3.6 (~30B, cannot remember exactly). My only prior LLM experience was with models &lt;8gb, primarily llama3.1.<p>My technical-expert twin played around with these LLMs, for about an hour, and then <i>correctly reasoned &quot;it&#x27;s able to be WRONG, faster.&quot;</i><p>This seems apt. My next LLM machine will be closer to 96gb+ vRAM.
            • selectodude20 minutes ago
              Once I get some kind of settlement after getting beaten up by a cop my first purchase will be some RTX Pro 6000s.
        • throwaway274482 hours ago
          A) the macos value add is enormous if you have any investment in the ecosystem, B) for me at least a GPU is completely useless for anything but being a token generator.
          • bigyabai2 hours ago
            &gt; for me at least a GPU is completely useless for anything but being a token generator.<p>No thanks to the &quot;macos value add&quot; that forces you to use Metal while Valve customers frolick in Protonland.
            • throwaway274481 hour ago
              &gt; No thanks to the &quot;macos value add&quot; that forces you to use Metal while Valve customers frolick in Protonland.<p>Crossover works on macos, too. So does moltenvk, so does vanilla wine, etc etc. You can run most games without a hitch these days (allegedly, according to &#x2F;r&#x2F;macgaming). But I don&#x27;t play video games so a GPU would probably be better off in some kid&#x27;s computer.
        • Eisenstein2 hours ago
          A 5090 has a 1.79TB&#x2F;s memory bandwidth. Qwen 3.8 27B NVFP4 is 22GB. You cannot generate tokens faster than the weights can traverse the GPU memory, so that makes max generation speed without MTP to be 81T&#x2F;s. Say MTP is giving you 0.5 acceptance rate (very good), that is 1.5 * 81 is 121T&#x2F;s. Even with a perfect acceptance rate you would only get 162T&#x2F;s.
          • beastman822 hours ago
            Off the top of my head, I&#x27;m guessing we&#x27;re missing sparse attention. But I&#x27;ll run your challenge through and see where the gaps are. I promise I&#x27;m telling the truth :)
        • mathisfun1232 hours ago
          same reason they spend huge amounts of money on rolexes when seikos work better (the tech crowd isn&#x27;t immune from vanity).
          • throwaway274482 hours ago
            If you seriously think apple products are nothing but a status item, you&#x27;re deluding yourself and probably have been for decades.
            • _hugerobots_2 hours ago
              This 1000%. Data centres don&#x27;t equate to medium sized labs and businesses. A stack of Macs is up and running without digging trenches, an electrician on staff and a department of PhDs to justify the spend.
              • bigyabai1 hour ago
                It&#x27;s likely that a stack of Macs will draw more power for slower prefill&#x2F;decode than equivalently priced Nvidia GPUs. If power efficient inference is the goal, Macs are a non-starter.
                • _hugerobots_1 hour ago
                  So if it isn&#x27;t a comparative ability, now it&#x27;s a power cost issue? This reads like goal post moving.
            • mathisfun1231 hour ago
              If you seriously think apple cares about anything other than cell phones, you&#x27;re deluding yourself and probably have been for decades.
              • throwaway274481 hour ago
                ...did you mean profit? I don&#x27;t think they&#x27;re manufacturing iphones just on the hope they delight you. This is also true of Google et al.<p>I don&#x27;t get these weird parasocial emotional attachments&#x2F;beefs people have with brands. Talk to a therapist.
                • mathisfun12326 minutes ago
                  brother my point is they don&#x27;t care about their product offerings outside of their phones. this post&#x2F;thread is about one of their product offerings which is not a phone which is inferior to their competitors&#x27;. simple.
                  • selectodude18 minutes ago
                    My M1 Pro MBP is 6 years old and continues to be the best computer I own, so if that’s Apple not trying, god help everybody else once they do.
                    • mathisfun1234 minutes ago
                      Jesus Christ reading comprehension has completely fallen off a cliff - we&#x27;re talking specifically about GPU performance here. Do you understand the words that are coming out of my mouth?
              • tom_1 hour ago
                They&#x27;ve been selling phones for less than 20 years at this point? Though I suppose 1.9 <i>is</i> not equal to 1, so it gets the plural.
      • liuliu2 hours ago
        Both are probably single-token decode performance, which is reasonable to show. Otherwise agree RTX 5090 should shinebetter with NVFP4.
      • searealist38 minutes ago
        ... or with llama.cpp with MTP.
      • ActorNightly1 hour ago
        [flagged]
        • bee_rider1 hour ago
          I guess it is possible, but Apple has had very vocal fans for decades. I suspect, rather than astroturfing, it is just people who are in their ecosystem.
        • abletonlive49 minutes ago
          Tok&#x2F;sec is 0 on a 3090 for most of the models that the mac can run
          • ActorNightly21 minutes ago
            Running very large models on Mac is unusable at 10 tok&#x2F;sec. You get more average inference over the day using free Google Gemini.<p>And for the price of a Mac that can run a large model, you can get 2 3090s humming along running a small model so fast that it can simulate a lot of the behavior in large models just through sheer number of context it generates. For example, editing code means that by the time your large model on your Mac is finished writing a file, the smaller models have generated the code, written the code to file, ran it, and debugged any issues.<p>So given that, which one of these is true about you?<p>1. You are paid by Apple to push marketing on HN<p>2. You are a hardcore Apple fanboy and just think that owning a Mac studio is a flex
    • karmakaze35 minutes ago
      I really appreciate seeing these dense model numbers. For a large unified memory system though I expect that MoE numbers are what people are more interested in.<p>These numbers could and should get much better. As an example I can run Qwen3.8-27B-MXFP4 (W4A8) on 2x AMD R9700 that gets 260+ tokens&#x2F;sec to start and slows down to ~110 tokens&#x2F;sec over 128k context and can do the max 256k. These are for batch size 1 and throughput goes higher with batching. This is due to speculative decoding, efficient all-reduce inter-gpu compression, and custom GEMM kernels for the specific hardware. Note each R9700 only has 644 GB&#x2F;s memory bandwidth.
    • nacs2 hours ago
      That&#x27;s a dense model. Of course it will do worse.<p>Now try running that Qwen 3.8 Next model on the 5090 and tell me what TPS you get (hint: it&#x27;s near 0 since it doesnt fit the 32GB VRAM on 5090 vs the 256 in OPs M5).
      • peri-cl2 hours ago
        Surprisingly, the Reddit crowd are reporting 50–60 tokens&#x2F;s (for the 32 GiB 5090 + 128 GiB RAM)—on par with the M5 Ultra benchmarks, despite both the PCIe bottleneck and much smaller DDR5 bandwidth,<p><a href="https:&#x2F;&#x2F;old.reddit.com&#x2F;r&#x2F;LocalLLaMA&#x2F;comments&#x2F;1wl06np&#x2F;qwen38flashnext_on_1x_rtx_5090_tg50_ts_pp2300_ts&#x2F;" rel="nofollow">https:&#x2F;&#x2F;old.reddit.com&#x2F;r&#x2F;LocalLLaMA&#x2F;comments&#x2F;1wl06np&#x2F;qwen38f...</a><p>(Note it&#x27;s a sparse MoE with only 6B active).
        • nacs2 hours ago
          Good to know thanks.<p>That&#x27;s with CPU offload to a DDR5 6000 RAM though which is around $3-4k at least.
    • redox993 hours ago
      A dense 27B doesn&#x27;t really make sense for the Mac. A MoE makes way more sense when you have modest bandwidth but lots of memory.
      • tcdent1 hour ago
        A dense model (up to the amount of memory available) actually does make the most sense on unified memory architectures. But when you hit the limit of what you can hold in memory, you reach the limitation of the platform.<p>Whereas a hybrid architecture with distinct DRAM and VRAM with sparse MoE, you can leverage two different bit rates depending on the actual need for constant access to common layers versus sparse access to infrequent layers and arbitrage the difference in cost for each of those in distinct classes of hardware.
      • peri-cl3 hours ago
        They do MoE. They benchmarked GLM 5.3-flash (320B &#x2F; 18B), and Qwen 3.8-flash-next (125B &#x2F; 6B). The dense Qwen is only focused (I assume) because it&#x27;s about the only thing that fits on a 5090, that they can compare the two heads on.
    • peri-cl3 hours ago
      Those are some incredible graphs, that leap in prompt processing going from M3 to M5.<p>Also: ~30 token&#x2F;s on GLM 5.3-flash, locally. (That&#x27;s roughly Opus 4.8-tier. I think).<p>&#x2F;meta Here&#x27;s a CSS filter that stops those nuisance chart animations,<p><pre><code> macstories.net##*:style(animation: none !important; transition: none !important)</code></pre>
    • alex7o1 hour ago
      On my m5 max 27b model does 75tps on 256k ctx and starts at 80 on the 8k ctx when you add <a href="https:&#x2F;&#x2F;huggingface.co&#x2F;collections&#x2F;z-lab&#x2F;dflash-2" rel="nofollow">https:&#x2F;&#x2F;huggingface.co&#x2F;collections&#x2F;z-lab&#x2F;dflash-2</a> to it. So yeah base might be 30tps (I used iq4) but mtp or dflash help a lot and should be used when checking what is useful and what is not for running models as it is not fare to judge without them.
    • RationPhantoms3 hours ago
      Thank you for this. I wish Apple focused their silicon design on improving the TTFT metrics but coming from an M3 Pro, it still looks laggard compared to Nvidia&#x27;s TensorCores in the 5090.<p>Maybe Apple is an acquisition away from changing that balance.
      • wlesieutre3 hours ago
        The rumor on Apple&#x27;s processor roadmap is that they&#x27;re skipping other M6 variations (all previous generations had Pro and Max, a few had Ultra) in order to focus on the M7 generation for AI reasons. What exactly the M7 improvements are who knows.<p><a href="https:&#x2F;&#x2F;www.macrumors.com&#x2F;2026&#x2F;06&#x2F;25&#x2F;2027-macs-m7-chips&#x2F;" rel="nofollow">https:&#x2F;&#x2F;www.macrumors.com&#x2F;2026&#x2F;06&#x2F;25&#x2F;2027-macs-m7-chips&#x2F;</a>
        • kridsdale13 hours ago
          I think that comes down to TSMC. Nvidia apparently booked out the whole A18 or 16 node. Apple is on 2nm right now and M7 will jump right to A14. According to my quick AI research anyway.
          • dagmx2 hours ago
            That sounds a lot like AI fantasy slop.<p>Apple just shifted to N2. They’re not going to be doing another major shift right away.<p>And TSMCs own roadmap would put your hallucination years away at best for a a product that follows a roughly annual cadence <a href="https:&#x2F;&#x2F;www.tomshardware.com&#x2F;tech-industry&#x2F;semiconductors&#x2F;tsmc-unveils-process-technology-roadmap-through-2029-a12-a13-n2u-announced-a16-slips-to-2027" rel="nofollow">https:&#x2F;&#x2F;www.tomshardware.com&#x2F;tech-industry&#x2F;semiconductors&#x2F;ts...</a>
        • GeekyBear58 minutes ago
          &gt; What exactly the M7 improvements are who knows.<p>&gt; Apple&#x27;s planned M7 Ultra chip is being designed to support up to 1.5 TB of unified memory and to push AI performance toward the class of Nvidia&#x27;s Blackwell accelerators<p><a href="https:&#x2F;&#x2F;www.tomshardware.com&#x2F;tech-industry&#x2F;semiconductors&#x2F;apples-rumored-m7-ultra-targets-1-5tb-of-memory-and-blackwell-class-ai" rel="nofollow">https:&#x2F;&#x2F;www.tomshardware.com&#x2F;tech-industry&#x2F;semiconductors&#x2F;ap...</a>
      • aurareturn25 minutes ago
        M6 got another prompt processing boost. Likely no M6 Ultra though because Apple is reportedly going all in on AI performance in M7 generation.
    • api2 hours ago
      I assume those are non-batched. I think the M series GPU can do 4X to 8X depending on model quant, which means if you can batch queries you&#x27;ll get almost 4X to 8X performance.
    • jmyeet3 hours ago
      The selling point of the M5 Ultra Mac Studio is that you can run much larger models that the 5090 can&#x27;t without swapping. NVidia aggressively segments the market on VRAM for this reason. That&#x27;s why a 5090 has an MSRP of ~$2k (but good luck getting one for less than $4k) while a 6000 Pro, which is basically a 5090 with 96GB of RAM has now soared beyond $15k where 3-6 months ago it was more like $10-11k. A 6000 Pro has the same memory bandwidth but slightly more CUDA units (IIRC ~24k vs ~21k).<p>This advantage won&#x27;t be apparent with a 27B model. The 256GB MS can probably run the newer Flash models locally, something you can&#x27;t do on a 5090.<p>I don&#x27;t think we&#x27;ll get a successor to the 5090 until late 2028, maybe even 2029. I&#x27;m basing this on the launch date of the 5000 series and that we haven&#x27;t got a midcycle refresh yet. Rumor has it the chips are ready but the 3GB RAM modules are 3-4x the price of the 2GB modules used on the current cards.<p>Apple should see a Mac Studio major update in 2028. That might even force NVidia&#x27;s hand. But it&#x27;s really impossible to say what the state of the market will be 2-3 years from now. It may have completely crashed. I suspect not however.<p>The interesting thing will be when the bandwidth demands start forcing HBM memory onto these home&#x2F;enthusiast solutions.
      • pama3 hours ago
        But what about builds that combine 8 of the 5090 with infiniband between boxes? Wouldn&#x27;t that be comparable to the mac in terms of price and potentially beat it by a lot in terms of performance for the large MoE? I understand the space&#x2F;heat&#x2F;noise considerations, but price wise it may still not make as much sense as people think. (Agreed that it is hard to get the NVIDIA hardware and the 6000 pro are priced less competitively).
        • throw0101c22 minutes ago
          &gt; <i>But what about builds that combine 8 of the 5090 with infiniband between boxes?</i><p>Why Infiniband (&quot;IB&quot;)? If it&#x27;s for RDMA, that is possible with certain Ethernet cards&#x2F;chipsets as well. Certainly Mellanox, but Broadcom:<p>* <a href="https:&#x2F;&#x2F;techdocs.broadcom.com&#x2F;us&#x2F;en&#x2F;storage-and-ethernet-connectivity&#x2F;ethernet-nic-controllers&#x2F;bcm957xxx&#x2F;adapters&#x2F;RDMA-over-Converged-Ethernet.html" rel="nofollow">https:&#x2F;&#x2F;techdocs.broadcom.com&#x2F;us&#x2F;en&#x2F;storage-and-ethernet-con...</a><p>and Intel as well:<p>* <a href="https:&#x2F;&#x2F;www.intel.com&#x2F;content&#x2F;www&#x2F;us&#x2F;en&#x2F;support&#x2F;articles&#x2F;000031906&#x2F;ethernet-products&#x2F;700-series-controllers-up-to-40gbe.html" rel="nofollow">https:&#x2F;&#x2F;www.intel.com&#x2F;content&#x2F;www&#x2F;us&#x2F;en&#x2F;support&#x2F;articles&#x2F;000...</a><p>Link level flow control or priority flow control needs to be supported on the switch ports as well.
        • wmf2 hours ago
          No, $40K is not comparable to $10K.
        • prmoustache2 hours ago
          Sounds like nice utility bill in the making.
        • kridsdale13 hours ago
          While that sounds super awesome, How many people are actually going to build and maintain that vs a box you can grab at the mall that fits in a lunchbox?
        • jmyeet2 hours ago
          I can&#x27;t speak to Infiniband pricing for something like that. It seems like the cheap option is 56&#x2F;100Gbps with used Enterprise equipment. You&#x27;d need 8 HCAs, DAC cabling and a switch but even then you&#x27;re into thousands of dollars. If you want 200Gbps+ it gets into the tens of thousands (AFAICT).<p>Each PC is probably going to cost ~$6k and you&#x27;re talking about 8000W of electricity draw. That&#x27;s going to consume multiple 20A circuits even at 240V. And the electricity ain&#x27;t free either. A Mac Studio seems to draw ~500W max.<p>Oh and the Mac Studio has an upgrade route to run 1T+ models too by chaining them together with TB5 chaining. OSX supports RDMA this way. That&#x27;s comparable bandwidth to the 100Gbps Infiniband option.<p>So you&#x27;re talking about $50-60k of hardware and more power draw and more heat for something that will I&#x27;m sure beat the MS M5U option but at huge cost. Also, at that kind of price point, I&#x27;m likely to get a workstation PC and put 2 (or possibly 3) 6000 Pros in it.
    • traceroute662 hours ago
      Not forgetting of course that an RTX5090 is what 600W+ ? And the Mac is probably half that at most ?
      • washadjeffmad1 hour ago
        Certainly not forgetting wattage. A 5090 is 575W. The M5 Ultra Studio is 480W.<p>nvidia-smi -pl 450 for like a 4% reduction in throughput. I tend to set it around 350W because it&#x27;s a comfortable temperature blowing on my legs under the desk without warming my office in the summer.<p>I put together this system two years ago, so it&#x27;s a little out of date, but it only cost $3000 for the same performance and capability as an Ultra. I don&#x27;t think I would spend $7000 to save 100W, though.
        • TacticalCoder24 minutes ago
          &gt; nvidia-smi -pl 450 for like a 4% reduction in throughput.<p>Yeah people don&#x27;t pay enough attention to those settings IMO. The first thing I do when I set up a new machine (or upgrade my OS) is to restore all my powersaving configs.<p>For example I&#x27;ve got all but one of my virtual desktops that put the CPU in powersave mode: I don&#x27;t need max Ghz when browsing the Web, not even on demand. But when I switch to the virtual desktop where my development environment is, then I want power on demand.<p>Now I don&#x27;t do it to save the planet: I do it because I love a quieter computing experience (coupled with Be Quiet! PSU and Noctua fans, this makes for a very quiet computer). That it consumes less electricity is a nice side-benefit.
      • beastman822 hours ago
        sure. so is 2x power worth 10x perf? I think it is in most cases.
      • ActorNightly1 hour ago
        When you are doing matrix math, compute is compute. Apple cant be more efficient due to physics. The only reason Macs are more efficient in general is that they have tightly bundled hw and sw for specific tasks.
    • GeekyBear1 hour ago
      The next Ultra, supposedly on deck in 2028:<p>&gt; Apple&#x27;s planned M7 Ultra chip is being designed to support up to 1.5 TB of unified memory and to push AI performance toward the class of Nvidia&#x27;s Blackwell accelerators, according to a new Bloomberg report published by Mark Gurman...<p>Apple plans to release a base M6 chip this fall for entry-level Macs... a base M7 in the first half of 2027, M7 Pro and M7 Max at the end of 2027, and the M7 Ultra in 2028.<p><a href="https:&#x2F;&#x2F;www.tomshardware.com&#x2F;tech-industry&#x2F;semiconductors&#x2F;apples-rumored-m7-ultra-targets-1-5tb-of-memory-and-blackwell-class-ai" rel="nofollow">https:&#x2F;&#x2F;www.tomshardware.com&#x2F;tech-industry&#x2F;semiconductors&#x2F;ap...</a>
  • srcreigh4 hours ago
    This is great as a first look, but the author is not a developer, so we don&#x27;t yet know whether a dev can be as productive with local models on M5 Mac Studio compared to a 20x subscription plan.<p>I&#x27;m also curious about any new low hanging optimization opportunities in the kernels for this new hardware.<p>It&#x27;s already clear to me that M5 Mac Studio is more cost-effective than anything you can run on open router, assuming decent utilization.<p>The M5 Mac Studio will be the most cost effective way to run uncensored cyber capable open agents.<p>An exciting tipping point will be if programmers can get an Astra-Ultra like experience all week with this hardware. That would be a real sense where this hardware exceeds the value of even 20x cloud subscriptions.
    • zozbot2342 hours ago
      Astra-Ultra? Even the largest open model to date (Kimi K3) is nowhere close to Astra level, and it will be quite slow even on the highest-spec M5 Ultra, with achievable speeds of about 0.5 tok&#x2F;s at most due to having to stream weights from SSD (~13 GB&#x2F;s on the highest storage capacity M5 Max machines so far). This is OK for doing simple Q&amp;A in the background but it&#x27;s far from a genuine coding experience. You&#x27;d have to test batching of multiple thinking streams in order to try and raise overall tok&#x2F;s via layer-wise reuse of the streamed weights (and this is where the &quot;Ultra&quot; part sort of becomes relevant; Kimi series models have good support for agent swarms) but this would decrease single-session performance even further. It would only be usable for background jobs, though the hardware would then have a chance of paying for itself if it was fully used on a 24&#x2F;7 basis.
      • srcreigh3 minutes ago
        &gt; You&#x27;d have to test batching of multiple thinking streams in order to try and raise overall tok&#x2F;s via layer-wise reuse of the streamed weights<p>isn’t this very straightforward to do..? I thought batching for Qwen models is already proven out.<p>&gt; but this would decrease single-session performance even further<p>Well let’s take Qwen 3.8 27B. Throughput for M3 at 8 agents is 4x compared to single agent. [1]<p>It’s really not clear to me that 8 concurrent agents at half speed will be worse task completion latency than 1 agent.<p>And that’s M3 studio benchmarks, not even M5 ultra, and without the many software improvements we will see<p>If you haven’t tried Qwen 3.8 27B xhigh on a task you might not get the hype. Idk.<p>If you’ve <i>tried doing this and don’t like it</i> sure, and be specific about what isn’t effective, but let’s not speculate.<p>[1]: <a href="https:&#x2F;&#x2F;omlx.ai&#x2F;benchmarks&#x2F;performance&#x2F;69kzkrv8?utm_source=chatgpt.com" rel="nofollow">https:&#x2F;&#x2F;omlx.ai&#x2F;benchmarks&#x2F;performance&#x2F;69kzkrv8?utm_source=c...</a>
    • slowin3 hours ago
      &gt; <i>This is great as a first look, but the author is not a developer, so we don&#x27;t yet know whether a dev can be as productive with local models on M5 Mac Studio compared to a 20x subscription plan.</i><p>Local models are definitely not as productive as SOTA, sadly it&#x27;s not close yet. I do think someday they will be &quot;good enough&quot; to use, but they aren&#x27;t today. Even the SOTA models <i>barely</i> code well, with Opus 4.5 being the first, good coding model.<p>That being said, I think it&#x27;s absolutely imperative that we keep pushing local model performance. We need to continue to advance technology there <i>and</i> ensure that the model labs don&#x27;t do regulatory capture in the name of &quot;safety&quot; (or anything else).
      • nowittyusername1 hour ago
        With the latest codex (weekly quota burn) fiasco I tried open weight alternatives for the first time. And tyeah... open weight models cant compete with likes of astra yet. But, my hope is that by the time I get my Mac studio at end of november an open weight models would have closed the gap (which i think is realistic at the speed of progress). Now its true a better gpt version will also be available then but it also seems the gap is shrinking with time so theres that.
      • _hugerobots_1 hour ago
        Local models can be widely used as productive assets. Yes the infrastructure of SOTA API models is engineered specifically for you to be that utility, but the blanket statement that local isn&#x27;t up to par is intensely short sighted. Billions of tokens per month on local pays for the hardware when compared to sota costs per month.
        • slowin1 hour ago
          I believe they can currently be used productively for non-coding tasks (classification, light summary)... but they definitely are not even close to SOTA when it comes to software development.
          • _hugerobots_1 hour ago
            Defining productivity is a use-case scenario, and a wildly generalized assumption for most people in this argument. Local infrastructure doesn&#x27;t need to be sota for absolutely every single need for a dev lab, but it absolutely can be delivered with non-api frontier class models.
            • slowin1 hour ago
              Just to be clear, I&#x27;m specifically talking about coding. I think local models can help with productivity today, just not coding.<p>I&#x27;m also a <i>huge</i> fan of local models and think it&#x27;s absolutely imperative that they continue to advance so we can move off of the Anthropic&#x2F;OpenAI hosted models. It&#x27;s important to accurately asses where we are in that journey though.
              • _hugerobots_41 minutes ago
                Like the other commenter, I&#x27;m confused about the &#x27;just not coding&#x27; conclusion. I&#x27;m using Qwen 27B on a 5090 at &gt; 100tk&#x2F;s with 150k context (which isn&#x27;t enough admittedly), and DeepSeek v4 Flash with 1million context on a gb10&#x2F;spark. Both of which are performing surface level, and deep needle precision infrastructure architecture. They code 24-7, stupendously.
              • srcreigh1 hour ago
                I think the issue is generalization, if you were more specific about which local models aren’t good enough for which tasks compared to which frontier models in your experience, it’d be a lot more informative
                • slowin22 minutes ago
                  I can&#x27;t just go into any codebase and ask a local model to &quot;Implement this feature: xxx&quot; and get acceptable output. I hope to someday soon though!
  • sajithdilshan4 hours ago
    On Apple website it says 512GB memory option is available in October. I guess bumping to that one would cost additional 4-6k US$. So an Ultra with 2TB storage would be north of 15k US$.<p>That’s like 12 years worth of OpenAI Pro subscriptions
    • 1122334 hours ago
      Hard to guess, it can go either way. If you will need to be in a syndicate to use non-sterilized models, that mac makes sense. But if there is mandatory registration of personal cyberarms, you risk going to mines once they check you purchases. You could try to play normie and pretend you simply wanted to show off, by keeping your actual work on external disk, but that leaves traces on system. Counting on someone in the Gap renting you gray iron works as long as you can swap credits. Still, this gear is tiny. Put it in your e-car, with uplink, and leave it at uncle&#x27;s farm. Discreet.
      • woah37 minutes ago
        It was a dark rainy night in Neo-Tokyo as Blake puffed on his vapor cartridge and watched the Mac dealers prowl below. Almost 15k Union Credits to get one of them to meet you in an e-cafe with a fully loaded M5, but man, the inference rush from one of those things was something else.
      • Razengan3 hours ago
        I gotta have some of what you had :)
        • kridsdale12 hours ago
          I thought it was a fun bit of cyberpunk fiction. Those who downvoted him seem to have taken it at face value?<p>I appreciate the reference to RUSH: Red Barchetta in the final line.
    • nowittyusername1 hour ago
      512 option isnt worth it imo, you get severe slowdowns when weights are that large. 256 is the sweet spot, you can run large open weight models at decent speeds for full private inference.
      • throw0101c45 minutes ago
        &gt; <i>512 option isnt worth it imo, you get severe slowdowns when weights are that large.</i><p>I think most people are getting 512 for running Chrome with a bunch of tabs open. &#x2F;s
    • geodel4 hours ago
      Agreed.<p>Specially since one can pay half right now to OpenAI and sign a 12 year iron clad contract for uninterrupted service delivery of OpenAI Pro.
      • Kurtz793 hours ago
        I think we all expect the heavy subsidized subscriptions to end or significantly increase in price at some point, but it could be years from now and I&#x27;d rather spend a similar figure on an hypotetical Mac Studio M8 Ultra, or whatever more advanced competitor that will have likley appeared by that time.<p>A more apples-to-apples comparison would be with API cost in OpenRouter at the same tok&#x2F;s rate for the same models that you can run locally, maybe.
        • qwytw33 minutes ago
          &gt; heavy subsidized subscriptions to end<p>Is there evidence that&#x27;s true though? I mean gross margins on subscriptions being negative since the API is seemingly very profitable (if the price is compared with the cost of serving very large open models).<p>As long as there is pressure from other providers serving cheaper models that are somewhat competitive without having to incur any of the R&amp;D costs raising prices will be tricky.
        • BatFastard1 hour ago
          &gt;A more apples-to-apples comparison<p>Don&#x27;t you mean an Apple to NVidea comparison?
      • vardump4 hours ago
        I hope that was sarcasm.
      • patrickmcnamara3 hours ago
        HN always has these completely contrived counterarguments. What is actually going to realistically happen that will prevent use of an LLM provider? Did you think that the OP literally meant the 12 years or maybe it was just to show how expensive using a Mac Mini as an alternative is?
        • geodel2 hours ago
          &gt; how expensive using a Mac Mini as an alternative is?<p>I think it goes without saying. And it is eminently evident over last couple of decades that from compute to storage to meals 3rd part providers have saved billions upon billions of dollars to enterprises and individuals alike by providing these essential services.
        • kridsdale12 hours ago
          Mass revolts of the peasantry burning down data centers and cutting fiber lines.
          • geodel2 hours ago
            Yes, it feels like that. Whereas frontier labs are pushing the frontier of human knowledge, selflessly working towards pulling humanity from dark ages. Ignorant peasants trying to burn the modern civilization down. Don&#x27;t they know data centers and fiber lines are lifeline of modern economy?
    • ericmay3 hours ago
      Just commenting here because you&#x27;re discussing hardware: I thought the test results from the SSD published in this article [1] were pretty interesting. Maybe that&#x27;s old news though.<p>[1] <a href="https:&#x2F;&#x2F;www.macworld.com&#x2F;article&#x2F;3238319&#x2F;mac-studio-m5-max-review.html" rel="nofollow">https:&#x2F;&#x2F;www.macworld.com&#x2F;article&#x2F;3238319&#x2F;mac-studio-m5-max-r...</a>
    • simonw4 hours ago
      Yeah, anyone who thinks local AI is going to save them money is likely to be disappointed, at least if they want to run models that are even remotely capable.<p>Plenty of other reasons to get excited about local AI, but I don&#x27;t think cost is one of them.
      • criddell3 hours ago
        Maybe you are using a local model to go after some Millennium Prize problem and you don&#x27;t want OpenAI to take your work and use it to win the prize for themselves? $15k might be a bargain.<p>And, yes, I know a current local model wasn&#x27;t going to solve the Navier-Stokes problem, but I&#x27;m just using it as an example where privacy might be valuable.
        • simonw3 hours ago
          Agreed, plenty of other reasons to get excited about local AI.
      • hgoel2 hours ago
        Despite being on a site called Hacker News, we seem to often overlook the simple aspect of wanting local AI hardware to hack (not necessarily in the cybersecurity sense) with. I got my local AI hardware because it&#x27;s an enjoyable hobby for me.
        • ionwake58 minutes ago
          apparently if you ever point out HN starts for hackernews and thus expect related attitudes you get downvoted by shocked ( what I guess are zoomers and not bots ) that desperately opine the name is a random abberation doesn&#x27;t mean anything and one should not deviate from our corporate overlods in any manner.
  • tempoponet4 hours ago
    While I know it&#x27;s not apples to apples, the target comparison right now is 2x DGX Sparks. Similar price, 256gb. The conversation has focused on memory bandwidth vs. compute in agentic loops, so for most people the raw numbers will mean less than the &quot;time per task&quot; in coding benchmarks.<p>This is a great article and bodes well for the M5, but we should expect more like this comparing to other platforms before we truly understand where it fits.
    • _hugerobots_1 hour ago
      Speed vs task-completion is a new conversation and a great point. Whereas the cost to compute doesn&#x27;t exist in a vacuum, making mistakes costs less, is easier to maintain with granularity and a whole host of other factors when you own the lab.
  • akozak2 hours ago
    &quot;a total cost of $0&quot; Uhh ... how much is that hardware?
    • novaleaf2 hours ago
      Another comment approximates at around USD$15k, so yeah, not zero.
  • ApolloFortyNine4 hours ago
    The model being tested is 18k as configured.<p>I didn&#x27;t expect this to make the 5090 to look like a good deal.
    • nacs2 hours ago
      5090 has 32GB VRAM.<p>It&#x27;d be silly to buy the 18k model to run a tiny model like Qwen 27B. You use models like GLM Flash and Qwen Next which won&#x27;t fit on a single 5090.
      • asimovDev1 hour ago
        can run multiple subagents of Qwen 27B though, right? Unless I am fundamentally misunderstanding how VRAM constraints work
        • Eisenstein33 minutes ago
          You might be. Running another agent doesn&#x27;t load a set of new weights. It creates a new KV cache for the agent and adds the prompts to the queue. Its just another inference turn.
      • orsorna2 hours ago
        Is it that silly? You could run multiple 27B models in parallel.
        • peri-cl2 hours ago
          You actually don&#x27;t need more RAM to batch multiple inference tasks of the same model.<p>(Each task needs its own context, but the (e.g.) 27B of constant parameters isn&#x27;t duplicated).
  • liuliu2 hours ago
    When people benchmark MLX related quant models, they really need to publish numbers on benchmarks. You cannot take this as it is what you get of the original models. MLX uses pretty simple quantization methods so at lower bits without QAT, it is just not as good quality as llama.cpp ones.
  • hamiltont57 minutes ago
    Once you hit the memory you need, generation speed is mainly set by bandwidth, and every Ultra from M1 thru M3 has ~800 GB&#x2F;s. IMO best ROI for most people is &#x27;cheapest used Ultra with enough RAM&#x27;<p>I setup an eBay alert and picked up a used M2 Ultra that has delivered good ROI (at least, far better than 15k for comparable-for-my-use-case performance)
    • peri-cl49 minutes ago
      I think M1 through M3 were compute bottlenecked in prompt processing (hence the very large gap between M3 and M5, in this page&#x27;s benchmarks, that&#x27;s not explained by memory bandwidth alone).<p>For <i>generation</i> speed in isolation, yes.
      • GeekyBear37 minutes ago
        The M5 generation added tensor instructions to the GPU cores.
    • Lwerewolf55 minutes ago
      This one is 2x m5 max, so ~1.2TB&#x2F;sec.
  • SamuelAdams2 hours ago
    I think Apple is really sleeping on making this run a Linux server. These things are very capable and draw very little wattage when idle. It would make an excellent homelab device, but MacOS currently holds it back in this regard.
    • flounder32 hours ago
      <a href="https:&#x2F;&#x2F;mac.getutm.app&#x2F;" rel="nofollow">https:&#x2F;&#x2F;mac.getutm.app&#x2F;</a>
    • jjtheblunt16 minutes ago
      i use linux a ton too, but still wonder what you want in a Linux server that macos as a BSD server does not have.
  • crossroadsguy2 hours ago
    My mac is 5 years old. I don&#x27;t think I can comfortably buy a new one right now. It has a 16GB unified RAM. Honestly that would be enough for so many local models that I want to use but can&#x27;t use. Because RAM usage (even with literally every single user installed app quit&#x2F;stopped) the RAM usage is very high that I can barely safely get 6-7 GB (I am supposed to get ~10 GB, but it goes up and down real fast!). That&#x27;s a shame. If only I could install an alternative OS that uses very little amount of RAM :-)
    • mjlee3 minutes ago
      How are you measuring memory usage? top tells me that 45&#x2F;48GB is &quot;used&quot;, but Activity Monitor shows me that 24GB is cached files.<p>I&#x27;d be quite surprised if Mac OS alone needs more than 8GB, given that they sell the Neo with 8GB of RAM today.
  • theplumber2 hours ago
    At this point I think I will get the DGX gb300 workstation though I will wait a bit more for the cold season. It is double the price but at least is the real thing
  • addaon3 hours ago
    Ordered one for OpenFOAM. Excited for it. Will be nice to not have my laptop running CFD 24 hours a day, but my M1 Max is currently my fastest machine… I’m expecting about 3.5x from the M5 Ultra.
  • crorella2 hours ago
    What are good options to run local models nowadays? Something good for coding and personal assistant kind of things
  • kokonokko13374 hours ago
    &gt; &quot;It also happens to be a Mac, with an operating system that looks nice and doesn’t suck&quot;<p>Yes Apple has some of the best hardware out there, albeit overpriced. But the software is such a hindrance and I can&#x27;t take anyone that states otherwise seriously. If only it had proper Linux support (and the Asahi people do an amazing job but you can reverse-engineer only so many stuff with limited funding, and then you have to do it again for new models). MacOS is good if you just want to have a standard experience, which to be fair is most people. It&#x27;s good for just setting up an LLM server I guess since the hardware is a perfect fit. I wouldn&#x27;t touch it otherwise.
    • steve19771 hour ago
      What exactly is missing from macOS that makes you feel the need for Linux?<p>I get it on Windows systems, at least when someone wants to use Linux-type tooling. But macOS already supports pretty much all of that natively?
      • RunSet1 hour ago
        &gt; What exactly is missing from macOS that makes you feel the need for Linux?<p>For starters, the source code.
        • steve19771 hour ago
          And why would you need that to run LLMs?<p>Apart from that, for the UNIX part, the source is available for quite a few components:<p><a href="https:&#x2F;&#x2F;github.com&#x2F;apple-oss-distributions" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;apple-oss-distributions</a><p>notably also the kernel<p><a href="https:&#x2F;&#x2F;github.com&#x2F;apple-oss-distributions&#x2F;xnu" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;apple-oss-distributions&#x2F;xnu</a>
          • throw0101c42 minutes ago
            &gt; <i>Apart from that, for the UNIX part, the source is available for quite a few components:</i><p>Strictly speaking, Apple can claim to ship a UNIX® operating system:<p>* <a href="https:&#x2F;&#x2F;www.opengroup.org&#x2F;openbrand&#x2F;register&#x2F;apple.htm" rel="nofollow">https:&#x2F;&#x2F;www.opengroup.org&#x2F;openbrand&#x2F;register&#x2F;apple.htm</a><p>* <a href="https:&#x2F;&#x2F;www.opengroup.org&#x2F;openbrand&#x2F;register&#x2F;" rel="nofollow">https:&#x2F;&#x2F;www.opengroup.org&#x2F;openbrand&#x2F;register&#x2F;</a>
            • steve197740 minutes ago
              Yeah for a while they used that in marketing copy actually, but it&#x27;s been a while I think.
          • Gracana32 minutes ago
            &gt; And why would you need that to run LLMs?<p>kokonokko1337 already said it was good enough to run LLMs, presumably RunSet isn&#x27;t saying the source code is needed to run an inference server.
  • BatchJob3 hours ago
    If you are buying expensive hardware to run LLMs &quot;on your own machine&quot; you will soon find your ladder is on the wrong wall.
  • devy3 hours ago
    This dream machine costs over $15k (not including the Apple Studio Display)? Nah, that dream is SO OUT OF TOUCH!
    • aenis26 minutes ago
      The irony here is, thats hobby hardware. You spend 20k and can run slow hobby models that are barely capable of anything unsupervised.<p>Entry level serious hardware starts at 100k, and a bit better but still almost-useful grade is 200k (8x rtx pro, plus a nice epyc pairing). Thats the sort of thing a salaried expert lets their employer buy them for sort of serious work.<p>Anything really serious is well north of 1M - not including the housing and commercial grade mains connection. And at best that buys fast Kimi K3 or GLM.
  • saejox1 hour ago
    i can buy a house with that amount of money. it used to be car money.
  • 12kaj22 hours ago
    The Year Of Local AI will be here no later than 2040, coinciding with the Year Of The Linux Desktop.
    • prmoustache2 hours ago
      The year of linux on the Desktop was 26 years ago for me.
  • snarfy4 hours ago
    $12,299
    • andrekandre3 hours ago
      5 years of (200&#x2F;month) tokens at that price, meanwhile an rtx 5090 pc is about half that… hmm<p>but i wonder how much these token costs are sustainable or not, it may be in the long term cheaper to have your own hardware if token costs go up (and hopefully hardware gets cheaper again)
      • chasd002 hours ago
        The token price isn&#x27;t the only reason to run a model locally though. You can do additional training to specialize or remove censorship that may be a no-no per TOS with cloud GPUs.
      • Eisenstein29 minutes ago
        $2200 for a 64GB VRAM machine if you are willing to do a bit of work.<p>* <a href="https:&#x2F;&#x2F;imgur.com&#x2F;mpdorVJ" rel="nofollow">https:&#x2F;&#x2F;imgur.com&#x2F;mpdorVJ</a>
  • slashtom1 hour ago
    Fantastic review, this is how it should be done with local AI.
  • cptskippy2 hours ago
    I think we&#x27;ll eventually get to the point where folks will have a local AI agent but I think people need to temper their expectations to a degree. You aren&#x27;t going to have data center level tok&#x2F;s from a box sitting under your desk and you don&#x27;t need instantaneous responses for many workloads. Having a local agent that can execute tasks over a couple days with your supervision that might otherwise take you weeks is perfectly acceptable.<p>However I also think that Agentic AI is very much not an out-of-the-box solution, local or otherwise, and it takes a high level of technical knowledge to create an effective AI agent. And there&#x27;s a problem now where most orchestration is fixed on what models are used for what tasks with no ability to weight constraints like cost, speed, and security.
  • villgax2 hours ago
    Lol, try generation of images &amp; videos on these, they ought to improve perf on Deep learning not just llms
  • WarmWash4 hours ago
    &gt;Let’s address the elephant in the room first: why bother with local AI at all when cloud frontier models are better and often faster?<p>Ehh, the actual elephant in the room is:<p>&quot;why bother with local AI at all when you can lease a GPU for $5&#x2F;hr?&quot;<p>To which the answer is you shouldn&#x27;t bother, unless you have a bunch of money to throw at hobby projects.
    • Youden1 hour ago
      $5&#x2F;hr = $3600&#x2F;mo.<p>Unless you only need the AI available some of the time, $5&#x2F;hr is pretty expensive. That&#x27;s an RTX Pro twice a year.<p>If you&#x27;re using it for discrete sessions of coding or something, that might make sense for you but if you&#x27;re using it for an always-on assistant, that pricing kinda sucks.
      • WarmWash14 minutes ago
        I would imagine extremely few people are utilizing an H200 for every hour of a month. Especially for something like an assistant<p>5090&#x27;s are like $0.20&#x2F;hr
    • chasd002 hours ago
      &gt; unless you have a bunch of money to throw at hobby projects.<p>there are lots of people with very expensive hobbies, see sailboat racing for example.
  • sghiassy3 hours ago
    Imagine spending a trillion dollars on data centers and then reading this article. Nightmare fuel for OpenAI
    • whalesalad3 hours ago
      For 99.99% of people, spending 15 grand on a Mac Studio just to run Qwen 3.8 locally is a non starter.
      • jmull3 hours ago
        It&#x27;s not the M5 Ultra itself, but the M7s or M9s that will do the damage.<p>99% of people will use whatever AI is free. The sophisticated, heavy users that are willing and able to pay a lot of money the ones that will be interested in controlling their inference bills.<p>Today, the sweet spot where an M5 Ultra makes sense is tiny. But we might expect that to grow a lot.
        • BatFastard1 hour ago
          Anthropic is reporting 100 Billion ARR.<p>Even if you could get a frontier model, you would not be able to run it on any Mac. So speculating on what M7 or M9 will achieve in 5 years (if we even still exist) seems pointless.
          • sghiassy1 hour ago
            Do you need a frontier model to write emails, check your calendar, search the web?<p>I don’t think Apple is going to lie down and cede AI to the cloud.
            • geodel36 minutes ago
              &gt; Do you need a frontier model to write emails, check your calendar, search the web?<p>How about writing mail to President and senators on AI doomsday scenario if frontier labs do not pace themselves?<p>That mini model on mac mini would scared to hell to do such thing. It need that rugged frontier model to speak truth to power.
      • sghiassy2 hours ago
        Yes, but in 7 years?
        • whalesalad1 hour ago
          In 7 years we will probably all be living under ground fighting skynet with plasma rifles made out of old microwave parts
          • fragmede1 hour ago
            You will. Some of us are going to be already ground into dust that the microwave parts are made out of. Others will be locked into our communism cubes with our daily allotment of entertainment and sustinece. let out into the sunlight for only 30 minutes per day.
      • beastman822 hours ago
        at 15 tok&#x2F;s
    • ajross52 minutes ago
      I don&#x27;t see how that math works? This is a $15k rig under benchmark and per the results it competes very acceptably against... one consumer GPU.<p>I really don&#x27;t see who buys this, except people who want the Studio for some other reason. But nothing in the story says you want to fill racks with these instead of Blackwell or TPU parts; it&#x27;s not even close.
    • CamperBob21 hour ago
      And nightmare fuel is just what they&#x27;ll be selling at the UN this week, for this very reason.<p>Sam&#x27;s address will probably be more riveting, imaginative, and terrifying than the last couple of Terminator screenplays. Legislators will lobby <i>him</i> to write the laws for them, and the ghost of Harlan Ellison will threaten to sue him.
  • itsmeduncan20 minutes ago
    [flagged]
  • rwissinger4 hours ago
    [dead]