17 comments

  • neomantra5 hours ago
    I maintain a fork of ds4 as shared libraries and thus can be used with other languages via FFI, along with public builds&#x2F;binaries [1]. I made ds4go [2] against ds4 using techniques inspired by yzma.<p>In addition to the library bindings, we have a small library of tools (workspace for view&#x2F;edit, scratchpad for persistence) and making your own is registering a Go function. And in recent weeks, I added the Vision and Qwen support, as ds4 added them.<p>Even if you don&#x27;t use the Go library, the ds4go binary makes it really easy to download the libraries off of HuggingFace with a TUI available vie Homebrew.<p>Here&#x27;s some TUI toy screenshots, sorry I still haven&#x27;t released that code; it&#x27;s of different quality than the others. [3]<p>EDIT: add ds4go TUI screenshot gist [4]<p>[1] <a href="https:&#x2F;&#x2F;github.com&#x2F;NimbleMarkets&#x2F;ds4&#x2F;releases&#x2F;tag&#x2F;v0.8.20260920" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;NimbleMarkets&#x2F;ds4&#x2F;releases&#x2F;tag&#x2F;v0.8.20260...</a><p>[2] <a href="https:&#x2F;&#x2F;github.com&#x2F;nimblemarkets&#x2F;ds4go#install" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;nimblemarkets&#x2F;ds4go#install</a><p>[3] <a href="https:&#x2F;&#x2F;gist.github.com&#x2F;neomantra&#x2F;ae47422c8daf7a458212c93992b3e078" rel="nofollow">https:&#x2F;&#x2F;gist.github.com&#x2F;neomantra&#x2F;ae47422c8daf7a458212c93992...</a><p>[4] <a href="https:&#x2F;&#x2F;gist.github.com&#x2F;neomantra&#x2F;40180ade13df93290250ce8c6d28c9f6" rel="nofollow">https:&#x2F;&#x2F;gist.github.com&#x2F;neomantra&#x2F;40180ade13df93290250ce8c6d...</a>
  • twoodfin6 hours ago
    <a href="https:&#x2F;&#x2F;github.com&#x2F;antirez&#x2F;ds4" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;antirez&#x2F;ds4</a><p>The project GitHub page is a much better introduction for the hn crowd.
  • xlayn5 minutes ago
    In case you like the store kv to disk so you can resume I keep this branch of llama.cpp that includes that same functionality<p><a href="https:&#x2F;&#x2F;github.com&#x2F;alainnothere&#x2F;llama.cpp&#x2F;commits&#x2F;disk-cache-eviction" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;alainnothere&#x2F;llama.cpp&#x2F;commits&#x2F;disk-cache...</a><p>And you know it&#x27;s load bearing each of the load baerings parts that bear some load and load a bear... you fight a bear because it took a load... or something like that...
  • simoiacos6 hours ago
    Nothing comparable but inspired from DwarfStar I wrote a little inference engine for Intel Xe-LP (no XMX) 32GB laptops. The only model supported right now is a quantized Gemma-4, but I don&#x27;t exclude in the future to support other MoE of similar size. Too bad we have no Qwen 3.8 35B-A3B yet.<p>I&#x27;m also looking into expanding the protocol and the engine to support various steering techniques.<p><a href="https:&#x2F;&#x2F;github.com&#x2F;simoneiacomino&#x2F;xenolith" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;simoneiacomino&#x2F;xenolith</a>
    • aziis984 hours ago
      Just tried this on my Intel Ultra 7 255H, I also only have an iGPU. This does ~22tps! Love this.<p>I just had to do a little patch to support my iGPU device that is a bit newer than Intel Xe-LP, maybe I&#x27;ll do a PR.<p>On a side note the other day I was experimenting with Sonnet 5.5. I gave it the llama cpp repo and told it to extract in a single file inference for a single model + backend (qwen3.5 4b mtp + sycl) and (after a long time) it actually worked! It produced a ~1400 lines file with no deps. I need to check the quality of inference yet but I think this is still a great achievement.<p>I&#x27;m pretty sure 2027 will be a very interesting year for local models and inference.
      • simoiacos3 hours ago
        Please open a PR! I was too conservative with the supported devices.<p>If your GPU supports XMX we could also explore using it to improve the prefill kernel, but I don&#x27;t have the hardware to test it myself.
    • ilaksh5 hours ago
      I wish someone would add Intel support to ds4. And also improve AMD support.<p>Maybe Intel and AMD should help them with that.
      • simoiacos4 hours ago
        Yeah I see the value but I built Xenolith to target smaller models.<p>I heard antirez saying that he designed DwarfStar also to be forked and tuned to everyone&#x27;s specific needs. Do you have a specific machine&#x2F;spec in mind?
        • ilaksh4 hours ago
          The recent Intel GPU&#x2F;AI cards. Really the same type of models as ds4
  • ttoinou4 hours ago
    Ive been using this since it was initially released with deepseek v4 flash, and it is absolutely the best launcher ever on my m5 max 128gb<p>Now Ive been running qwen 3.8 flash next for more than a week and it’s doing great, really fast and super long context windows. Sometimes the model is behaving stupidly by not remembering something I said earlier but it could be also a problem from the agentic AI harness. Im using oh my pi but Im wondering what people are using ds4 with here ?
  • cuttothechase1 hour ago
    Wondering how well this does with tool calling. Any one has any numbers or videos or anything using this?<p>From the github repo it seems like you really don&#x27;t need a big Mac with huge amounts of RAM but SSD is sufficient.<p>If this is anywhere near 50 TPS, that would be a game changer in the personal LLM space!
  • vlowther6 hours ago
    It is pretty nifty. I spend some time over last weekend implementing fused TQ to allow for 1m context lengths on a 128 gb MacBook M5 Max when using Qwen 3.8 flash next (<a href="https:&#x2F;&#x2F;github.com&#x2F;antirez&#x2F;ds4&#x2F;pull&#x2F;1115" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;antirez&#x2F;ds4&#x2F;pull&#x2F;1115</a> if you are interested). If I get bored I might port over the Metal kernels from oMLX -- the speed increase they have for the v0.7.0 release is amazeballs.
    • ttoinou4 hours ago
      I’m already able to use 1M context windows with the same machine than you and same model. Strange
      • vlowther3 hours ago
        Yeah, most of what I did was to add fused TQ support to leave more memory free for other nefarious purposes.
  • gchamonlive5 hours ago
    <p><pre><code> small native inference engine optimized first for DeepSeek V4 Flash (including the experimental vision model), DeepSeek V4.1 Flash (Metal, and text inference on CUDA), and additionally GLM 5.2 and 5.3, GLM 5.3 Flash and DeepSeek V4 PRO, and Qwen3.8 Flash Next (Metal and CUDA) </code></pre> This is local targeting high end consumer hardware like DGX Spark or AMD Ryzen AI Halo.<p>For our mere mortals that were kids not long ago and can&#x27;t really believe we&#x27;ve got our hands on a x090 series targeting Qwen3.8 27b, <a href="https:&#x2F;&#x2F;github.com&#x2F;noonghunna&#x2F;club-3090" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;noonghunna&#x2F;club-3090</a> is the way to go.<p>I&#x27;m maintaining a web frontend for this, trying to at least. You can follow it here: <a href="https:&#x2F;&#x2F;github.com&#x2F;gchamon&#x2F;club-3090-server" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;gchamon&#x2F;club-3090-server</a>
    • zozbot2344 hours ago
      First of all, a x090 series card is &quot;high end hardware&quot; in its own right these days. Secondly, I&#x27;d think you&#x27;d probably get more interesting results running a MoE model in CPU-MoE mode, i.e. with the shared parameters residing on GPU and sparse experts on CPU plus SSD offload. Yes it will be slower, but small dense models are just a dead end and not that interesting. (Note that prefill would still be sped up in this setting; the CPU&#x2F;GPU layer split in llama.cpp and the like applies to decode, but even a &quot;0 graphic layers&quot; setup does accelerate prefill.)
      • gchamonlive3 hours ago
        Qwen3.8 35ba3b is decidedly faster but also dumber, at least in my tests.
  • liuliu4 hours ago
    If you are interested in high-end models with high-end Apple Silicon, also try out Local Code: <a href="https:&#x2F;&#x2F;releases.drawthings.ai&#x2F;p&#x2F;public-beta-of-local-code-by-draw" rel="nofollow">https:&#x2F;&#x2F;releases.drawthings.ai&#x2F;p&#x2F;public-beta-of-local-code-b...</a> It is currently in TestFlight (and will open-source next week), supporting vision with DeepSeek 4.1 Flash, Qwen 3.8 27B and DeepSeek 4 Flash 0731 without vision. Custom quants &amp; SSD streaming to make these big models work with 64GiB and above devices (and of course, Qwen works with devices with 16GiB and above).
  • yieldcrv28 minutes ago
    I’m a little confused<p>ds4 is referring to “dwarfstar” “4” and references DeepSeek V4 most of the time<p>but its model agnostic-ish<p>and benchmarks compared to what? what do these large MoE models typically get in tokens per second?<p>I’m garnering this is just an easier way to load large models per expert on consumer hardware? as opposed to the hackier solutions?<p>I’m intruiged. Note that the blogpost says 64gb Macs are good minimums while the github says 96gb is a minimum
  • HoldOnAMinute5 hours ago
    How is this different from other LLM runners?
    • ilaksh4 hours ago
      Emphasis on performance and usable coding&#x2F;agentic ability for consumer AI hardware. Does not attempt to handle all models or hardware at once but rather focuses on optimizing the best options for that category of hardware.
    • simonw4 hours ago
      It&#x27;s more likely to work. Most LLM runners are meant to work with any model, which means there are all kinds of ways you might misconfigure them in a way that causes function tooling not to work, or performance to be less than you would like.<p>DwarfStar&#x27;s selling point is that it only supports a small set of carefully chosen models, but it supports them <i>really well</i>.
    • ttoinou2 hours ago
      Lots of small details are taken care of so it runs smoothly. For example ds4-agent is append only, never rewriting history of messages, keeping KV cache prefix reusable. Huge benefit
    • pydry4 hours ago
      My instinctive reaction from the readme is that it isnt. It&#x27;s apparently a vibe coded knock off of llama.CPP.
    • csmlab_notes4 hours ago
      [flagged]
  • pulkitsh12345 hours ago
    curious, why did antirez go with C instead of something like Rust ?
    • ilaksh5 hours ago
      Antirez has been writing C for a million years so is much more familiar with it than Rust.<p>Also the goal of the project is to squeeze the absolute maximum performance and capability possible out of limited hardware resources (compared to clusters of B200s or something).<p>Does Rust even give you good access to low-level code on different platforms? And if so, how much extra work do you need to do to make it acceptable to the compiler? And is that work worthwhile if you are not going to get the security guarantees of normal Rust code? Is it a worthwhile tradeoff when the goal is performance?<p>Those are real questions by the way, not rhetorical. If Rust could work well for this type of project then I would like to know.
      • Aeolos4 hours ago
        Yes, Rust gives you great access to low-level code on different platforms, including SIMD. It is also alias-free by default, and gives you excellent primitives to write multi-threaded code with compile-time correctness guarantees, which is how projects such as zlib-rs end up significantly faster than their C counterparts.[1]<p>It&#x27;s about as good as it can get for this kind of code.<p>[1] <a href="https:&#x2F;&#x2F;www.reddit.com&#x2F;r&#x2F;rust&#x2F;comments&#x2F;1ixt1ei&#x2F;zlibrs_is_faster_than_c_trifecta_tech_foundation&#x2F;" rel="nofollow">https:&#x2F;&#x2F;www.reddit.com&#x2F;r&#x2F;rust&#x2F;comments&#x2F;1ixt1ei&#x2F;zlibrs_is_fas...</a>
      • zozbot2344 hours ago
        &gt; Antirez has been writing C for a million years so is much more familiar with it than Rust.<p>This is explicitly an AI-coded project, Antirez argues that LLMs are worse at writing Rust than C because so much high quality systems code (think e.g. sendmail) that ends up in AI training sets is C, not Rust. Another related argument is that the more detailed syntax and compiler feedback found in Rust compared to C are really a negative for LLM workflows.<p>There&#x27;s plenty of room to disagree wrt. this of course: without the strong typing checks of Rust around e.g. indirect references, safety and correctness ends up being a global property in typical C programs, and LLMs are terrible wrt. reasoning about global properties. You&#x27;re better off forcing them to adapt to a different local syntax that does a more complete job of enforcing modularity, since this is comparatively foolproof.
        • cuttothechase1 hour ago
          Yes, it is AI-coded. But definitely not a one shot kind of a deal.<p>Much easier to work with a language you are most comfortable with right?
    • GTP5 hours ago
      Personal preference of the author, he made at least one video on YouTube on why he dislikes Rust. I think he finds it too cumbersome and not worth it when the software isn&#x27;t security-critical (not that I agree, just reporting what IIRC his stance is).
    • simoiacos4 hours ago
      He recently said that he finds Rust less ergonomic and that this also affects code written by LLMs, which he thinks excel at writing C partly because of the enormous, high-quality codebase they were trained on. He sees security-critical code as a reason to choose Rust.<p>The video is in Italian but has an auto-dubbed English audio track: <a href="https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=sOt0WpQG5eU\&amp;t=526s" rel="nofollow">https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=sOt0WpQG5eU\&amp;t=526s</a>
  • try-working4 hours ago
    There are insane speed improvements for local inference going around on X right now. They&#x27;ve popped up the last month and week.<p>Tensorfold is getting 100%+ speed increases on both prefill and decode for models like Qwen 27B. oMLX has followed them and have had similar improvements in the past week.<p>There&#x27;s lots of different techniques like letting CPU help with prefill, DFlash specualtive decoding etc.<p>I&#x27;m really excited for this as I&#x27;ll be receiving an M5U in about a month. Expect to be running Qwen 4 27B or Flash (it&#x27;s a 96gb machine), and they may come close in performance to DS 4&#x2F;4.1 Flash, and should be able to hit 100 tps. Local is really becoming viable, especially considering that GPT 6.1 has been running at 20ish tps the past week.
  • Almondsetat3 hours ago
    This website is pure slop. I&#x27;d ask @dang to just link the original repo
    • aeve8901 hour ago
      Right? Compare this with antirez&#x27;s blog lmao. The very author of an incredible piece of software using the most plain website possible, while a derivative post about the same tool it&#x27;s a slop fest with useless FX, cringe hackerman style palette and such. It&#x27;s just too funny.
  • doctorpangloss8 hours ago
    the problem is the dsv4 checkpoint so quantized isn&#x27;t very good
    • ilaksh5 hours ago
      Which ds4 checkpoint for which model exactly did you test? Don&#x27;t they have multiple different versions and quantization levels?
    • c0rruptbytes5 hours ago
      not my experience<p>the ds4 quants were very good beating the unsloth quants <a href="https:&#x2F;&#x2F;github.com&#x2F;michaelasper&#x2F;benchmarks&#x2F;blob&#x2F;main&#x2F;deepseek-v4-flash-0731-pi-on-slop-code-bench.md" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;michaelasper&#x2F;benchmarks&#x2F;blob&#x2F;main&#x2F;deepsee...</a>
      • dotancohen5 hours ago
        That&#x27;s quite the statement - unsloth quants are amazing.
  • 123-112927 hours ago
    [flagged]
  • fierycatnet6 hours ago
    Random comment but the name is funny to me, reminds me of Silicon Valley.<p>What are we going to name the company, how about Dwarfism 2.0? What happened to 1.0 Jared?
    • aprct2 hours ago
      I immediately thought of the unique ring from Diablo 2 named Dwarf Star.
    • seemaze5 hours ago
      Dwarf Star is better than Dirty Socks, or Dynamic Slinky.. definitely not the worst backronym.
    • dools5 hours ago
      Smallulator