3 comments

  • cmiles834 minutes ago
    Small open weight local models are the future.<p>While hosted mega models make headlines for doing cool stuff, the vast majority of applications for AI simply don&#x27;t need all that power, and thus cost. That’s a big part of why businesses are screaming that there’s no ROI from AI.<p>Brining this tech down into small local models is likely where this all converges for the vast majority of use cases and what solves the present ROI crisis for LLM-based AI.
    • scotty7928 minutes ago
      If you are into small local models I highly recommend vibe thinker. It&#x27;s a model trained specifically for reasoning. Basically a problem solver. When compared with other models, on math problems benchmarks, it&#x27;s closer to models hundred times its size than ten times its size which it beats comfortably.<p>It supports long contexts on limited VRAM and is blazing fast.<p><a href="https:&#x2F;&#x2F;github.com&#x2F;WeiboAI&#x2F;VibeThinker" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;WeiboAI&#x2F;VibeThinker</a>
  • kamranjon1 hour ago
    This seems really interesting - I was curious about this line from the website.<p>“The whole post-training stack in one CLI. Soup doctors your data pre-flight, picks the method, writes the config, derives evals from your own data, gates every save, and self-corrects reward hacking mid-run instead of just halting.”<p>How does soup auto tune the hyper parameters and make some of these more complex training decisions?
  • MakazhanAlpamys2 hours ago
    Author here.<p>The constraint everyone works around is that the frozen base has to fit in VRAM. But during LoRA the base is frozen — read, never written. It doesn&#x27;t need to live in VRAM, it needs to arrive before the matmul that uses it. So it sits in host RAM and streams into a small pool of pre-allocated VRAM buffers, one decoder layer at a time, prefetched one ahead on a dedicated CUDA stream. Peak VRAM becomes one layer instead of the whole model.<p>Measured on an RTX 3050 Laptop (4 GB, Windows): Llama-3.1-8B in NF4 at 119.6 tok&#x2F;s, 3.32 GB peak, 100% SM occupancy. Also Qwen2.5-3B with an un-quantized bf16 base at 143 tok&#x2F;s in 2.15 GB, which is CUDA OOM when trained resident on the same card. Overhead is 1.43x vs resident, measured at 0.5B — the only size on this card with a valid resident baseline, and I publish that baseline so you can check the division.<p>Most of the work wasn&#x27;t speed, it was correctness. Streaming fails silently: cut the autograd path and the loss still falls because the upper layers keep learning. So the bar was bit-exactness against a resident reference of the same numerics — max abs logit difference 0.0, across nine architecture families in two precisions, as a CI test rather than a one-off. That protocol caught a PEFT dispatch defect producing 0.94 logit divergence with byte-identical weights and adapters, no crash, no warning.<p>Not claiming anything above 8B — 14B NF4 needs ~7.5 GB page-locked against a measured 7.12 GB ceiling here, so I didn&#x27;t run it. All numbers are Windows, so pessimistic vs Linux.<p>Measurement records, including the ones I threw away: <a href="https:&#x2F;&#x2F;github.com&#x2F;MakazhanAlpamys&#x2F;Soup&#x2F;tree&#x2F;main&#x2F;benchmarks" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;MakazhanAlpamys&#x2F;Soup&#x2F;tree&#x2F;main&#x2F;benchmarks</a><p>Write-up: <a href="https:&#x2F;&#x2F;doi.org&#x2F;10.5281&#x2F;zenodo.21771064" rel="nofollow">https:&#x2F;&#x2F;doi.org&#x2F;10.5281&#x2F;zenodo.21771064</a><p>Happy to answer anything about the scheduler or the correctness protocol.