7 comments

  • augment_me2 hours ago
    Perf goes from 80% to 47% on Wikitext-2. Also no comparisons to FP4 solutions that are able to maintain or exceed perf on the same dataset 80% perf with a 4.25-4.5 big budget.<p>I think more meaningful thing here would be a hybrid solution that went down to sub-bit representations when the informational representation does not need it (for example later layers) that still maintains task performance
  • cpldcpu54 minutes ago
    I understand the obsession with low bit quantization, but it is empirically quite evident that it is not possible to compress models to less than 4 bit per weight without severe loss of capabilities².<p>It may be nice as an experiment, but it is obviously a very inefficient route for model training: spending all the flops on a saturated model only to prune its capabilties.<p>²As to why, I have seen few explanations. But the empirical evidence is there.
    • tempoponet35 minutes ago
      Influencers have latched onto the pitch that everyday people can run frontier models on an 8gb GPU while sticking it to the labs. There&#x27;s a large audience of people who haven&#x27;t had the hardware to test larger quants to see the difference.<p>They would be better served with smaller models that can reliably call tools and generate structured outputs without looping or totally hallucinating. These projects exist, but aren&#x27;t getting amplified.
      • gobdovan26 minutes ago
        &gt; These projects exist, but aren&#x27;t getting amplified.<p>Name them
  • big-chungus43 hours ago
    Can this produce a useful model? So far 1 bit quants have been less useful than smaller models that use the same memory
    • tcdent1 hour ago
      I don&#x27;t think it&#x27;s trying to be a useful implementation, but the significance they do provide is that they are able to improve on the relative loss at lower quants.<p>So, not something anyone would want to run currently, but an indicator that there is still more to squeeze out of lower precisions.<p>Trellis quantization is a far more approachable enhancement right now, but it doesn&#x27;t cross the 1-bit barrier (and perhaps doesn&#x27;t intend to).
    • GaggiX2 hours ago
      I recently found this 1.58-bit model for ASR and it&#x27;s surprising good (and very fast), that being said it&#x27;s not a LLM.<p><a href="https:&#x2F;&#x2F;huggingface.co&#x2F;moondream&#x2F;parakeet-redux" rel="nofollow">https:&#x2F;&#x2F;huggingface.co&#x2F;moondream&#x2F;parakeet-redux</a>
  • badatnames3 hours ago
    Their paper shows this comes with huge quality loss, but that doesn&#x27;t make it a negative result by any means
  • nbutton7622 hours ago
    Thought this was going to be on the original Little Bit paper, always nice to find out about a surprise sequel!
  • bArray2 hours ago
    Has anybody tested this? Are there any available computed models to test?
  • nico3 hours ago
    Has anyone tried this on apple silicon M1-5? Any benchmarks&#x2F;comps?