HN in RSCserver-reason-react
top.mdnew.mdbest.mdask.mdshow.mdjobs.md
← Back to stories

Sub-1-Bit LLM Compression via Latent Factorization

51 pointsby brainless 3 hours ago7 comments

Discussion

Loading discussion
  • badatnames · 50 minutes ago

    Their paper shows this comes with huge quality loss, but that doesn't make it a negative result by any means

  • nico · 49 minutes ago

    Has anyone tried this on apple silicon M1-5? Any benchmarks/comps?

  • big-chungus4 · 46 minutes ago

    Can this produce a useful model? So far 1 bit quants have been less useful than smaller models that use the same memory

    • GaggiX · 37 minutes ago

      I recently found this 1.58-bit model for ASR and it's surprising good (and very fast), that being said it's not a LLM. https://huggingface.co/moondream/parakeet-redux

  • nbutton762 · 28 minutes ago

    Thought this was going to be on the original Little Bit paper, always nice to find out about a surprise sequel!

  • augment_me · 22 minutes ago

    Perf goes from 80% to 47% on Wikitext-2. Also no comparisons to FP4 solutions that are able to maintain or exceed perf on the same dataset 80% perf with a 4.25-4.5 big budget. I think more meaningful thing here would be a hybrid solution that went down to sub-bit representations when the informational representation does not need it (for example later layers) that still maintains task performance

  • bArray · 5 minutes ago

    Has anybody tested this? Are there any available computed models to test?