Breaking the 1.58-bit Barrier for Ternary LLMs

(arxiv.org)

87 points | by matt_d 2 hours ago

6 comments

  • om8 34 minutes ago
    Ternary quantization does not make any sense. Vector quantization and trellis based methods are better in this region for PTQ.
    • janalsncm 2 minutes ago
      PTQ and vector quantization aren’t used for this because part of the point of ternary LLMs is to make them faster. In a ternary LLM every weight is an add, subtract, or no-op so it is fast on CPU.

      If you’re just using a code book to reconstruct a f16 model the only savings you can get are in sending it over the wire.

    • om8 33 minutes ago
      If you want sub-2 bit llm, get one that’s already trained in higher precision, and compress it with something like YAQA/QTIP with finetuning or PV-tuning + AQLM/HIGGS
  • plqbfbv 36 minutes ago
    Very interesting, I was just exploring this to hopefully fit one of the latest quantized models in 16GB of VRAM.
  • NooneAtAll3 36 minutes ago
    This is the only time "1.58 bit" phrase makes more sense than "1 trit"

    Who knew that if you actually look at information entropy you can pack stuff better!

  • Kevcmk 59 minutes ago
    Woah. Good science.
  • kadushka 15 minutes ago
    [dead]
  • kadushka 53 minutes ago
    [flagged]