7 comments

  • spider-mario 22 minutes ago
    > Second, besides noise (bars are Wilson 95% confidence intervals, very conservative for run-to-run noise), there is little difference down to 4-bit; only the 2-bit scores a bit lower.

    Confidence intervals have nothing to do with run-to-run variation. They have little to do with anything people usually ascribe to them (https://link.springer.com/article/10.3758/s13423-015-0947-8 ), but even less with run-to-run variation (https://link.springer.com/article/10.1007/s10654-016-0149-3 misconception 22).

    • jnwatson 5 minutes ago
      Mind blown. The more I read about statistics, the less I know.
  • purpleflame1257 27 minutes ago
    There's a real hole here at Q3. A critical breakpoint here is sub 16-GB cards, which covers the 5080, 5070 Ti, 5060ti, and several other cards from this generation and the last. It would be instructive to see where the quality knee is.
    • jadbox 4 minutes ago
      Q3 XL and Q3 XS are the two I'm trying to decide on
  • quietraster 1 minute ago
    the 4-bit matching bf16 on terminal-bench is a useful data
  • bellowsgulch 21 minutes ago
    Qwen3.8 27B seems like it was clearly supposed to be a high-end consumer open-weights model, but the t/s is so low for me on my old M1 Max 64GB that I hope others are getting use out of it.

    Unfortunately, the calculus has changed and it seems cheaper to me to just use MiMo V2.5 for pennies or DeepSeek V4 Flash instead of using Qwen anymore unless I need a local model specifically for doing reverse engineering work that gets otherwise rejected.

    • spider-mario 9 minutes ago
      > Qwen3.8 27B seems like it was clearly supposed to be a high-end consumer open-weights model, but the t/s is so low for me on my old M1 Max 64GB that I hope others are getting use out of it.

      Have you tried it with MTPLX? I get around 30 tok/s with it, also on an M1 Max with 64GB.

      • SwellJoe 0 minutes ago
        [delayed]
      • Xeoncross 8 minutes ago
        Nice, which model quantization is this? Is it on huggingface?
    • Xeoncross 9 minutes ago
      I leave it running at night. No danger of burning my token subscriptions and it has hours and hours to run slowly with a manager like: github.com/kunchenguid/gnhf
  • InvectusXIV 3 minutes ago
    [flagged]