17 comments

  • alentred 0 minutes ago
    I would be very interested in is a similar benchmark for *KV cache* quantizations.

    I use Qwen3.8 27B Q4_K_M for coding sometimes and therefore need a relatively long context. I settled on q8_0 because it is the only way to fit the model + 100k tokens into 24GB VRAM, but still wonder what am I loosing in quality, and what other options are there.

  • spider-mario 2 hours ago
    > Second, besides noise (bars are Wilson 95% confidence intervals, very conservative for run-to-run noise), there is little difference down to 4-bit; only the 2-bit scores a bit lower.

    Confidence intervals have nothing to do with run-to-run variation. They have little to do with anything people usually ascribe to them (https://link.springer.com/article/10.3758/s13423-015-0947-8 ), but even less with run-to-run variation (https://link.springer.com/article/10.1007/s10654-016-0149-3 misconception 22).

    • stared 4 minutes ago
      Point taken, but there is a much more fundamental issue with it - and precisely why I wrote "very conservative".

      It is a different problem if we pick two sets from the same data distribution, A and B, and first we have a score on A, then on B. Here we re-run on precisely the same set of Terminal Bench 2.1 problems. It may be that results are so random between runs that each single task has the same probability in a Bernoulli distribution. But more likely, many problems are easy (i.e. each run will solve them consistently), many are too hard (i.e. no run is going to solve them) and only a fraction is somehow in between.

      Maybe there is some good trick to find a proper distribution, but to my knowledge, we would need to run it at least two times on TB2.1 to get any more educated estimates. That said, I am open to new ideas.

    • maCDzP 3 minutes ago
      Thank you for these, coz I learned a lot! Great that they are open access.
    • jnwatson 1 hour ago
      Mind blown. The more I read about statistics, the less I know.
      • exogenousdata 1 hour ago
        “There are three kinds of lies: Lies, damned lies and statistics.” - Mark Twain (attributed but unsubstantiated to Benjamin Disraeli)
    • fr2029 1 hour ago
      the 2nd derivate of shannon covariance of noise begs to differ
  • sharmajai 52 minutes ago
    This confirms a theory I have to explain the minimal loss in quality when using lower quants (I use IQ3_XXS with an 8-bit KV cache) and the XHIGH (default) thinking level.

    It's well-known that while quantization affects the sampling probability distribution (given the same context, which next token is the most probable), Qwen 3.8 27b seems to offset that by just thinking more and as a result eventually finishing the task (benchmark or otherwise).

    So as long as the thinking (albeit longer) is sound, this leads to the same success rate (as shown in the article) but potentially at the cost of more tokens and hence more time.

    I think it'll be further useful to chart each quantization's used tokens as well, in addition to the success rate.

    Thanks for doing and sharing the research!

    • seemaze 45 minutes ago
      As they say, time is money.

      In the age of the rampocalypse, the peasants may not have a choice between the two.. time it is!

      • chmod775 7 minutes ago
        Smaller models are also generally faster, so thinking "more" may not matter and may even come out ahead.
      • conmod278 8 minutes ago
        Computer science has known the tradeoffs between memory and compute since ages ago. The same could be reflected here.
    • anon291 45 minutes ago
      I personally think thinking is basically variable but rate precision. If you are in a 4bit mode but need 2x as many tokens you're just doing fp8 with hoops( of course 4bit multiply is faster)
      • kennywinker 17 minutes ago
        Fair enough mental model, except my GPU can’t load the 8bit version and paging from disk makes it way more than 1/2 speed.
    • lowbloodsugar 39 minutes ago
      If it digs itself into a hole, try low or medium. In the rust coding benchmarks (on my machine) it did better on low and medium because xhigh never finished.
  • purpleflame1257 2 hours ago
    There's a real hole here at Q3. A critical breakpoint here is sub 16-GB cards, which covers the 5080, 5070 Ti, 5060ti, and several other cards from this generation and the last. It would be instructive to see where the quality knee is.
    • civvv 1 hour ago
      Running Q3 on my AMD RX 9070XT. 32k context and 32/TPS. Apart from the context window preventing it from doing any large tasks, this thing is seriously powerful. I could probably push it to 64k context. Local open models are the future, and I am definitely getting a more powerful card. Very fun!
      • Forgeties79 38 minutes ago
        What are you offloading to ram (or even CPU)? I’m using a 9080 (not XT) and having trouble with context/token rates
      • slim 1 hour ago
        Running Q3 on 5060ti with 64k context. It runs great
    • dofm 1 hour ago
      There is an interesting new dynamic 3 bit quantisation I have been meaning to test:

      https://huggingface.co/ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF

      Luke of Luke’s Dev Lab on YouTube had a look at it. It seems to outperform the typical 3-bit quantisation but whether it outperforms the new Unsloth dynamic I don’t know.

    • jadbox 1 hour ago
      Q3 XL and Q3 XS are the two I'm trying to decide on
  • Farmadupe 1 hour ago
    hmm, assuming that this article is part written by claude and part human-written, can anyone help me find a rule of thumb for "how to know if the article is worth reading"?

    Because on the one hand, the prose and the presentation is painful (narrating irrelevant points, nonlinear X-axes, ambiguous chart labels, etc etc),

    But on the other hand, the result that I'm assuming the author means to communicate ("on these evals, generation quality seems fairly good") sounds worthwhile to share?

    Because I really struggle with this question at the moment. Am I allowed to draw an adverse inference that "if the writeup presents irrelevant text side by side with the data, then this may be a sign that the author does not understand the task that they are attempting to write up"?

    • vardalab 22 minutes ago
      Just ask your freaking agent to read it for you and extract the information. That's what it's all about. Why would I be reading these articles other than information?
      • kennywinker 11 minutes ago
        Claude, read Love in the Time of Cholera for me and summarize the information contained. You are an expert book reader and understander. Make no mistakes.
    • stared 37 minutes ago
      I wrote this blog post myself, with AI for proofreading (typos and grammar, but not style).

      So, if there are irrelevant remarks, these are mine. :)

      Charts are vibe-coded - but it took quite a bit of hand-holding to get something decent. And the logarithmic scale for model size is my conscious choice.

    • cogman10 1 hour ago
      IMO, whether or not an LLM was used in the writing process doesn't really matter and I think it's a bit annoying that articles are being dismissed out of hand because of that.

      The line is "Is this an interesting and accurate article that concisely makes it's case".

      LLMs love to burn paragraphs writing about nothing which is why it's generally poor writing. Humans can do the same thing if they are trying to make very little information feel more substantial.

      I say, stop trying to determine if an LLM was used and start judging based on your subjective measure that you'd have used before LLMs became widespread.

      • Farmadupe 1 hour ago
        (If it helps, I ask my own question of myself too -- I mostly don't write code by hand any more as I find that an LLM writes it faster and with less bugs -- Is that therefore proof that my time was never worth my paycheck? I hope not but at the same time I would actually be proud if I had got away with being an accidental charlatan/fraudster at my employer's expense during my entire career)

        -----

        Similarly, if what I said really is true, I would be implying that LLMs are charlatan/fraudster detectors (to some statistical level). And I refuse on principle to believe that that is actually the case.

      • dofm 1 hour ago
        My main problem — which I am sure being middle-aged compounds — is that I struggle to retain information that an LLM has written or produced. I cannot explain why but it is a consistent problem.

        In a week’s time I might remember the substance of your comment and some of its shape as a matter of course. Nothing LLM-written that I see today will stick, no matter how curated it was.

        • cogman10 56 minutes ago
          Perhaps it's just a bias? You are already negatively biased against LLM writing so you disregard stuff you read when you suspect it's an LLM. This could also be a selection bias. It may be that you generally struggled to retain information but you are more aware of it when LLMs are involved.
      • lowbloodsugar 8 minutes ago
        Sure, but if you're going to publish it, at least run it through an edit prompt and tell it to remove clickbait "Its not X, its Y" rubbish. Like literally calling them clickbait has given me better results. Interestingly, I have a lot less trouble with the first draft with Qwen than with Opus.
      • JSR_FDED 1 hour ago
        Except that wasting the reader’s time became a lot easier with LLMs.
        • cogman10 59 minutes ago
          Yes, LLMs make it a lot easier to produce a lot more garbage. That's not an exception to my point. If something is well written then it doesn't waste the reader's time.
    • clircle 1 hour ago
      I think the advice is the same regardless of AI use: read articles written by authors that have a history of high quality writing.
    • JSR_FDED 1 hour ago
      You don’t need anyone’s permission. You have only so much attention, why spend it wading through slop?
  • kouteiheika 1 hour ago
    Note that these quants are not quantized uniformly, so 4-bit isn't actually a "true" 4-bit here, so these observations won't necessarily hold up to other quants which might be done differently.
    • wgd 11 minutes ago
      It looks like they tested Q4_K_M which should be just the standard K-quant without any imatrix calibration. The smaller ones are indeed dynamic though.
  • syntaxing 56 minutes ago
    I’m more curious how each 4 bit quant compares. It seems like NVFP4 outperforms Q4_K_M in terms of speed and top 1 but is only good for expensive Nvidia cards
  • dvh 1 hour ago
    Could this be used to estimate how many fingers LLM have?
  • rvba 32 minutes ago
    Those benchmarks are very interesting.

    But is there any model that actually works in a decent way at quantization of 1?

  • bellowsgulch 1 hour ago
    Qwen3.8 27B seems like it was clearly supposed to be a high-end consumer open-weights model, but the t/s is so low for me on my old M1 Max 64GB that I hope others are getting use out of it.

    Unfortunately, the calculus has changed and it seems cheaper to me to just use MiMo V2.5 for pennies or DeepSeek V4 Flash instead of using Qwen anymore unless I need a local model specifically for doing reverse engineering work that gets otherwise rejected.

    • spider-mario 1 hour ago
      > Qwen3.8 27B seems like it was clearly supposed to be a high-end consumer open-weights model, but the t/s is so low for me on my old M1 Max 64GB that I hope others are getting use out of it.

      Have you tried it with MTPLX? I get around 30 tok/s with it, also on an M1 Max with 64GB.

      • SwellJoe 1 hour ago
        Even at 30 t/s, 3.8 thinks so long, even on medium, it still takes 3x or more longer than any cloud model, in my testing.
      • Xeoncross 1 hour ago
        Nice, which model quantization is this? Is it on huggingface?
      • bellowsgulch 36 minutes ago
        Thanks, man! I’ll go use that now that I know. llama-server the last time I used it for inference with this model wasn’t able to produce work fast enough to reach those numbers.
    • Xeoncross 1 hour ago
      I leave it running at night. No danger of burning my token subscriptions and it has hours and hours to run slowly with a manager like: github.com/kunchenguid/gnhf
    • sroussey 1 hour ago
      Have you tried https://huggingface.co/prism-ml/Bonsai-27B-mlx-1bit ? PrismML is the only people i am aware of doing 1bit that is decent.
    • ThrowawayTestr 1 hour ago
      I treat it like image gen. Send a prompt then come back in 40 minutes.
  • quietraster 1 hour ago
    the 4-bit matching bf16 on terminal-bench is a useful data
  • zrail 1 hour ago
    I've been running Unsloth IQ3_S on my 5060ti with mmproj offloaded, getting 600-1000 prefill and 30-50 tg with this config:

           /data/llm/llama.cpp/build/bin/llama-server
            --threads 4
            --threads-batch 8
            --batch-size 4096
            --ubatch-size 256
            --port 9999
            --temp "1.0"
            --top-p "0.95"
            --top-k "20"
            --min-p "0.0"
            --presence-penalty "0.0"
            --reasoning auto
            --reasoning-preserve
            --reasoning-budget 4096
            --gpu-layers-draft all
            --spec-type draft-mtp,ngram-map-k4v,ngram-mod
            --spec-draft-n-max 3
            --spec-draft-p-min 0.75
            --spec-ngram-mod-n-match 24
            --spec-ngram-mod-n-min 4
            --spec-ngram-mod-n-max 16
            --spec-ngram-map-k4v-size-n 8
            --spec-ngram-map-k4v-size-m 16
            --spec-ngram-map-k4v-min-hits 1
            --n-gpu-layers all
            --ctx-size 131072
            --repeat-penalty 1.0
            --jinja
            --metrics
            --model /data/llm/models/unsloth/Qwen3.8-27B-UD-IQ3_S.gguf
            --chat-template-file /data/llm/models/qwen3.6-chat-template.jinja
            --fit off
            --flash-attn on
            --cors-origins localhost
            --mmproj /data/llm/models/unsloth/Qwen3.8/mmproj-BF16.gguf
            --no-mmproj-offload
            --parallel 1
            --kv-unified
            --cache-type-k q4_0
            --cache-type-v q4_0
            --cache-type-k-draft q4_0
            --cache-type-v-draft q4_0
  • john_rood 1 hour ago
    [flagged]
  • InvectusXIV 1 hour ago
    [flagged]
  • dotinvictim 1 hour ago
    local llm don't make sense currently consumer compute is not upto mark it may take atleast 7 more years to be usable
    • kennywinker 3 minutes ago
      It literally is usable now. A 5060 for $800 can run qwen3.8-27b 4bit at >40t/s, and the model beats opus 4.6 (max).