![]() Qwen3.8-27B, sözcükle toplama testinde düşük başarı gösterdiQwen3.8-27B shows low performance in word-based addition testHabere gitRead the article Türkçe (otomatik çeviri) English (automatic translation) |
ÖzetSummarySimon Willison, nicelemede Q4_K_M formatındaki Qwen3.8-27B modeli üzerinde kontrollü bir yerel kıyaslama yaptı. Modelden pozitif tam sayıları toplamasını ve yalnızca yazıyla sonuç vermesini istedi. Düşünme süreci olmadan model, sayısal doğrulukta yalnızca %23,57 aldı. Küçük sayılarda %97 olan doğruluk, 10-13 haneli sayılarda %6,44'e düştü. Düşünme süreci etkinleştirildiğinde ise tek denemede doğruluk 169 sorunun 167'sinde sağlandı.Simon Willison conducted a controlled local benchmark on the Q4_K_M quantized Qwen3.8-27B model. The model was asked to sum positive integers and provide the result in text only. Without a thinking process, the model achieved only 23.57% numerical accuracy. Accuracy was 97% for small numbers but dropped to 6.44% for 10-13 digit numbers. When the thinking process was enabled, accuracy was achieved in 167 out of 169 attempts in a single run. |
Neden ÖnemliWhy it mattersBu test, standart MMLU tarzı kıyaslamalarda görünmeyen, tekrarlanabilir bir hata modunu izole ediyor: sıkı bir çıktı formatı kısıtı altında çok haneli aritmetik. Format uyumu (%96,17) ile sayısal doğruluk arasındaki keskin fark, modelin 'görevi bildiğini' ancak hesaplama işlemini gerçekleştiremediğini gösteriyor. Bu ayrım, kesin hesaplamalar için büyük dil modellerine (LLM) güvenen ajanlar geliştiren herkes için önemli. Testin tüketici sınıfı bir DGX Spark üzerinde çalıştırılması, bulut API kullanıcılarından çok yerel LLM uygulayıcıları için doğrudan ilgili bir sonuç sunuyor.This test isolates a reproducible failure mode not visible in standard MMLU-style benchmarks: multi-digit arithmetic under strict output format constraints. The sharp contrast between format compliance (96.17%) and numerical accuracy shows that the model 'knows the task' but cannot perform the calculation. This distinction is crucial for anyone developing agents that rely on Large Language Models (LLMs) for precise calculations. Running the test on a consumer-grade DGX Spark makes the findings directly relevant to local LLM implementers rather than just cloud API users. |
Öne ÇıkanlarHighlights
|
EleştiriCritical takeQ4_K_M nicelemesi önemli bir karıştırıcı değişken: %6,44'lük başarısızlığı 'Qwen3.8-27B' modeline atfetmek, bunun yerine ağır sıkıştırılmış 4 bitlik bir kontrol noktasına atfetmek, bulguyu abartır. Ayrıca tek denemeli düşünme çalışması (167/169), 30 denemeli düşünmesiz ızgarayla doğrudan karşılaştırılamaz. Bu nedenle 'düşünme süreci sorunu çözüyor' anlatısı, başlığın ima ettiğinden daha zayıftır.The Q4_K_M quantization is a significant confounding variable: attributing the 6.44% failure rate to the 'Qwen3.8-27B' model rather than the heavily compressed 4-bit checkpoint exaggerates the finding. Additionally, the single-run thinking test (167/169) is not directly comparable to the 30-attempt non-thinking grid. Therefore, the narrative that 'the thinking process solves the problem' is weaker than the headline implies. |
