No. 79
2026-10-06
yapay zeka nabzı — ham değil, demlenmişthe AI pulse — brewed, not raw
 AI KRİTİKAI CRITIQUE
Görsel, 30 sabit çift içeren bir kelime ekleme görevinde Qwen3.8-27B'nin farklı rakam uzunlukları ve rakam sayılarındaki performansını gösteren bir çubuk grafik sunmaktadır; genel doğruluk oranı %23,57'dir.

Qwen3.8-27B, sözcükle toplama testinde düşük başarı gösterdiQwen3.8-27B shows low performance in word-based addition test

Habere gitRead the article    Türkçe (otomatik çeviri) English (automatic translation)

ÖzetSummary

Simon Willison, nicelemede Q4_K_M formatındaki Qwen3.8-27B modeli üzerinde kontrollü bir yerel kıyaslama yaptı. Modelden pozitif tam sayıları toplamasını ve yalnızca yazıyla sonuç vermesini istedi. Düşünme süreci olmadan model, sayısal doğrulukta yalnızca %23,57 aldı. Küçük sayılarda %97 olan doğruluk, 10-13 haneli sayılarda %6,44'e düştü. Düşünme süreci etkinleştirildiğinde ise tek denemede doğruluk 169 sorunun 167'sinde sağlandı.Simon Willison conducted a controlled local benchmark on the Q4_K_M quantized Qwen3.8-27B model. The model was asked to sum positive integers and provide the result in text only. Without a thinking process, the model achieved only 23.57% numerical accuracy. Accuracy was 97% for small numbers but dropped to 6.44% for 10-13 digit numbers. When the thinking process was enabled, accuracy was achieved in 167 out of 169 attempts in a single run.

Neden ÖnemliWhy it matters

Bu test, standart MMLU tarzı kıyaslamalarda görünmeyen, tekrarlanabilir bir hata modunu izole ediyor: sıkı bir çıktı formatı kısıtı altında çok haneli aritmetik. Format uyumu (%96,17) ile sayısal doğruluk arasındaki keskin fark, modelin 'görevi bildiğini' ancak hesaplama işlemini gerçekleştiremediğini gösteriyor. Bu ayrım, kesin hesaplamalar için büyük dil modellerine (LLM) güvenen ajanlar geliştiren herkes için önemli. Testin tüketici sınıfı bir DGX Spark üzerinde çalıştırılması, bulut API kullanıcılarından çok yerel LLM uygulayıcıları için doğrudan ilgili bir sonuç sunuyor.This test isolates a reproducible failure mode not visible in standard MMLU-style benchmarks: multi-digit arithmetic under strict output format constraints. The sharp contrast between format compliance (96.17%) and numerical accuracy shows that the model 'knows the task' but cannot perform the calculation. This distinction is crucial for anyone developing agents that rely on Large Language Models (LLMs) for precise calculations. Running the test on a consumer-grade DGX Spark makes the findings directly relevant to local LLM implementers rather than just cloud API users.

Öne ÇıkanlarHighlights

  • Düşünme süreci olmadan %23,57 doğruluk, düşünme süreciyle 167/169 doğruluk: düşünce zinciri adımı, aritmetik işin neredeyse tamamını yapıyor.23.57% accuracy without a thinking process versus 167/169 accuracy with it: the chain-of-thought step performs nearly all of the arithmetic work.
  • Doğruluk, 1-3 haneli sayılarda %97,04'ten 10-13 haneli sayılarda %6,44'e düşüyor. Bu %90'ın üzerindeki keskin düşüş, görevin zorluğundan çok operatör uzunluğuyla paralel gidiyor.Accuracy drops from 97.04% for 1-3 digit numbers to 6.44% for 10-13 digit numbers. This drop of over 90% correlates more with operand length than with task difficulty.
  • %96,17 format uyumu, modelin 'yalnızca yazıyla' talimatına güvenilir şekilde uyduğunu ancak yine de sayıyı yanlış hesapladığını gösteriyor. Format uyumu, hesaplama doğruluğu anlamına gelmiyor.96.17% format compliance indicates the model reliably followed the 'text only' instruction but still calculated the number incorrectly. Format compliance does not imply calculation accuracy.

EleştiriCritical take

Q4_K_M nicelemesi önemli bir karıştırıcı değişken: %6,44'lük başarısızlığı 'Qwen3.8-27B' modeline atfetmek, bunun yerine ağır sıkıştırılmış 4 bitlik bir kontrol noktasına atfetmek, bulguyu abartır. Ayrıca tek denemeli düşünme çalışması (167/169), 30 denemeli düşünmesiz ızgarayla doğrudan karşılaştırılamaz. Bu nedenle 'düşünme süreci sorunu çözüyor' anlatısı, başlığın ima ettiğinden daha zayıftır.The Q4_K_M quantization is a significant confounding variable: attributing the 6.44% failure rate to the 'Qwen3.8-27B' model rather than the heavily compressed 4-bit checkpoint exaggerates the finding. Additionally, the single-run thinking test (167/169) is not directly comparable to the 30-attempt non-thinking grid. Therefore, the narrative that 'the thinking process solves the problem' is weaker than the headline implies.

Bu kritik, yerel yapay zeka modeli Qwen3.8-27B tarafından yazılmıştır.This critique was written by the local AI model Qwen3.8-27B.
Bültene dönBack to the issue