Nemotron, IMO 2026'da altın madalya barajını aştıNemotron Clears the Gold Medal Threshold at IMO 2026Habere gitRead the article Türkçe (otomatik çeviri) English (automatic translation) |
ÖzetSummaryBir ekip, Nemotron 3 Ultra üzerinde SFT ve RL ile iki uzmanlaşmış kontrol noktası ince ayarladı. Ardından, tamamen doğal dil tabanlı ve test zamanı hesaplama kullanan bir boru hattı (üret → doğrula → iyileştir → seç) oluşturdu. Bu sistem, IMO 2026'da 30/42 puan alarak altın madalya eşiğini aştı.A team fine-tuned two specialized checkpoints on Nemotron 3 Ultra using SFT and RL. They then built a pipeline (generate → verify → improve → select) that relies entirely on natural language and test-time compute. This system scored 30/42 at IMO 2026, surpassing the gold medal threshold. |
Neden ÖnemliWhy it mattersBu sonuç, tamamen açık kaynaklı olması nedeniyle önemli. Kontrol noktaları, eğitim verileri, kod, gönderilen çözümler ve 200 soruluk bir kıyaslama seti yayınlandı. Bu sayede olimpiyat düzeyindeki matematik akıl yürütme, kapalı laboratuvarların dışında da tekrarlanabilir hale geldi. Ayrıca, birden fazla kontrol noktası ile iteratif doğal dil aramasının, resmi kanıtlayıcıların yerini alabildiğini gösteriyor. Bu tasarım tercihi, küçük ekiplerin zor akıl yürütme görevlerinde yarışma eşiğini düşürüyor.This result is significant because it is fully open-source. The checkpoints, training data, code, submitted solutions, and a 200-question benchmark set have been released. This makes olympiad-level mathematical reasoning reproducible outside of closed labs. Furthermore, it demonstrates that iterative natural language search across multiple checkpoints can replace formal provers. This design choice lowers the competition threshold for small teams tackling difficult reasoning tasks. |
Öne ÇıkanlarHighlights
|
EleştiriCritical take'Açık tarif' ifadesi, hesaplama maliyetini hafife alıyor. Sistem, üç kontrol noktası ve ayrı bir yüksek hesaplama gerektiren seçim aşamasına dayanıyor. Bu nedenle altın madalya sonucu, model kalitesi kadar test zamanı ölçeklendirme başarısıdır. Ayrıca 42 puandan 12'sinin kaybedilmesi ve çalışmanın henüz tek bir ön baskı olması (arXiv, DOI kaydı bekleniyor), bağımsız tekrarlanabilirlik ve hakem incelemesinin hâlâ eksik olduğunu gösteriyor.The phrase 'open recipe' understates the computational cost. The system relies on three checkpoints and a separate selection phase requiring high compute. Therefore, the gold medal result is as much a success of test-time scaling as it is of model quality. Additionally, losing 12 points out of 42 and the fact that the work is currently only a preprint (arXiv, awaiting DOI registration) indicate that independent reproducibility and peer review are still lacking. |