Meta, kod optimizasyonunda pekiştirmeli öğrenmenin neden başarısız olduğunu gösterdiMeta 新论文:RL 优化代码为何常失败Meta'nun yeni araştırması, doğru ve hızlı programları ödüllendirmenin tek başına işe yaramadığını, çünkü çalışma zamanının gürültülü ve seyrek bir sinyal olduğu; testler, sandbox, ödül mekanizması veya GRPO'daki küçük hataların doğruluğu bozabileceği ve bu nedenle güvenilir hız artışı sağlamak için test, zamanlama altyapısı, ödüller ve GRPO'nun birlikte yeniden tasarlanması gerektiğini ortaya koyuyor.Meta's new research shows that rewarding correct and fast programs alone is insufficient because runtime is a noisy and sparse signal; tests, sandboxes, reward mechanisms, or small errors in GRPO can corrupt accuracy, and therefore, to achieve reliable speed improvements, the testing framework, timing infrastructure, rewards, and GRPO must be redesigned together. Bu, AiPulse bülteni için derlenmiş bir haber özetidir — orijinal makale değil. Orijinal haber aihot.virxact.com adresinde.This is a news summary compiled for the AiPulse newsletter — not the original article. The original article is at aihot.virxact.com. Orijinal haberi okuRead the original article AI Kritik →AI Critique → Türkçe (otomatik çeviri) English (automatic translation) |