HABERNEWS
![]() Microsoft ve Hugging Face, ajanların tutarlılığını ölçen ThinkingBox kıyaslamasını yayınladıMicrosoft and Hugging Face Release ThinkingBox Benchmark to Measure Agent ConsistencyThinkingBox, yapay zeka ajanlarının tek seferlik başarılarını değil, 20 tekrarlı çalışmada veritabanı durumundaki tutarlılıklarını ölçerek güvenilirliklerini test ediyor. Bu kıyaslama, ajanların hata oranlarını ve modellerin istikrarlı performansını değerlendirmek için tasarlandı.ThinkingBox tests the reliability of AI agents by measuring the consistency of database states in 20 repeated runs, rather than their one-off successes. This benchmark is designed to evaluate the error rates of agents and the stable performance of models. Bu, AiPulse bülteni için derlenmiş bir haber özetidir — orijinal makale değil. Orijinal haber huggingface.co adresinde.This is a news summary compiled for the AiPulse newsletter — not the original article. The original article is at huggingface.co. Orijinal haberi okuRead the original article AI Kritik →AI Critique → Türkçe (otomatik çeviri) English (automatic translation) |
