Türkçe altyazı Turkish subtitles
0:00 Uzun bağlam ve önbellek destekli üretim, büyük dil modellerine harici bilgiye erişim sağlayan iki yöntemdir ve birbirini tamamlayan, anlaşılması değerli bir ilişki içindedir. Long context and cache augmented generation are two ways to give a large language model access to external knowledge and they actually build on each other in a way that's worth understanding. 0:13 Yani bir LLM, yani büyük dil modeli, yalnızca eğitim verisindeki bilgileri bilir. So an LLM, a large-language model, it only knows what was in its training data. 0:20 Eğer model, özel bir belgedeki veriler üzerinde akıl yürütmesi veya çeyreklik finansal sonuçlara bakması gerekiyorsa, bunu çıkarım sırasında o bilgiye erişim yolu bulması gerekir. So if it needs to reason over some data that's in a private document or maybe it needs to take a look at quarters financials, well it needs a way to access that information at inference time. 0:37 Çoğu insan RAG, yani Arama Destekli Üretim'den haberdardır. Now most people have heard of RAG, Retrieval Augmented Generation. 0:43 Aman Tanrım, RAG'ı daha önce birkaç kez ele almış gibi hissediyorum. My goodness, I feel like we've covered RAG a few times before. 0:49 RAG, bu sorunu vektör veritabanları ve gömme modelleri gibi araçları kullanan bir arama hattıyla çözer. Now RAG solves this problem with a retrieval pipeline, meaning using things like vector databases and embedding models. 0:58 Ancak CAG ve uzun bağlam, bir yapay zeka modeline bu harici bilgiyi sağlamanın iki başka yoludur. But CAG and long context are two other ways to provide an AI model with this external knowledge. 1:05 Peki bunlar nasıl çalışıyor? So how do they work? 1:07 Uzun bağlamın çok daha basit bir şeyden ibaret olduğunu söyleyebiliriz; temelde her şeyi bağlam penceresine sığdırmak demek. Well, it doesn't get much more simple than long context, which is basically to say, just stuff everything into the context window. 1:17 Bunu isteme gönderin, bu şekilde yapay zeka modeline iletin. Send it into the prompt, send it to the AI model that way. 1:21 Bu basit, ama her şey sığacak mı? Now that's simple, but will everything fit? 1:24 Bağlam pencereleri oldukça hızlı büyüyor. Well, context windows have been growing pretty fast. 1:28 Bunu hızlı bir şema ile gösterelim. So let me illustrate that with a quick diagram. 1:31 Bu eksende bağlam penceresinin boyutunu, şu eksende ise zamanı gösterelim. So on this axis here, we're gonna put the size of the context window and then this axis there, this will just be time. 1:41 Tamam, bunu haritalandıralım. Okay, so let's kind of map this out. 1:43 2020'de başlarsak, GPT-3 bin token işleyebiliyordu. So if we start in 2020, GPT-3, that could handle thousand tokens. 1:55 Bu yaklaşık 10 sayfa metne denk geliyor. That's maybe 10 pages of text. 1:59 2023'e ilerlersek, GPT-4 Turbo bunu 128.000 token'a, yani yaklaşık 300 sayfaya kadar önemli ölçüde artırdı. If we go forward to 2023, well that's where GPT-4 Turbo pushed that quite significantly to 128,000 tokens, which is roughly 300 pages. 2:14 Ardından 2024'te Google Gemini 1.5 Pro geldi. Then by 2024 Google Gemini 1.5 Pro came along. 2:22 Burada ne var? What do we have here? 2:23 İki milyon token, bu belki de 20 tam boy roman eder ve bu eğilim hâlâ yükseliyor. Two million tokens, that's maybe 20 full length novels, and the trend is still climbing. 2:32 Dolayısıyla strateji şu hale geliyor: bağlam penceresi yeterince büyükse, geri getirme hattını tamamen atlayın. So the strategy becomes, if the context window is big enough, then just skip the retrieval pipeline entirely. 2:41 Yani tüm malzemelerimizi, harici bilgi olarak sağlamak istediğimiz tüm belgeleri alıyoruz ve bunlar, büyük dil modeline göndermek istediğimiz sorguyla birlikte isteme ekleniyor ve hepsi bu bağlam penceresinde saklanıyor. So we just get all of our stuff, we get all our documents that we want to provide as external knowledge, and they go into the prompt along with the query that we want to send to the large language model and it's all stored in this context window. 3:00 Yapay zeka modeline her şeyi okumasını bırakıyoruz, çünkü artık çok büyük bir bağlam penceresine sahibiz ve ardından o modelden bir yanıt çıkıyor. We let the AI model just read the entire thing because we've got such a big context window to go from now and then a response comes out from that model. 3:12 Şimdi belirli iş yükleri için bu yöntem gerçekten çok iyi çalışıyor. Now for certain workloads this works really really well. 3:16 Modelin her şeye erişimi var, bu yüzden geri getirme adımında yanlış belgeleri çekme veya RAG'ın yapabileceği gibi ilgili bir şeyi kaçırma riski yok. The model has access to everything so there's no risk of the- retrieval step, pulling the wrong documents, or missing something relevant, like rag might do. 3:26 Ama maliyetler var. But there are costs. 3:28 Somut maliyetler. Literal costs. 3:30 Çünkü çoğu LLM API'sinin fiyatlandırması token sayısına göre ölçekleniyor. Because pricing for most LLM APIs scales with token counts. 3:34 Yani bu 200.000 token'luk bir bağlam ise, her tekil sorguda bunu ödemek zorunda kalırsınız ve bu çok hızlı bir şekilde birikebilir, gecikme de öyle. So if this is 200,000 tokens of context, you're going to have to pay that on every single query, and that can add up pretty fast, as does latency. 3:44 Büyük bir bağlam, yanıt başına daha fazla işleme süresi gerektirebilir. A large context can more processing time per response. 3:48 Ayrıca daha ince bir sorun da var, LLM'lerin "ortada kaybolma" etkisi denen bir şeyi var. Then there's a subtler problem as well, LLMs have what's called the lost in the middle effect. 3:55 Bağlam penceresinin başlangıcındaki performans genellikle oldukça güçlüdür ve sonundaki bilgiler de genellikle geri getirilebilir, So performance at the beginning of a context window, well that generally is is pretty strong, and then information at the end can generally be retrieved as well, 4:08 ama uzun bağlam penceresinin ortasına gömülü olan bilgilerde doğruluk önemli ölçüde düşebilir. but the information buried in the Middle of the long context window the accuracy can drop significantly. 4:16 Model, merkeze kıyasla kenarlara daha iyi dikkat eder. The model attends to the edges better than it does the center. 4:21 Ama belki de en büyük sorun, her sorgunun tüm bu belgeleri sıfırdan yeniden işlemesidir. But maybe the biggest issue is that every query reprocesses all of these documents from scratch. 4:28 Yani aynı belge seti hakkında 10 soru sormak, modelin bu belgeleri 10 kez okuması anlamına gelir. So 10 questions about the same set of documents mean the model reads these documents 10 times. 4:34 Bu da oldukça açık bir soruyu gündeme getiriyor. Which raises a pretty obvious question. 4:38 Peki ya model belgeleri sadece bir kez okuyup sonra onları hatırlayabilseydi? What if the model could read the documents just once and then remember them? 4:43 İşte Önbellek Destekli Üretim'in, yani CAG'ın arkasındaki fikir bu ve onu anlamak için anahtar olan şey Anahtar Değer Önbelleği olarak adlandırılan bir mekanizmadır. Well, that's the idea behind Cache Augmented Generation, or CAG, and the key to understanding it is something called Key Value Cache. 4:56 Büyük dil modelleri metin işlerken, dönüştürücünün her katmanı anahtar ve değer matrislerini hesaplar. So Key Value When a large language model processes text, each layer of the transformer computes what are called key and value matrices. 5:11 Bunlar temelde modelin çalışma belleğidir ve modelin bugüne kadar okuduğu her şeyi nasıl anladığını ve kodladığını temsil eder. And these are basically the model's working memory and they represent how the model has understood and encoded everything it's read so far. 5:19 Normalde bunlar her istekte sıfırdan hesaplanır, ancak CAG bunu bir kez yapın ve sonucu yeniden kullanın der. Now, normally these get computed fresh on every single request, but Cag says, do it once and reuse the result. 5:30 CAG'ın nasıl çalıştığını üç farklı aşamada düşünebiliriz. Now we can think about how CAG works in three different phases. 5:35 Birinci aşama, bilgi hazırlığıdır. And phase number one, that's knowledge preparation. 5:40 İlgili belgeler, modelin bağlam penceresine sığacak şekilde oluşturulur ve biçimlendirilir. So the relevant documents, they get created and formatted to fit within the model's context window. 5:45 Bu belgeler şirket politikaları, ürün dokümantasyonu veya bilgi tabanının ne olduğu her ne ise olabilir. And those documents can be whatever you like, company policies, product documentation, whatever the knowledge base is. 5:52 İkinci aşamada önceden hesaplama aşamasına geçeriz; yani model tüm bu belgeleri işler ve okuduğu her şeyin içsel temsili olan KV önbelleğini oluşturur. Then in phase two, we move on to pre-computation, which is to say, the model processes all of those documents and generates the KV cache, which is the internal representation of everything it's read. 6:07 Ardından bu önbellek bir yere kaydedilir. And then that cache gets saved somewhere. 6:09 Yani kalıcı hale getirilir. So it's persisted. 6:10 Disk veya belleğe. So to disk or to memory. 6:12 Üçüncü aşama ise aslında çıkarım aşamasıdır. And then phase number three, that's actually the inference phrase. 6:17 Burada olan şudur: bir sorgu gelir ve bu sorguyu işlemek için bir yapay zeka modeline gönderilir. So what happens here is a, you know, a query comes in and it's sent into an AI model to process that query. 6:29 Sonra model, tüm belgeleri yeniden okumak yerine, bu önceden hesaplanmış KV önbelleğini yükler, yeni soruyu ekler ve ardından cevabı üretir. Then instead of rereading all of the documents, the model just loads this pre-computed KV cache, it appends the new question and then it generates the answer coming out. 6:43 Ağır hesaplama zaten ikinci aşamada yapıldığı için, üçüncü aşama çok daha hızlı olacaktır. So the heavy computation already happened in phase two, so phase three here is going to be a lot faster. 6:50 Nitekim araştırmalar, yaklaşık 10 kat hızlanma göstermiştir. And in fact Research has showed something like a 10x speedup. 6:56 KV önbelleğinde küçük veri kümeleri kullanıldığında ve daha büyük olanlarda ise her seferinde tam bağlamı yeniden işlemeye kıyasla 40 kat gibi daha da yüksek hızlanmalar elde edilmektedir. When we use small data sets in the KV cache and even more with larger ones, something like 40X compared to reprocessing the full context every time. 7:07 CAG'ın sınırlılıkları vardır ve en belirgini, tüm bilgi tabanının hâlâ bağlam penceresine sığması gerektiğidir; kaynak belgeler değiştiğinde ise tüm KV önbelleğinin yeniden hesaplanması gerekir. Now, CAG does have limitations and the obvious one is that the entire knowledge base still has to fit within the context window and when the source documents change, well, the entire KVCache has to be recomputed. 7:24 Yani veri sık sık değişiyorsa. So if the data changes frequently. 7:27 Sürekli yeniden önbelleğe alma maliyeti, sağladığınız gecikme tasarrufunu yemeye başlar. The cost of constantly re-caching starts in, it starts to eat into any latency savings that you are making. 7:33 Yani aslında, CAG'nin en iyi çalıştığı durum, kullandığı bilgi tabanının stabil olmasıdır. So really, CAG works best when the knowledge base here that it's using is stable. 7:39 Uzun bağlam ile CAG arasındaki fark, hesaplamanın ne zaman yapıldığına bağlıdır. So the difference between long context and CAG, well, it comes down to when the computation happens. 7:45 Uzun bağlamda, model bağlam penceresine koyduğumuz her belgeyi işlemek zorundadır. So with long context, the model has to process every document that we've put into the context window. 7:54 Bunu yapması gerekir. It needs to do it. 7:55 Gelen her sorguda, her seferinde. Every time on every query that comes in. 7:59 Kurulumu basittir, yönetilmesi gereken bir şey yoktur, ancak maliyet ve gecikme yükü her çıkarım sırasında ortaya çıkar. I mean, it's simple to set up, there's nothing to manage, but the cost and the latency hit come on every inference. 8:07 CAG'da ise model aynı belgelerin tümünü işler, ancak bunu yalnızca önceden hesaplama sırasında bir kez yapar. With CAG, the model processes all of those same documents, but it only does it once during pre-computation. 8:18 Bu KV önbelleğini kaydeder ve ardından her sonraki sorgu sadece önbelleği yükleyip soruyu yanıtlar. It saves that KV cache, and then each subsequent query just loads the cache and answers the question. 8:24 İlk sorgu uzun bağlam kadar pahalıdır, ancak ikinci, onuncu veya yüzüncü sorguda CAG kendini göstermeye başlar. Now the first query is just as expensive as long context, but the second query, the tenth query, the hundredth query, that's where tag starts paying off. 8:35 Uzun bağlam da belirli bir sorgu türü için çok iyi bir seçenektir, o da şudur. Now, long context is also a really good fit for a specific type of query, and that is. 8:44 Bir şeyi yalnızca bir kez çağırdığımız sorgular; örneğin tek bir belgenin analiz edilmesi veya hakkında birkaç sorunun yanıtlanması. Queries where we just call something one time, so analyzing a single document, maybe answering a couple of questions about it. 8:51 Sadece bir kez sorgulanacak bir şey için önbellek önceden hesaplamak mantıksızdır. There's no point pre-computing a cache for something that's only going to be queried one time. 8:56 Oysa CAG, bu istekler tekrar tekrar gerçekleşiyorsa çok daha iyidir. Whereas CAG is much better if those requests happen over and over. 9:02 Tekrarlayan sorgular, stabil bir bilgi tabanı üzerinde CAG'nin uzun bağlama karşı parladığı alandır; örneğin şirket politikaları hakkında soruları yanıtlayan bir İK sohbet robotu, çünkü bilgi nadiren değişir ve önbellek geçerliliğini korur. So repeated queries are really where CAG shines against a stable knowledge base, so an HR chat bot answering questions about company policies, for example, because the knowledge doesn't change very often and the cache stays valid. 9:18 Dolayısıyla ilk sorgudan sonraki her sorgu hızlı ve ucuz olacaktır. So every query after the first one is gonna be fast and cheap. 9:22 Gerçek dünyada CAG'yi kullanılabilir kılan bir parça daha var. Now, there is one more piece to this that makes CAG a practical thing to use in the real world. 9:29 Buna istem önbellekleme denir. That is called prompt caching. 9:33 Artık büyük LLM sağlayıcıları, istem önbellekleme özelliğini API'lerinde sunmaktadır. Now major LLM providers now offer prompt cashing as a feature of their APIs. 9:37 Temel fikir, aslında bir hizmet olarak CAG'dir; bu gerçek bir kısaltma olmasa da durum budur. And the idea is essentially CAG as a service which is not a real acronym but that's what it is. 9:46 Tüm belgeleri içeren uzun bir sistem istemi gönderirsiniz ve sağlayıcı KV önbellek yönetimini arka planda halleder. So you send a long system prompt containing all the documents and the provider handles the KEV cash management behind the scenes. 9:55 Aynı istem önekinin paylaşıldığı sonraki isteklerde belge işleme tamamen atlanır ve bunun ekonomik boyutu oldukça büyüktür, çünkü önbellekten okumalar, o tokenleri sıfırdan işlemeye kıyasla %90'a varan indirimlerle çok daha ucuz olabilir. The subsequent requests that share the same prompt prefix they skip the document processing entirely, and the the economics of this are pretty significant because cash reads can come at a big discounts, something like a 90% discount compared to processing those tokens fresh. 10:15 Dolayısıyla istem önbellekleme, RAG gibi ilginç bir araştırma fikrini, geliştiricilerin önbellek altyapısını kendileri yönetmek zorunda kalmadan kullanabileceği bir şeye dönüştürüyor. So prompt caching takes what was an interesting research idea, which is CAG, and it turns it into something any developer can use without having to manage the cache infrastructure themselves. 10:28 Modellere harici bilgi verme konusunda tüm bir video çektik ve tek bir RAG tabanlı geri getirme hattı bile kurmadık. So an entire video about giving models external knowledge, and we didn't build a single retrieval pipeline using RAG. 10:38 Tüm bunların ardından RAG hâlâ gerekli mi? Whether RAG is still needed after all of this? 10:42 Bu, bu kanaldaki başka bir videonun tam olarak başlığı. Well, that's literally the title of another video on this channel. 10:45 O videoyu da kontrol edin. Go check that one out. Altyazı bilgisayarımda üretildi ve videoyla ilerler; bir satıra tıklayın, o ana atlasın. Subtitles generated on my computer; they follow the video — click a line to jump there.
Bu video, LLM'lerin harici bilgiye erişimini sağlayan CAG ve uzun bağlam tekniklerinin RAG ile karşılaştırmasını ve bağlam penceresi kapasitesindeki zaman içindeki gelişimini açıklıyor. This video explains the comparison of CAG and long-context techniques, which enable LLMs to access external information, with RAG, and the development of context window capacity over time.
Bu video youtube.com üzerinde yayımlandı; buradaki oynatıcı YouTube’undur. Türkçe özet ve altyazı AiPulse için hazırlanmıştır. This video is published on youtube.com ; the player here is YouTube’s. The Turkish summary and subtitles are prepared for AiPulse .
YouTube’da izle Watch on YouTube AI Kritik → AI Critique →