Türkçe altyazı Turkish subtitles
0:00 Yani, bir gazeteci olduğunuzu ve belirli bir konu hakkında bir makale yazmak istediğinizi hayal edin. So imagine you're a journalist and you want to write an article on a specific topic. 0:07 Şimdi bu konu hakkında oldukça iyi bir genel fikriniz var, ama biraz daha araştırma yapmak istiyorsunuz. Now you have a pretty good general idea about this topic, but you'd like to do some more research. 0:14 Bu yüzden yerel kütüphanenize gidiyorsunuz. So you go to your local library. 0:19 Şimdi bu kütüphane, birçok farklı konuda binlerce kitaba sahip. Now, this library has thousands of books on multiple different topics. 0:27 Peki, bir gazeteci olarak konuya uygun kitapları nasıl belirleyeceksiniz? But how do you know, as a journalist, which books are relevant for your topic? 0:32 O zaman kütüphaneciye gidiyorsunuz. Well, you go to the librarian. 0:34 Kütüphaneci, kitapların içeriği ve kütüphanedeki bilgiler konusunda uzmandır. Now, the librarian is the expert on what books contain, which information in the library. 0:40 Yani gazetecimiz, belirli konulardaki kitapları bulmak için kütüphaneciye sorgu gönderiyor. So, our journalist queries the librarian to retrieve books on certain topics. 0:47 Ve kütüphaneci o kitapları bulup gazeteciye geri sunar. And the librarian produces those books and provides them back to the journalist. 0:52 Ancak kütüphaneci makale yazma konusunda uzman değildir ve gazeteci de en güncel ve ilgili bilgileri bulma konusunda uzman değildir. Now, the librarian isn't the expert on writing the article, and the journalist isn't the expert on finding the most up-to-date and relevant information. 1:00 Fakat ikisinin birleşimiyle işi halledebiliriz. But with the combination of the two, we can get the job done. 1:04 Luv, bu aslında RAG yani Retrieval Augmented Generation sürecine çok benziyor; burada büyük dil modelleri, bir soruyu yanıtlamak için temel veri ve bilgi kaynaklarını sağlamak üzere vektör veritabanlarına başvurur. Luv, this sounds like a lot like the process of RAG, or Retrieval Augmented Generation, where large language models call on vector databases to provide key sources of data and information to answer a question. 1:17 Bağlantıyı göremiyorum. I'm not seeing the connection. 1:19 Biraz daha iyi anlamama yardımcı olabilir misin? Can you help me understand a little bit better? 1:21 Tabii ki. Sure. 1:22 Yani bir kullanıcı var. So we have a user. 1:25 Senin senaryonda bu gazeteci. In your scenario, it's that journalist. 1:34 Ve bir soruları var. And they have a question. 1:37 Peki, hangi tür soruları sormak istersiniz? So what types of questions would you want to ask? 1:40 Belki bunu daha çok bir iş bağlamına taşıyabiliriz. Maybe we can make this more of a business context. 1:43 Evet, diyelim ki bu bir iş analisti. Yeah, so let's say this is a business analyst. 1:45 Ve diyelim ki şu soruyu sormak istiyor: "Kuzeydoğu bölgesindeki müşterilerden Q1'deki gelir neydi?" And let's say they want to ask, "What was revenue in Q1 from customers in the northeast region?" 1:52 İşte bu senin sorunun. Right, so that's your prompt. 1:57 Tamam, bu kullanıcı hakkında birkaç soru var. Okay, so a couple of questions on that user. 1:59 Kişi olması gerekiyor mu yoksa başka bir şey de olabilir mi? Does it have to be a person or could it be something else too. 2:02 Evet. Bu mutlaka bir kullanıcı olmak zorunda değil. Yeah. So this doesn't necessarily have to be a user. 2:05 Bir bot da olabilir ya da başka bir uygulama da. It could be a bot or it could be another application. 2:11 Konuştuğumuz soru bile, "Kuzeydoğu bölgesinden Q1'deki gelirimiz neydi?" Even the question that we're talking about, "What was our revenue in Q1 from the northeast?" 2:16 Bu sorunun ilk kısmı, genel bir LLM'in anlaması oldukça kolay. You know, the first part of that question, it's pretty easy for, you know, a general LLM to understand, right? 2:22 Gelirimiz neydi? What was our revenue? 2:23 Ancak ikinci kısmı, Q1'de kuzeydoğu bölgesindeki müşterilerden. But it's that second part in Q1 from customers in the northeast. 2:28 Bu LLM'lerin eğitildiği bir şey değil. That's not something that LLMs are trained on, right. 2:30 İşimize çok özgü ve zaman içinde değişiyor. It's very specific to our business and it changes over time. 2:35 Bu yüzden bu kısmı ayrı ayrı ele almamız gerekiyor. So we have to treat those separately. 2:37 Peki, isteğin o bölümünü nasıl yönetiriz? So how do we manage that part of the request? 2:41 Kesinlikle. Exactly. 2:43 Belirli bir soruya cevap vermek için potansiyel olarak birden fazla veri kaynağına ihtiyacımız olacak. You'll need multiple different sources of data potentially to answer a specific question, right? 2:48 Bu soru bir PDF'den, başka bir iş uygulamasından ya da bazı görsellerden gelebilir; doğru veriye ihtiyacımız var. Whether that's maybe a PDF, or another business application, or maybe some images, whatever that question is, we need the appropriate data in order to provide the answer back. 3:01 Hangi teknoloji bu verileri birleştirip LLM'imizde kullanmamıza olanak tanır? What technology allows us to aggregate that data, and use it for our LLM. 3:07 Bu verileri alıp bir vektör veritabanına koyabiliriz. Yeah, so we can take this data and we can put it into what we call a vector database. 3:15 Bir vektör veritabanı, yapılandırılmış ve yapılandırılmamış verilerin matematiksel bir temsili olup, bir dizi içinde gördüklerimize benzer. a vector database is a mathematical representation of structured and unstructured data similar to what we might see in an array. 3:24 Anladım, ve bu diziler makine öğrenimi ya da üretken yapay zeka modelleri için sadece temel yapılandırılmamış veriye göre daha uygun ve anlaşılır. Gotcha, and these arrays are better suited or easier to understand for machine learning or generative AI models versus just that underlying unstructured data. 3:34 Tam olarak. Exactly. 3:35 Vektör veritabanımızı sorguluyoruz, değil mi? We query our vector database, right? 3:37 Ve bize, talep ettiğimiz ilgili verileri içeren bir embedding döner. And we get back an embedding that includes the relevant data for which we're prompting. 3:45 Sonra bunu tekrar orijinal soruya ekliyoruz, değil mi? And then we include it back into the original prompt, right? 3:47 Evet, tam olarak. Yeah, exactly. 3:49 Bu, soruya geri beslenir. That feeds back into the prompt. 3:51 Ve bu noktaya geldiğimizde denklemin diğer tarafına geçeriz; yani büyük bir dil modeline. And then once we're at this point, we move over to the other side of the equation, which is a large language model. 3:58 Anladım, yani vektör embeddinglerini içeren bu sorgu şimdi büyük dil modeline beslenir ve model, orijinal sorumuzun cevabını üretir. Gotcha, so that that prompt, that includes the vector embeddings now are fed into the large language model, which then produces the output with the answer to our original question 4:12 Kaynağa dayalı, güncel ve doğru verilerle. with sourced, up-to-date and accurate data. 4:15 Tam olarak. Ve bu, onun kritik bir yönüdür. Exactly. And that's a crucial aspect of it. 4:18 Yeni veriler vektör veritabanına geldikçe, Q1 performansıyla ilgili sorularınıza geri dönen şeyler güncellenir; gelen yeni verilerle embedding'ler de yenilenir. As new data comes in to this vector database, where things that are updated back to your relevant question around performance in Q1, as new data comes in, those embeddings are updated. 4:30 Dolayısıyla aynı soru ikinci kez sorulduğunda, LLM'ye daha ilgili veri sağlarız ve model cevabı üretir. So when that question is asked the second time, we have more relevant data in order to provide back to the LLM who then generates the output in the answer. 4:39 Tamam, çok güzel. OK, very cool. 4:40 Shawn, bu çok fazla kütüphaneci ve gazetecimizle yaptığım orijinal benzetmeye benziyor, değil mi? So Shawn, this sounds a lot like my original analogy there with the librarian and our journalist, right? 4:46 Gazeteci, kütüphanedeki bilgilerin doğru ve güvenilir olduğuna güvenir. So the journalists trusts that the information in the library is accurate and correct. 4:52 Şimdi gördüğüm zorluklardan biri, kurumsal müşterilerle konuşurken bu tür teknolojiyi müşteri odaklı ve iş kritik uygulamalara entegre etmeye dair endişeleri. Now, one of the challenges that I see is when I'm talking to enterprise customers is they're concerned about deploying this kind of technology into customer facing, business critical applications. 5:04 Yani uygulama geliştiriyorlar, müşteri siparişlerini alıyorlar, iadeleri işliyorlarsa; bu teknolojilerin hayal ürünü sonuçlar ya da hatalı veriler üretebileceğinden endişe ediyorlar, değil mi? So if they're building applications, taking customer orders, processing refunds, they're worried that these kinds of technologies can produce hallucinations or inaccurate results, right? 5:16 Ya da bir tür önyargıyı sürdürmek. Or perpetuate some kind of bias. 5:18 Bu endişeleri hafifletmek için neler yapılabilir? What are some things that can be done to help mitigate some of these concerns? 5:23 Bu harika bir nokta, Luv. That brings up a great point Luv. 5:24 Bu taraftan ve diğer taraftan gelen veriler, prompt oluşturup yanıt aldığımızda çıktımız için son derece önemlidir. Data that comes in on this side, but also on this side, is incredibly important to the output that we get when we go to make that prompt and get that answer back. 5:33 Yani gerçekten doğru: "Çöp girer, çöp çıkar", değil mi? So it really is true: "Garbage in and garbage out", right? 5:37 Bu yüzden vektör veritabanına iyi veri girdiğinden emin olmamız gerekiyor. So we need to make sure we have good data that comes into the vector database. 5:40 Verinin temiz, yönetişimli ve düzgün bir şekilde yönetildiğinden emin olmalıyız. We need to make sure that data is clean, governed and managed properly. 5:45 Anladım, yani duyduğum şey şu: yönetişim ve veri yönetimi vektör veritabanı için elbette çok önemli, değil mi? Gotcha, so what I'm hearing is that things like governance and data management are of course crucial to the vector database, right? 5:59 Yani modelin içine akıp giden gerçek bilginin (örneğin konuştuğumuz örnek promptta iş sonuçları) yönetişimli ve temiz olduğundan emin olmak, aynı zamanda büyük dil modeli tarafında da kritik. So making sure that the actual information that's flowing through into the model, such as the business results in the sample prompt we talked about is governance and clean, but also crucially, on the large language model side, 6:13 Büyük bir dil modelinin kara kutu yaklaşımını kullanmadığından emin olmamız gerekiyor, değil mi? we need to make sure that we're not using a large language model that takes a black box approach, right? 6:19 Yani eğitilirken kullanılan temel veriyi aslında bilmediğimiz bir model. So, a model where you don't actually know what is the underlying data that went into training it, right? 6:25 Orada herhangi bir fikri mülkiyet olup olmadığını bilmiyorsun. You don't know if there's any intellectual property in there. 6:28 Oradaki yanlışlıkları ya da çıktılarında önyargıyı sürdürecek veri parçalarını bilmiyorsun. You don't know if there's inaccuracies in there or you don't know if there are, pieces of data that will end up perpetuating bias in your output results. 6:37 Değil mi? Right? 6:37 Bir işletme olarak ve marka itibarını korumaya çalışan bir şirket olarak, şeffaf bir şekilde eğitilmiş LLM'ler kullandığımızdan emin olmak son derece kritik. So as a business, and as a business that's trying to manage and uphold their brand reputation, it's absolutely critical to make sure that we're taking an approach that uses LLMs that are transparent in how they were trained, 6:55 Ve orada olması gereken verinin dışına çıkan yanlışlıklar ya da veri olmadığından %100 emin olabiliriz, değil mi? and we can be 100% certain that there aren't any inaccuracies or data that's not supposed to be in there, right? 7:05 Evet, tam olarak. Yeah, exactly. 7:06 Özellikle bir marka olarak doğru cevapları almak çok önemli. It's incredibly important, especially as a brand, that we get the right answers. 7:10 Etki sonuçlarını gördük ve özellikle "Q1'de gelirimiz neydi?" sorusuna geri dönersek, We've seen the results of impact, and especially back to our original question around "what was our revenue in Q1", right? 7:17 Bu sorunun bir LLM'imizden gelen sonuçlarla etkilenmesini istemiyoruz. We don't want that to be impacted by the results of a question that comes from, you know, that prompts one of our LLMs. 7:24 Kesinlikle, kesinlikle. Exactly, exactly. 7:25 Çok güçlü bir teknoloji. So very powerful technology. 7:26 Ama beni kütüphaneye geri götürüyor. But it makes me think back to the the library. 7:30 Gazetecimiz ve kütüphanemiz, ikisi de kütüphanedeki veri ve kitaplara güveniyor. Our journalist and librarian, they both trust the data and the books that are in the library. 7:34 İş dünyasında bu tür üretken yapay zeka kullanım senaryolarını oluştururken de aynı güvene sahip olmalıyız. We have to have that same kind of confidence when we're building out these types of generative AI use cases for business as well. 7:40 Tam olarak, Luv. Exactly, Luv. 7:41 Dolayısıyla yönetişim, yapay zeka ve ayrıca veri ile veri yönetimi bu süreçte son derece önemli. So governance, AI, but also data and data management are incredibly important to this process. 7:48 En iyi sonucu elde etmek için üçünün de gerekli. We need all three in order to get the best result. Altyazı bilgisayarımda üretildi ve videoyla ilerler; bir satıra tıklayın, o ana atlasın. Subtitles generated on my computer; they follow the video — click a line to jump there.
RAG, büyük dil modellerinin vektör veritabanlarından ilgili bilgileri çekerek sorulara yanıt vermesini sağlayan bir yöntemdir. RAG is a method that enables large language models to answer questions by retrieving relevant information from vector databases.
Bu video youtube.com üzerinde yayımlandı; buradaki oynatıcı YouTube’undur. Türkçe özet ve altyazı AiPulse için hazırlanmıştır. This video is published on youtube.com ; the player here is YouTube’s. The Turkish summary and subtitles are prepared for AiPulse .
YouTube’da izle Watch on YouTube AI Kritik → AI Critique →