No. 30
2026-08-18
yapay zeka nabzı — ham değil, demlenmişthe AI pulse — brewed, not raw
 VİDEOVIDEO
Türkçe altyazıTurkish subtitles
  1. Vektör veritabanı nedir?What is a vector database?
  2. Şöyle derler: bir resim bin kelimeye bedeldir.Well, they say a picture is worth a thousand words.
  3. O halde bir taneyle başlayalım.So let's start with one.
  4. Eğer fark edemediyseniz, bu bir dağ manzarasında gün batımının fotoğrafı.Now in case you can't tell, this is a picture of a sunset on a mountain vista.
  5. Güzel.Beautiful.
  6. Şimdi bu bir dijital görüntü ve onu saklamak istediğimizi varsayalım.Now let's say this is a digital image and we want to store it.
  7. Bunu bir veritabanına koymak istiyoruz ve burada geleneksel bir ilişkisel veritabanı kullanacağız.We want to put it into a database and we're going to use a traditional database here called a relational database.
  8. Şimdi bu ilişkisel veritabanına bu resimle ilgili neyi saklayabiliriz?Now what can we store in that relational database of this picture?
  9. İlk olarak, gerçek resim ikili verisini veritabanına koyabiliriz; yani bu gerçek görüntü dosyası, ama aynı zamanda resim hakkında bazı temel üst verileri de saklayabiliriz.Well we can put the actual picture binary data into our database to start with, so this is the actual image file but we can also store some other information as well like some basic metadata about the picture so that would be.
  10. Dosya biçimi ve oluşturulma tarihi gibi şeyler.things like the file format and the date that it was created, stuff like that.
  11. Ayrıca bu resme manuel olarak eklenmiş etiketler de koyabiliriz.And we can also add some manually added tags to this as well.
  12. Yani 'gün batımı', 'manzara' ve 'turuncu' gibi etiketler ekleyebiliriz; bu bize resmi temel bir şekilde bulma imkanı verir, ancak resmin genel anlamsal bağlamını büyük ölçüde kaçırır.So we could say, let's have tags for sunset and landscape and orange, and that sort of gives us a basic way to be able to retrieve this image, but it kind of largely misses the images overall semantic context.
  13. Örneğin, benzer renk paletlerine sahip resimleri ya da arka planda dağ manzaraları olan görüntüleri bu bilgilerle nasıl sorgularsınız?Like how would you query for images with similar color palettes for example using this information or images with landscapes of mountains in the background for example.
  14. Bu kavramlar, yapılandırılmış alanlarda iyi temsil edilmez ve bilgisayarların veriyi saklama şekli ile insanların bunu anlama biçimi arasındaki bu kopukluğun bir adı vardır.Those concepts aren't really represented very well in these structured fields and that disconnect between how computers store data how humans understand it has a name.
  15. Bu, anlamsal boşluk olarak adlandırılır.It's called the semantic gap.
  16. Şimdi geleneksel veritabanı sorguları, örneğin 'select * where color = orange', bu tür sorgular yetersiz kalır çünkü yapılandırılmamış verinin nüanslı çok boyutlu doğasını yakalayamaz.Now traditional database queries like select star where color equals orange, it kind of falls short because it doesn't really capture the nuanced multi-dimensional nature of unstructured data.
  17. İşte burada vektör veritabanları devreye girer; veriyi matematiksel vektör gömülüleri olarak temsil eder.Well, that's where vector databases come in by representing data as mathematical vector embeddings.
  18. Vektör gömülüleri aslında bir sayı dizisidir.and what vector embeddings are, it's essentially an array of numbers.
  19. Bu vektörler, verinin anlamsal özünü yakalar; benzer öğeler vektör uzayında birbirine yakın konumlandırılırken, farklı öğeler uzak konumda bulunur ve vektör veritabanlarıyla benzerlik aramaları matematiksel işlemler olarak yapılabilir.Now these vectors, they capture the semantic essence of the data where similar items are positioned close together in vector space and dissimilar items are positioned far apart, and with vector databases, we can perform similarity searches as mathematical operations,
  20. Birbirine yakın vektör gömülüleri aramak, anlamsal olarak benzer içeriği bulmaya karşılık gelir.looking for vector embeddings that are close to each other, and that kind of translates to finding semantically similar content.
  21. Şimdi tüm çeşitlerdeki yapılandırılmamış verileri bir vektör veri tabanında temsil edebiliriz.Now we can represent all sorts of unstructured data in a vector database.
  22. Buraya ne koyabiliriz?What could we put in here?
  23. Tabii ki, dağ gün batımı gibi görüntü dosyalarını.Well image files of course like our mountain sunset.
  24. Buraya bir metin dosyası da koyabiliriz ya da hatta ses dosyalarını bile depolayabiliriz.We could put in a text file as well or we could even store audio files as well in here.
  25. Bu yapılandırılmamış veri ve bu karmaşık nesneler aslında vektör gömmelerine dönüştürülür ve bu vektör gömmeleri daha sonra vektör veri tabanına depolanır.Well this is unstructured data and these complex objects They are actually transformed into vector embeddings, and those vector embeddings are then stored in the vector database.
  26. Peki bu vektör gömmeleri nasıl görünür?So what do these vector embeddings look like?
  27. Şöyle ki, sayı dizileri vardır ve bu dizi içinde her konum bir öğrenilmiş özelliği temsil eder.Well, I said there are arrays of numbers and there are arrays of numbers where each position represents some kind of learned feature.
  28. O halde basitleştirilmiş bir örnek alalım.So let's take a simplified example.
  29. Buradaki dağ resmimizi hatırlıyor musunuz?So remember our mountain picture here?
  30. Evet, bunu bir vektör gömmesi olarak temsil edebiliriz.Yep, we can represent that as a vector embedding.
  31. Şimdi, dağ için vektör gömmesinin birinci boyutunun 0.91 olduğunu, sonraki boyutun 0.15 ve üçüncü boyutun 0.83 olduğunu varsayalım, ve böyle devam eder.Now, let's say that the vector embedding for the mountain has a first dimension of say 0.91, then let's say the next one is 0.15, and then there's a third dimension of 0.83 and kind of so forth.
  32. Tüm bunlar ne anlama geliyor?What does all that mean?
  33. Birinci boyuttaki 0.91, önemli yükseklik değişikliklerini gösterir çünkü bu dağlar.Well, the 0.91 in the first dimension, that indicates significant elevation changes because, hey, this is the mountains.
  34. İkinci boyuttaki 0.15, az sayıda kentsel öğe olduğunu gösterir; burada çok fazla bina görmüyorsunuz, bu yüzden skor düşük.Then 0.15 The second dimension here, that shows few urban elements, don't see many buildings here, so that's why that score is quite low.
  35. Üçüncü boyuttaki 0.83, gün batımı gibi güçlü sıcak renkleri temsil eder.0.83 in the third dimension, that represents strong warm colors like a sunset and so on.
  36. Diğer birçok boyut da eklenebilir.All sorts of other dimensions can be added as well.
  37. Şimdi bunu farklı bir resimle karşılaştırabiliriz.Now we could compare that to a different picture.
  38. Peki ya bu, sahilde bir gün batımı?What about this one, which is a sunset at the beach?
  39. O halde sahil örneği için vektör gömmelerine bir göz atalım.So let's have a look at the vector embeddings for the beach example.
  40. Bu da bir dizi boyuta sahip olacaktır.So this would also have a series of dimensions.
  41. İlk boyutun 0.12 olduğunu varsayalım, ardından 0.08 var ve sonunda 0.89 var; daha sonra da başka boyutlar geliyor.Let's say the first one is 0.12, then we have a 0.08, and then finally we have a 0.89 and then more dimensions to follow.
  42. Burada bazı benzerliklere dikkat edin.Now, notice how there are some similarities here.
  43. Üçüncü boyut, 0.83 ve 0.89, oldukça benzer.The third dimension, 0.83 and 0.89, pretty similar.
  44. Bunun nedeni ikisinin de sıcak renkler taşıması.That's because they both have warm colors.
  45. İkisi de gün batımı fotoğrafları, ancak ilk boyut burada oldukça farklı çünkü bir plajın yükselti değişiklikleri dağlara göre çok azdır.They're both pictures of sunsets, but the first dimension that differs quite a lot here because a beach has minimal elevation changes compared to the mountains.
  46. Bu oldukça basitleştirilmiş bir örnek.Now this is a very simplified example.
  47. Gerçek makine öğrenmesi sistemlerinde vektör gömmeleri genellikle yüzlerce hatta binlerce boyut içerir ve ayrıca söylemeliyim ki bu tür bireysel boyutlar nadiren bu kadar net yorumlanabilir özelliklere karşılık gelir, ama fikri anladınız.In real machine learning systems vector embeddings typically contain hundreds or even thousands of dimensions and I should also say that individual dimensions like this they rarely correspond to such clearly interpretable features, but you get the idea.
  48. Ve bu tüm soruyu ortaya çıkarıyor: Bu vektör gömmeleri aslında nasıl oluşturuluyor?And this all brings up the question of how are these vector embeddings actually created?
  49. Cevap, büyük veri setleri üzerinde eğitilmiş gömme modelleri aracılığıyla.Well, the answer is through embedding models that have been trained on massive data sets.
  50. Her veri türünün kullanabileceğimiz kendine özgü bir gömme modeli vardır.So each type of data has its own specialized type of embedding model that we can use.
  51. Size bazı örnekler vereceğim.So I'm gonna give you some examples of those.
  52. Örneğin, Clip.For example, Clip.
  53. Görseller için Clip'i kullanabilirsiniz.You might use Clip for images.
  54. Metinle çalışıyorsanız GloVe'yi, sesle çalışıyorsanız Wav2vec'i kullanabilirsiniz. Bu süreçler hepsi oldukça benzer.if you're working with text, you might use GloVe, and if you're working with audio, you might use Wav2vec These processes are all kind of pretty similar.
  55. Temelde, veriler birden fazla katmandan geçer.Basically, you have data that passes through multiple layers.
  56. Gömme modelinin katmanlarından geçerken, her katman giderek daha soyut özellikler çıkarır.And as it goes through the layers of the embedding model, each layer is extracting progressively more abstract features.
  57. Görseller için erken katmanlar kenarlar gibi temel şeyleri algılayabilir, daha derin katmanlara geldiğimizde ise belki tüm nesneler gibi daha karmaşık şeyleri tanırız.So for images, the early layers might detect some pretty basic stuff, like let's say edges, and then as we get to deeper layers, we would recognize more complex stuff, like maybe entire objects.
  58. Metin için erken katmanlar baktığımız kelimeleri, tek tek sözcükleri belirlerken, daha derin katmanlar bağlamı ve anlamı çözer.perhaps for text these early layers would figure out the words that we're looking at, individual words, but then later deeper layers would be able to figure out context and meaning,
  59. Bu temelde nasıl çalışır: Buradaki daha derin katmandan gelen yüksek boyutlu vektörleri alırız ve bu yüksek boyutlu vektörler genellikle girişin temel özelliklerini yakalayan yüzlerce hatta binlerce boyuta sahiptir.and how this essentially works is we take the high dimensional vectors from this deeper layer here, and those high dimensional vectors often have hundreds or maybe even thousands of dimensions that capture the essential characteristics of the input.
  60. Şimdi vektör gömmeleri oluşturuldu.Now we have vector embeddings created.
  61. Geleneksel ilişkisel veritabanlarıyla mümkün olmayan güçlü işlemler yapabiliyoruz; örneğin benzerlik araması sayesinde bir sorgu öğesine en yakın vektörleri bulup benzer öğeleri tespit edebiliyoruz.We can perform all sorts of powerful operations that just weren't possible with those traditional relational databases, things like similarity search, where we can find items that are similar to a query item by finding the closest vectors in the space.
  62. Ancak veritabanınızda milyonlarca vektör ve bu vektörler yüzlerce hatta binlerce boyuttan oluşuyorsa, sorgu vektörünüzü veritabanındaki her bir vektörle etkili ve hızlı bir şekilde karşılaştırmanız mümkün değil.But when you have millions of vectors in your database and those vectors are made up of hundred or maybe even thousands of dimensions, you can't effectively and efficiently compare your query vector to every single vector in the database.
  63. Bu çok yavaş olurdu.It would just be too slow.
  64. Bu işlemi gerçekleştirmek için bir süreç var ve buna vektör indeksleme deniyor.So there is a process to do that and it's called vector indexing.
  65. Vektör indeksleme, yaklaşık en yakın komşu (ANN) algoritmalarını kullanıyor; bu algoritmalar tam en yakın eşleşmeyi bulmak yerine, çok muhtemel en yakın eşleşmelerden bazılarını hızlıca tespit ediyor.Now this is where vector indexing uses something called approximate nearest neighbor or ANN algorithms and instead of finding the exact closest match these algorithms quickly find vectors that are very likely to be among the closest matches.
  66. Bunun için bir dizi yaklaşım bulunuyor.Now there are a bunch of approaches for this.
  67. Örneğin HNSW (Hierarchical Navigable Small World), benzer vektörleri bağlayan çok katmanlı grafikler oluştururken, IVF (Inverted File Index) vektör uzayını kümelere bölüp sadece en ilgili küme(leri) içinde arama yapıyor.For example, HNSW, that is Hierarchical Navigable Small World that creates multi-layered graphs connecting similar vectors, and there's also IVF, that's Inverted File Index, which divides the vector space into clusters and only searches the most relevant of those clusters.
  68. Bu indeksleme yöntemleri, bir miktar doğruluk kaybını büyük ölçüde artan arama hızıyla takas ediyor.These indexing methods, they basically are trading a small amount of accuracy for pretty big improvements in search speed.
  69. Vektör veritabanları, RAG (retrieval-augmented generation) adlı bir özelliğin temelini oluşturuyor; burada veritabanları belgelerin ve bilgi tabanlarının bölümlerini gömme (embedding) olarak saklıyor.Now, vector databases are a core feature of something called RAG, retrieval augmented generation, where vector databases store chunks of documents and articles and knowledge bases as embeddings and
  70. Kullanıcı bir soru sorduğunda, sistem vektör benzerliğini karşılaştırarak ilgili metin parçalarını buluyor.then when a user asks a question, the system finds the relevant text chunks by comparing vector similarity?
  71. Bu parçalar daha sonra büyük bir dil modeline (LLM) beslenerek, geri getirilen bilgilerle yanıtlar üretiliyor.and feeds those to a large language model to generate responses using the retrieved information.
  72. İşte vektör veritabanları bu şekilde çalışıyor.So that's vector databases.
  73. Hem yapılandırılmamış verileri depolamak hem de onları hızlı ve anlamsal bir şekilde geri almak için bir alan sunuyor.They are both a place to store unstructured data and a place to retrieve it quickly and semantically.
Altyazı bilgisayarımda üretildi ve videoyla ilerler; bir satıra tıklayın, o ana atlasın.Subtitles generated on my computer; they follow the video — click a line to jump there.

Vektör Veritabanı Nedir? Semantik Arama ve AI Uygulamalarını GüçlendirmeWhat is a vector database? Enhancing semantic search and AI applications

Vektör veritabanları, görseller gibi yüksek boyutlu verileri vektörlere dönüştürerek semantik bağlamı yakalar ve geleneksel veritabanlarının 'semantik boşluk' sorununu çözer.Vector databases capture semantic context by converting high-dimensional data such as images into vectors, solving the 'semantic gap' problem of traditional databases.

Bu video youtube.com üzerinde yayımlandı; buradaki oynatıcı YouTube’undur. Türkçe özet ve altyazı AiPulse için hazırlanmıştır.This video is published on youtube.com; the player here is YouTube’s. The Turkish summary and subtitles are prepared for AiPulse.
YouTube’da izleWatch on YouTube    AI Kritik →AI Critique →
Bültene dönBack to the issue