Измерение сходства эмбеддингов
Разберитесь в метриках расстояния и сходства, лежащих в основе векторного поиска, и узнайте, как выбрать подходящую метрику.
«Измерение сходства эмбеддингов» — бесплатный урок LangChain / RAG / Vector DBs на CoddyKit. Это урок 4 из 4. Ты можешь прочитать весь урок бесплатно ниже — а потом практиковать его прямо в браузере с встроенным редактором кода и ИИ-репетитором 24/7. Это часть пути обучения LangChain / RAG / Vector DBs, и твой прогресс синхронизируется между веб-версией и приложением CoddyKit. Курс LangChain / RAG / Vector DBs содержит 4 уроков всего.
Части этого урока еще не переведены и отображаются на английском.
From Vectors to Meaning
An embedding maps text to a list of numbers in high-dimensional space. Texts with similar meaning land close together. To rank results we need a way to measure that closeness.
Cosine Similarity
Cosine similarity measures the angle between two vectors, ignoring their length. It ranges from -1 (opposite) to 1 (identical direction).
import numpy as np
def cosine(a, b):
a, b = np.array(a), np.array(b)
return a.dot(b) / (np.linalg.norm(a) * np.linalg.norm(b))
print(cosine([1, 0], [1, 1])) # ~0.707Euclidean Distance
Euclidean (L2) distance is the straight-line distance between two points. Smaller means more similar. Unlike cosine, it is sensitive to magnitude.
import numpy as np
def l2(a, b):
return np.linalg.norm(np.array(a) - np.array(b))
print(l2([0, 0], [3, 4])) # 5.0Dot Product
The dot product multiplies matching dimensions and sums them. For normalized vectors it equals cosine similarity, which is why many stores normalize first.
import numpy as np
def dot(a, b):
return float(np.array(a).dot(np.array(b)))
print(dot([1, 2, 3], [4, 5, 6])) # 32.0Normalization
Dividing a vector by its length gives a unit vector. After normalization, dot product and cosine similarity become equivalent, simplifying the math.
import numpy as np
def normalize(v):
v = np.array(v, dtype=float)
return v / np.linalg.norm(v)
print(normalize([3, 4])) # [0.6 0.8]Choosing a Metric
Most modern text embedding models are trained for cosine similarity. Use cosine unless your provider documentation recommends otherwise.
- Cosine: direction matters, length ignored
- L2: absolute position matters
- Dot: cosine on normalized data
Similarity vs. Distance
Beware the inversion: higher cosine = more similar, but higher L2 = less similar. Vector stores expose this difference, sometimes returning a score you must interpret.
Why High Dimensions Help
Embeddings often have hundreds or thousands of dimensions. More dimensions give the model room to separate subtle differences in meaning, at the cost of more storage and compute.
Setting Metric in a Store
When creating a collection you declare the metric. Many libraries default to cosine.
import chromadb
client = chromadb.Client()
col = client.create_collection(
name="docs",
metadata={"hnsw:space": "cosine"}
)Ranking Search Results
Search computes the chosen metric between the query embedding and every stored vector, then returns the top-k closest. The metric directly shapes which documents win.
query_vec = embed("refund policy")
scored = [(cosine(query_vec, d.vec), d) for d in docs]
scored.sort(reverse=True)
top3 = scored[:3]Pitfall: Mixing Models
Vectors from different embedding models live in different spaces and are not comparable. Always embed your query with the same model you used to index the documents.
Quick Check
Test your grasp of similarity metrics.
Recap
You explored how similarity is measured:
- Cosine compares direction (most common for text)
- Euclidean compares position
- Dot product equals cosine on normalized vectors
- Always query and index with the same model
Часто задаваемые вопросы
Урок «Измерение сходства эмбеддингов» бесплатный?
Да — полный текст урока «Измерение сходства эмбеддингов» бесплатно доступен здесь в веб-версии. Чтобы практиковать его интерактивно (встроенный редактор кода и ИИ-репетитор 24/7) и разблокировать остальной курс LangChain / RAG / Vector DBs, подпишись на CoddyKit PRO. Курс LangChain / RAG / Vector DBs содержит 4 уроков всего.
Чему я научусь в уроке «Измерение сходства эмбеддингов»?
Разберитесь в метриках расстояния и сходства, лежащих в основе векторного поиска, и узнайте, как выбрать подходящую метрику. Ты практикуешь LangChain / RAG / Vector DBs с помощью реального кода, который запускаешь прямо в браузере, и ИИ-репетитор 24/7 отвечает на твои вопросы во время урока.
Нужен ли мне опыт, чтобы начать LangChain / RAG / Vector DBs?
Предыдущий опыт не требуется. LangChain / RAG / Vector DBs на CoddyKit структурирован для всех уровней — от новичков до продвинутых, поэтому ты можешь начать отсюда или с самого начала и учиться в своем темпе. Это урок 4 из 4.
Сколько времени занимает урок «Измерение сходства эмбеддингов»?
Большинство уроков CoddyKit занимают около 5–10 минут. Каждый из них компактный и интерактивный, поэтому ты постоянно делаешь прогресс и продолжаешь с того же места в веб-версии и приложении.
Можно ли писать и запускать код в этом уроке LangChain / RAG / Vector DBs?
Да. Каждый урок LangChain / RAG / Vector DBs включает встроенный редактор кода, поэтому ты пишешь и запускаешь реальный код прямо в браузере и получаешь моментальную обратную связь от AI — локальная установка не требуется.
Все уроки этого курса
- Текстовые эмбеддинги
- Введение в векторные базы данных
- Хранение и извлечение эмбеддингов
- Измерение сходства эмбеддингов