0Pricing
AI Engineering Academy · 课时

稠密检索与稀疏检索:权衡取舍

了解稠密嵌入何时会错过精确的关键词匹配、BM25 何时会错过语义改写,以及为什么将两者结合通常始终优于单独使用任何一种方法。

稠密检索与稀疏检索:权衡取舍 是 CoddyKit 上的免费 AI Engineering Academy 课时。 这是第 1 节课,共 4 节。 你可以在下方免费阅读本课时的完整内容 — 然后在浏览器中使用内置代码编辑器和全天候 AI 导师进行实践。 这是 AI Engineering Academy 学习路径的一部分,你的进度在网页和 CoddyKit 应用中同步。 AI Engineering Academy 课程共包含 4 节课。

本课时的部分内容尚未翻译,以英文显示。

Two Fundamentally Different Retrieval Signals

Modern retrieval systems rely on two distinct signals: dense retrieval encodes meaning into continuous vector spaces, while sparse retrieval counts exact term occurrences. These signals are complementary, not interchangeable. Understanding their individual strengths and weaknesses is the first step toward building a system that uses both effectively.

How Dense Embeddings Work

Dense retrieval maps both the query and each document into a high-dimensional vector using a neural encoder. Similarity is measured by cosine distance or dot product between vectors. Because the encoder was trained on large text corpora, semantically related phrases end up near each other in vector space even if they share no common words — this is the key advantage of dense retrieval.

from openai import OpenAI
import numpy as np

client = OpenAI()

def embed(text: str) -> list[float]:
    resp = client.embeddings.create(
        model='text-embedding-3-small',
        input=text,
    )
    return resp.data[0].embedding

def cosine_similarity(a, b):
    a, b = np.array(a), np.array(b)
    return float(np.dot(a, b) / (np.linalg.norm(a) * np.linalg.norm(b)))

q = embed('How do I cancel my subscription?')
d = embed('Steps to unsubscribe from the service')
print(cosine_similarity(q, d))  # high similarity despite different words

Where Dense Retrieval Fails

Dense models struggle with rare terms that were underrepresented during encoder training. A query containing a specific product model number like RTX-4090-Ti-OC, a medical drug name, or a proprietary internal identifier will often fail to match the correct document because the encoder has no learned representation for that token sequence. The vector just lands somewhere unhelpful in embedding space.

How Sparse BM25 Retrieval Works

BM25 (Best Matching 25) is a probabilistic ranking function that scores documents based on how often query terms appear in the document, normalized for document length and dampened by term frequency saturation. It produces a sparse score vector — most dimensions are zero because documents contain only a small fraction of the vocabulary.

# BM25 scoring formula (conceptual)
# score(D, Q) = sum over query terms t of:
#   IDF(t) * (tf(t,D) * (k1 + 1)) / (tf(t,D) + k1 * (1 - b + b * |D|/avgdl))

# k1 controls term frequency saturation (typically 1.2-2.0)
# b controls document length normalization (typically 0.75)
# IDF(t) = log((N - df(t) + 0.5) / (df(t) + 0.5))

# N = total documents, df(t) = documents containing term t
# tf(t,D) = frequency of t in document D, |D| = doc length, avgdl = average doc length

BM25 Strengths: Exact Terms and Jargon

BM25 excels at queries involving exact technical terms, product names, error codes, and numeric identifiers that should match precisely. A query for ORA-01017 (an Oracle error code) will rank documents containing that exact string far above documents that merely discuss database authentication in general terms. This is impossible for a dense model that has never seen that specific code.

from rank_bm25 import BM25Okapi

corpus = [
    'Oracle database ORA-01017 invalid username or password logon denied',
    'Database authentication and connection troubleshooting guide',
    'How to resolve login errors in Oracle and MySQL databases',
]

tokenized_corpus = [doc.lower().split() for doc in corpus]
bm25 = BM25Okapi(tokenized_corpus)

query = 'ORA-01017 error fix'
scores = bm25.get_scores(query.lower().split())
print(dict(zip(range(len(corpus)), scores)))
# doc 0 scores highest because it contains ORA-01017

Where BM25 Fails: Paraphrases and Synonyms

BM25 is blind to semantic paraphrasing. A document about 'automobile engine repair' will score zero for a query about 'car motor maintenance' because none of the exact words overlap. This vocabulary mismatch problem, sometimes called the lexical gap, means pure keyword search misses huge amounts of relevant content that simply uses different words to express the same idea.

from rank_bm25 import BM25Okapi

corpus = [
    'automobile engine repair and maintenance tips',
    'car motor maintenance guide for beginners',
    'vehicle powertrain service intervals',
]
tokenized = [doc.split() for doc in corpus]
bm25 = BM25Okapi(tokenized)

scores = bm25.get_scores(['car', 'motor', 'maintenance'])
print(scores)
# doc 1 scores high, doc 0 and 2 score lower despite being semantically related

Benchmark Evidence: Hybrid Wins Consistently

Benchmarks on BEIR, MS MARCO, and enterprise Q&A datasets consistently show that hybrid retrieval outperforms either dense or sparse alone by 5-15 percent on NDCG@10. The improvement is largest on datasets with a mix of factual lookups (where BM25 helps) and paraphrase queries (where dense embeddings help). No single retrieval method dominates across all query types.

Query Type Analysis: Which Retriever Wins

You can predict which retriever will perform better by analyzing the query type. Dense retrieval wins on conceptual questions, paraphrases, and broad topic queries. BM25 wins on queries containing proper nouns, version numbers, code snippets, acronyms, and rare technical terms. Hybrid always wins when query type is unknown in advance — which is almost always true in production.

# Query type heuristics
def predict_retriever_advantage(query: str) -> str:
    tokens = query.split()
    has_numbers = any(t[0].isdigit() for t in tokens)
    has_uppercase_acronyms = any(t.isupper() and len(t) > 2 for t in tokens)
    is_short = len(tokens) <= 4

    if has_numbers or has_uppercase_acronyms:
        return 'BM25 likely wins (exact terms)'
    elif is_short:
        return 'Dense likely wins (semantic matching needed)'
    else:
        return 'Hybrid recommended (mixed signals)'

Score Incompatibility Problem

Combining dense and sparse results is non-trivial because their scores are on incompatible scales. Cosine similarity produces values between -1 and 1, while BM25 produces unbounded positive scores that depend on corpus size. You cannot simply add them. The standard solution is to use rank-based fusion rather than score-based fusion — merging ranked lists instead of raw scores.

Practical Decision: When to Use Each

Use dense-only retrieval when your corpus is in a narrow domain with consistent vocabulary and you need semantic generalization across paraphrases. Use BM25-only when queries are primarily lookup-style with exact identifiers and your dataset is small enough that brute-force is feasible. Use hybrid in all production RAG systems where query types vary — the overhead is modest and the recall improvement is significant.

Performance and Infrastructure Trade-offs

Dense retrieval requires GPU-accelerated approximate nearest neighbor search or a vector database, which adds infrastructure cost. BM25 runs entirely on CPU with an inverted index and is extremely fast. Hybrid retrieval requires both infrastructure components plus a fusion step. The added complexity is justified by the recall improvement for most production use cases, but must be weighed against your infrastructure budget.

Quick Check

Test your understanding of dense versus sparse retrieval trade-offs from this lesson.

Lesson Recap

In this lesson you learned: dense retrieval captures semantic meaning but fails on rare exact terms, BM25 sparse retrieval handles exact keywords but misses paraphrases, and hybrid retrieval consistently outperforms either method alone across diverse query types. Their scores are incompatible and must be merged via rank fusion rather than score addition. Next up we implement BM25 keyword search in Python.

常见问题解答

「稠密检索与稀疏检索:权衡取舍」课时是免费的吗?

是的 — 「稠密检索与稀疏检索:权衡取舍」的完整文本可在网页上免费阅读。要进行交互式练习(内置代码编辑器和全天候 AI 导师)并解锁 AI Engineering Academy 课程的其余内容,请升级到 CoddyKit PRO。 AI Engineering Academy 课程共包含 4 节课。

「稠密检索与稀疏检索:权衡取舍」这节课中我会学到什么?

了解稠密嵌入何时会错过精确的关键词匹配、BM25 何时会错过语义改写,以及为什么将两者结合通常始终优于单独使用任何一种方法。 你通过在浏览器中直接运行的动手代码来练习 AI Engineering Academy,全天候 AI 导师会在你学习这节课的过程中回答你的问题。

学习 AI Engineering Academy 需要有经验吗?

无需任何先前经验。CoddyKit 上的 AI Engineering Academy 课程适合初学者到高级学习者,你可以从这里开始或从头开始,按照自己的节奏学习。 这是第 1 节课,共 4 节。

「稠密检索与稀疏检索:权衡取舍」课时需要多长时间?

大多数 CoddyKit 课程大约需要 5–10 分钟。每节课都很精短且互动,所以你能稳步进步,并在网页和应用中从离开的地方继续。

我能在这节 AI Engineering Academy 课中编写并运行代码吗?

能。每节 AI Engineering Academy 课都包含内置代码编辑器,你可以在浏览器中直接编写并运行真实代码,并获得即时 AI 反馈 — 无需本地设置。

此课程中的所有课时

  1. 稠密检索与稀疏检索:权衡取舍
  2. 实现 BM25 关键词搜索
  3. 使用倒数排名融合合并分数
  4. 在 Pinecone 和 pgvector 中实现混合搜索
← 返回 AI Engineering Academy