查询改写与重新排序
探索优化用户查询并对检索到的文档重新排序的技术,以提高其对 LLM 的相关性。
查询改写与重新排序 是 CoddyKit 上的免费 LLM Apps in Production (RAG + Vector DB + Caching) 课时。 这是第 1 节课,共 4 节。 你可以在下方免费阅读本课时的完整内容 — 然后在浏览器中使用内置代码编辑器和全天候 AI 导师进行实践。 这是 LLM Apps in Production (RAG + Vector DB + Caching) 学习路径的一部分,你的进度在网页和 CoddyKit 应用中同步。 LLM Apps in Production (RAG + Vector DB + Caching) 课程共包含 4 节课。
本课时的部分内容尚未翻译,以英文显示。
Optimizing Queries for RAG
Welcome to advanced RAG techniques! In this lesson, we'll explore two powerful methods to make your Retrieval Augmented Generation (RAG) system even smarter: Query Rewriting and Reranking.
These techniques help ensure your LLM gets the most relevant information possible, leading to better and more accurate responses.
Why Raw Queries Fall Short
When a user asks a question, their initial query might not be perfect for searching your knowledge base. It could be:
- Too short or vague: Lacking specific keywords.
- Ambiguous: Having multiple possible meanings.
- Missing synonyms: Not using the exact terms found in your documents.
This can lead to your retriever fetching less relevant documents.
Understanding Query Rewriting
Query rewriting is the process of modifying the user's original query before it's sent to your document retriever.
The goal is to transform the query into a more effective search term that is more likely to match relevant documents in your vector database.
Techniques for Rewriting Queries
Query rewriting can involve several strategies:
- Query Expansion: Adding synonyms or related terms to broaden the search.
- Query Rephrasing: Changing the query's structure or wording to improve clarity.
- Query Decomposition: Breaking a complex, multi-part query into simpler, individual sub-queries.
Often, another LLM is used to perform these rewriting tasks.
Code: Simple Query Rewriting
Here's a conceptual Python example of how a simple query expansion might work. In a real system, an LLM would do the heavy lifting.
class QueryRewriter:
def rewrite(self, query):
# Simulate an LLM or a rule-based system
if "LLM performance" in query:
return query + " large language model efficiency optimization"
if "vector db" in query:
return query + " vector database semantic search"
return query
rewriter = QueryRewriter()
user_query_1 = "improve LLM performance"
rewritten_1 = rewriter.rewrite(user_query_1)
print(f"Original 1: {user_query_1}")
print(f"Rewritten 1: {rewritten_1}\n")
user_query_2 = "how to use vector db"
rewritten_2 = rewriter.rewrite(user_query_2)
print(f"Original 2: {user_query_2}")
print(f"Rewritten 2: {rewritten_2}")Why We Need Reranking
Even after a great initial search (perhaps with a rewritten query!), the top 'N' documents returned by your retriever might not be perfectly ordered by relevance.
The retriever's job is often to find potential matches. Reranking steps in to refine this order, ensuring the absolute best documents are at the very top.
The Reranking Process
Reranking works like this:
- Your initial retriever fetches a larger set of candidate documents (e.g., top 50).
- A specialized reranker model then takes each of these candidate documents, along with the original user query, and provides a more precise relevance score.
- The documents are then sorted again based on these new, more accurate scores.
This ensures the most relevant documents are passed to the LLM.
Specialized Reranking Models
Unlike a retriever that often uses embeddings for approximate similarity, rerankers typically use more sophisticated models, often called cross-encoders.
- Cross-encoders take both the query AND a document as input.
- They consider the interaction between the query and document terms directly.
- This allows for a much more nuanced understanding of relevance, though it's computationally more intensive, hence why it's only applied to a smaller subset of documents.
Code: Simulating Reranking
This example shows how a reranker might re-score and reorder an initially retrieved list of documents based on their relevance to the query.
class Reranker:
def rerank(self, query, documents):
# Simulate a cross-encoder model scoring documents
scores = {}
for doc in documents:
if "vector database" in doc.lower() and "fast" in query.lower():
scores[doc] = 0.95 # Highly relevant
elif "llm" in doc.lower() and "improve" in query.lower():
scores[doc] = 0.85
elif "database" in doc.lower():
scores[doc] = 0.7
else:
scores[doc] = 0.3 # Less relevant
# Sort documents by score in descending order
sorted_docs = sorted(documents, key=lambda d: scores.get(d, 0), reverse=True)
return sorted_docs
reranker = Reranker()
user_query = "How to build a fast vector database?"
initial_docs = [
"Introduction to LLMs",
"Building a scalable vector database",
"Optimizing LLM inference",
"Fast data ingestion for databases"
]
reranked_docs = reranker.rerank(user_query, initial_docs)
print(f"Original documents: {initial_docs}\n")
print("Reranked documents (most relevant first):")
for doc in reranked_docs:
print(f"- {doc}")The Power of Combination
The true power comes from combining both techniques:
- First, Query Rewriting creates a better search query.
- Then, your retriever uses this improved query to fetch a broader, more relevant set of documents.
- Finally, Reranking fine-tunes the order of these documents, ensuring the LLM receives the absolute best context to generate a response.
This multi-stage approach significantly boosts the quality and accuracy of your RAG system.
Check Your Understanding
Time to test what you've learned about optimizing RAG through query rewriting and reranking.
Recap & Next Steps
Great job! In this lesson, you learned about Query Rewriting and Reranking.
- Query Rewriting modifies the user's input to create a more effective search query.
- Reranking reorders initially retrieved documents using a more precise model to surface the most relevant ones.
Together, these techniques significantly enhance the quality and accuracy of your RAG applications. Next, we'll explore even more advanced RAG architectures!
用 AI 导师学习 LLM Apps in Production (RAG + Vector DB + Caching) — 免费
在浏览器中编写并运行真实代码,获得全天候 AI 导师的即时帮助,并在网页或应用中继续学习。
- 课程
- 12
- 课程
- 48
常见问题解答
「查询改写与重新排序」课时是免费的吗?
是的 — 「查询改写与重新排序」的完整文本可在网页上免费阅读。要进行交互式练习(内置代码编辑器和全天候 AI 导师)并解锁 LLM Apps in Production (RAG + Vector DB + Caching) 课程的其余内容,请升级到 CoddyKit PRO。 LLM Apps in Production (RAG + Vector DB + Caching) 课程共包含 4 节课。
「查询改写与重新排序」这节课中我会学到什么?
探索优化用户查询并对检索到的文档重新排序的技术,以提高其对 LLM 的相关性。 你通过在浏览器中直接运行的动手代码来练习 LLM Apps in Production (RAG + Vector DB + Caching),全天候 AI 导师会在你学习这节课的过程中回答你的问题。
学习 LLM Apps in Production (RAG + Vector DB + Caching) 需要有经验吗?
无需任何先前经验。CoddyKit 上的 LLM Apps in Production (RAG + Vector DB + Caching) 课程适合初学者到高级学习者,你可以从这里开始或从头开始,按照自己的节奏学习。 这是第 1 节课,共 4 节。
「查询改写与重新排序」课时需要多长时间?
大多数 CoddyKit 课程大约需要 5–10 分钟。每节课都很精短且互动,所以你能稳步进步,并在网页和应用中从离开的地方继续。
我能在这节 LLM Apps in Production (RAG + Vector DB + Caching) 课中编写并运行代码吗?
能。每节 LLM Apps in Production (RAG + Vector DB + Caching) 课都包含内置代码编辑器,你可以在浏览器中直接编写并运行真实代码,并获得即时 AI 反馈 — 无需本地设置。
此课程中的所有课时
- 查询改写与重新排序
- 多阶段与智能体式 RAG 模式
- 处理复杂文档结构
- 自查询与引用