多查询检索策略
通过生成用户查询的多个视角,并结合不同搜索的结果,提高检索召回率
多查询检索策略 是 CoddyKit 上的免费 LangChain / RAG / Vector DBs 课时。 这是第 1 节课,共 4 节。 你可以在下方免费阅读本课时的完整内容 — 然后在浏览器中使用内置代码编辑器和全天候 AI 导师进行实践。 这是 LangChain / RAG / Vector DBs 学习路径的一部分,你的进度在网页和 CoddyKit 应用中同步。 LangChain / RAG / Vector DBs 课程共包含 4 节课。
本课时的部分内容尚未翻译,以英文显示。
Boosting RAG with Multi-Query
Sometimes, a single search query isn't enough to find all the relevant information. Multi-query retrieval is a technique that helps your RAG system cast a wider net.
It generates several different versions of your original question. This helps ensure you don't miss important context, leading to more comprehensive answers.
Why One Query Isn't Enough
Imagine asking "What are the benefits of RAG?". A single search might only pick up documents directly matching those exact words. This can limit the context available to your LLM.
- It might miss documents using terms like "advantages of RAG" or "why use RAG".
- It could also miss related concepts that provide crucial background.
This limitation can lead to incomplete or less accurate answers from the LLM.
Expanding Your Search Horizon
Before advanced tools, people would manually brainstorm related queries to ensure a broader search. For example:
- Original: "How does RAG improve LLM accuracy?"
- Expanded: "What are RAG benefits for LLMs?", "RAG's impact on factual correctness", "How RAG reduces hallucinations".
While effective, this manual process is tedious and hard to scale for complex systems. We need an automated solution!
LLMs as Query Generators
Large Language Models (LLMs) are excellent at understanding context and generating variations of text. We can leverage an LLM to automatically create several alternative questions from an initial user query.
- This uses the LLM's natural language understanding abilities.
- It automates the query expansion process efficiently.
These diverse queries then provide multiple 'angles' for searching your knowledge base.
LangChain's MultiQueryRetriever
LangChain provides a powerful component called MultiQueryRetriever. This tool automates the entire multi-query process for you, making it easy to integrate into your RAG pipeline.
- It uses an LLM to generate diverse queries from your initial input.
- It then runs these multiple queries against your vector store.
- Finally, it combines the results into a single, comprehensive set of documents for the LLM.
MultiQueryRetriever in Action
Let's see how to set up MultiQueryRetriever in Python. You'll need an LLM and an existing retriever (e.g., from a vector store).
This example mocks a vector store for demonstration. In a real application, your vector store would be pre-populated with your documents.
from langchain_community.chat_models import ChatOpenAI
from langchain.retrievers.multi_query import MultiQueryRetriever
from langchain_community.vectorstores import FAISS
from langchain_community.embeddings import OpenAIEmbeddings
from langchain_core.documents import Document
# This is a mock setup for demonstration
# In a real app, your vector store would be populated
embeddings = OpenAIEmbeddings()
vectorstore = FAISS.from_documents(
[
Document(page_content="RAG improves LLM factual accuracy."),
Document(page_content="Retrieval Augmented Generation reduces hallucinations."),
Document(page_content="The advantages of RAG include up-to-date information."),
Document(page_content="RAG systems combine retrieval with generation."),
], embeddings
)
retriever = vectorstore.as_retriever()
# Initialize a chat model (replace with your actual LLM)
# For local testing, consider using a local LLM or mock
llm = ChatOpenAI(temperature=0) # Placeholder
# Create the MultiQueryRetriever
multi_query_retriever = MultiQueryRetriever.from_llm(
retriever=retriever, llm=llm
)
print("MultiQueryRetriever initialized successfully!")
# Example usage would follow: multi_query_retriever.get_relevant_documents(query)The Multi-Query Workflow
Here's a simplified breakdown of what MultiQueryRetriever does behind the scenes:
- Initial Query: You provide one query, e.g., "What are RAG's advantages?"
- Query Generation: The LLM takes your query and generates 2-4 alternative queries (e.g., "Benefits of RAG", "Why use RAG?", "How RAG improves LLMs").
- Parallel Retrieval: Each generated query is sent to your underlying retriever (e.g., vector store) simultaneously.
- Result Combination: All retrieved documents from these multiple searches are collected and often de-duplicated.
- Final Context: This combined set of documents is then passed to your main LLM for generating the final answer.
Merging Retrieved Documents
After multiple queries fetch documents, how do we best combine them to form the final context?
- Unique Documents: The simplest approach is to gather all documents and remove duplicates. This ensures variety without redundancy.
- Re-ranking: For more advanced control, you can apply a re-ranking model to the combined set. This model scores documents based on their relevance to the original query, ensuring the most important ones are prioritized (a topic for a future lesson!).
Why Use Multi-Query Retrieval?
Implementing multi-query strategies offers significant benefits for your RAG system:
- Improved Recall: You're more likely to find all relevant pieces of information, even if they use different phrasing.
- Richer Context: The LLM receives a broader and more diverse set of documents, leading to more comprehensive and accurate answers.
- Reduced Hallucinations: With better context, the LLM is less likely to "make things up" due to lack of information.
- Handles Ambiguity: Helps when the user's initial query is slightly ambiguous or could have multiple interpretations.
Multi-Query Check
Multi-query retrieval aims to improve retrieval by generating multiple perspectives of a user's query. Which of the following best describes the core problem it solves?
Multi-Query Recap
In this lesson, we explored Multi-Query Retrieval. We learned that a single query can often miss valuable context, and how Large Language Models (LLMs) can generate multiple, diverse queries to overcome this.
LangChain's MultiQueryRetriever automates this entire process, significantly improving the recall and richness of the context provided to your RAG system. This ultimately leads to more comprehensive and accurate answers.
常见问题解答
「多查询检索策略」课时是免费的吗?
是的 — 「多查询检索策略」的完整文本可在网页上免费阅读。要进行交互式练习(内置代码编辑器和全天候 AI 导师)并解锁 LangChain / RAG / Vector DBs 课程的其余内容,请升级到 CoddyKit PRO。 LangChain / RAG / Vector DBs 课程共包含 4 节课。
「多查询检索策略」这节课中我会学到什么?
通过生成用户查询的多个视角,并结合不同搜索的结果,提高检索召回率 你通过在浏览器中直接运行的动手代码来练习 LangChain / RAG / Vector DBs,全天候 AI 导师会在你学习这节课的过程中回答你的问题。
学习 LangChain / RAG / Vector DBs 需要有经验吗?
无需任何先前经验。CoddyKit 上的 LangChain / RAG / Vector DBs 课程适合初学者到高级学习者,你可以从这里开始或从头开始,按照自己的节奏学习。 这是第 1 节课,共 4 节。
「多查询检索策略」课时需要多长时间?
大多数 CoddyKit 课程大约需要 5–10 分钟。每节课都很精短且互动,所以你能稳步进步,并在网页和应用中从离开的地方继续。
我能在这节 LangChain / RAG / Vector DBs 课中编写并运行代码吗?
能。每节 LangChain / RAG / Vector DBs 课都包含内置代码编辑器,你可以在浏览器中直接编写并运行真实代码,并获得即时 AI 反馈 — 无需本地设置。