查询与答案生成
开发处理用户查询、检索相关上下文并使用 LLM 综合生成答案的逻辑
查询与答案生成 是 CoddyKit 上的免费 LangChain / RAG / Vector DBs 课时。 这是第 2 节课,共 4 节。 你可以在下方免费阅读本课时的完整内容 — 然后在浏览器中使用内置代码编辑器和全天候 AI 导师进行实践。 这是 LangChain / RAG / Vector DBs 学习路径的一部分,你的进度在网页和 CoddyKit 应用中同步。 LangChain / RAG / Vector DBs 课程共包含 4 节课。
本课时的部分内容尚未翻译,以英文显示。
Querying RAG: The Answer Flow
After integrating RAG components, the next step is to use them to answer user questions. This lesson covers the full process from a user's query to a generated answer.
We'll focus on the 'query-time' logic: how your system takes a question, finds relevant context, and synthesizes a coherent response using an LLM.
Understanding the User Query
A RAG system starts with a user's question, just like a search engine. This raw input is the trigger for the entire process.
- It defines what information needs to be retrieved.
- It guides the LLM on what kind of answer to generate.
No special formatting is typically needed at this initial stage; it's just plain text.
Setting Up Your Retriever
To find relevant documents, you need a retriever. This component knows how to query your vector store. You typically obtain it from your VectorStore instance.
The as_retriever() method creates this component, and you can configure parameters like k (number of top documents to fetch).
from langchain_community.vectorstores import InMemoryVectorStore
from langchain_core.documents import Document
from langchain_core.embeddings import Embeddings
from typing import List
# A simple mock for embeddings
class MockEmbeddings(Embeddings):
def embed_documents(self, texts: List[str]) -> List[List[float]]:
return [[i * 0.1] * 10 for i in range(len(texts))]
def embed_query(self, text: str) -> List[float]:
return [0.5] * 10
def main():
# Create a dummy vector store with some content
embeddings = MockEmbeddings()
docs = [
Document(page_content="The capital of France is Paris.", metadata={"source": "wiki"}),
Document(page_content="Eiffel Tower is in Paris, France.", metadata={"source": "travel"})
]
vectorstore = InMemoryVectorStore.from_documents(docs, embeddings)
# Create a retriever from the vector store
retriever = vectorstore.as_retriever(search_kwargs={"k": 1})
print("Retriever created successfully!")
if __name__ == "__main__":
main()Fetching Contextual Documents
Once you have a retriever, you can invoke it with the user's query. It will perform a similarity search in your vector store and return the most relevant Document objects.
These documents form the context that will be passed to the LLM.
from langchain_community.vectorstores import InMemoryVectorStore
from langchain_core.documents import Document
from langchain_core.embeddings import Embeddings
from typing import List
class MockEmbeddings(Embeddings):
def embed_documents(self, texts: List[str]) -> List[List[float]]:
return [[i * 0.1] * 10 for i in range(len(texts))]
def embed_query(self, text: str) -> List[float]:
return [0.5] * 10
def main():
embeddings = MockEmbeddings()
docs = [
Document(page_content="The capital of France is Paris.", metadata={"source": "wiki"}),
Document(page_content="Eiffel Tower is in Paris, France.", metadata={"source": "travel"}),
Document(page_content="London is the capital of the UK.", metadata={"source": "wiki"})
]
vectorstore = InMemoryVectorStore.from_documents(docs, embeddings)
retriever = vectorstore.as_retriever(search_kwargs={"k": 2})
user_query = "What is the capital of France?"
retrieved_docs = retriever.invoke(user_query)
print(f"Retrieved {len(retrieved_docs)} documents:")
for doc in retrieved_docs:
print(f"- {doc.page_content[:50]}...")
if __name__ == "__main__":
main()Preparing Context for the LLM
LLMs usually prefer a single string of text as context. The retrieved Document objects need to be combined into a coherent format.
A common approach is to concatenate their page_content fields, perhaps with separators, and include source metadata if desired.
- Ensures all context fits within the LLM's token window.
- Presents a clean input for the LLM to reason over.
Crafting the RAG Prompt
The prompt template is crucial. It instructs the LLM on how to use the provided context to answer the user's question. It typically includes placeholders for both the context and the question.
A well-designed prompt guides the LLM to be factual and avoid hallucination.
from langchain_core.prompts import ChatPromptTemplate
def main():
# Define a RAG-specific prompt template
rag_prompt = ChatPromptTemplate.from_messages([
("system", "You are an AI assistant for Q&A. Use the context to answer. If you don't know, say that you don't know."),
("human", "Context: {context}\nQuestion: {question}")
])
print("RAG Prompt Template created!")
# Example of how it formats:
# print(rag_prompt.format(context="some info", question="a query"))
if __name__ == "__main__":
main()Connecting the Generation Engine
The final step in generating an answer is to pass the prepared context and the user's question to a Large Language Model. LangChain allows you to easily plug in various LLM providers.
For this example, we'll use a mock LLM to demonstrate the integration without needing an API key.
from langchain_core.language_models import BaseChatModel
from langchain_core.messages import BaseMessage, AIMessage
from typing import List, Any
# A simple mock LLM
class MockChatLLM(BaseChatModel):
def invoke(self, input: Any, config: Any = None) -> BaseMessage:
# Simulate LLM response based on input
if "Paris" in str(input):
return AIMessage(content="Paris is the capital of France.")
elif "London" in str(input):
return AIMessage(content="London is the capital of the UK.")
else:
return AIMessage(content="I don't have enough info to answer.")
async def ainvoke(self, input: Any, config: Any = None) -> BaseMessage:
return self.invoke(input, config) # Simple async pass-through
@property
def _llm_type(self) -> str:
return "mock-chat-llm"
def main():
llm = MockChatLLM()
print("Mock LLM initialized!")
# Example invocation (not part of the RAG chain yet)
response = llm.invoke("Tell me about Paris.")
print(f"LLM Response: {response.content}")
if __name__ == "__main__":
main()Assembling the End-to-End RAG Chain
Now, we combine the retriever, prompt template, and LLM using LangChain Expression Language (LCEL) to create a powerful, flexible RAG chain. This chain handles the entire flow.
We'll use RunnablePassthrough to manage inputs and StrOutputParser to extract the final text answer.
from langchain_core.runnables import RunnablePassthrough, RunnableLambda
from langchain_core.output_parsers import StrOutputParser
from langchain_core.prompts import ChatPromptTemplate
from langchain_core.documents import Document
from langchain_community.vectorstores import InMemoryVectorStore
from langchain_core.embeddings import Embeddings
from langchain_core.language_models import BaseChatModel
from langchain_core.messages import BaseMessage, AIMessage
from typing import List, Any
# Mock Embeddings
class MockEmbeddings(Embeddings):
def embed_documents(self, texts: List[str]) -> List[List[float]]:
return [[i * 0.1] * 10 for i in range(len(texts))]
def embed_query(self, text: str) -> List[float]:
return [0.5] * 10
# Mock LLM
class MockChatLLM(BaseChatModel):
def invoke(self, input: Any, config: Any = None) -> BaseMessage:
input_str = str(input)
if "Paris" in input_str and "capital of France" in input_str:
return AIMessage(content="Based on context, Paris is the capital of France.")
elif "Eiffel Tower" in input_str and "Paris" in input_str:
return AIMessage(content="The Eiffel Tower is in Paris, France.")
else:
return AIMessage(content="I don't have enough info in the context.")
async def ainvoke(self, input: Any, config: Any = None) -> BaseMessage:
return self.invoke(input, config)
@property
def _llm_type(self) -> str:
return "mock-chat-llm"
def main():
# 1. Setup Retriever
embeddings = MockEmbeddings()
docs = [
Document(page_content="The capital of France is Paris.", metadata={"source": "wiki"}),
Document(page_content="Eiffel Tower is in Paris, France.", metadata={"source": "travel"})
]
vectorstore = InMemoryVectorStore.from_documents(docs, embeddings)
retriever = vectorstore.as_retriever(search_kwargs={"k": 1})
# 2. Setup Prompt
rag_prompt = ChatPromptTemplate.from_messages([
("system", "You are an AI assistant. Use the following context to answer: {context}. If you don't know, say 'I don't know.'"),
("human", "Question: {question}")
])
# 3. Setup LLM
llm = MockChatLLM()
# 4. Define how to format retrieved documents
def format_docs(docs: List[Document]) -> str:
return "\n\n".join(doc.page_content for doc in docs)
# 5. Build the RAG chain
rag_chain = (
{"context": retriever | RunnableLambda(format_docs),
"question": RunnablePassthrough()}
| rag_prompt
| llm
| StrOutputParser()
)
print("RAG chain assembled!")
if __name__ == "__main__":
main()Querying Your RAG Application
With the RAG chain fully constructed, you can now invoke it with a user's question. The chain will internally handle retrieval, context formatting, prompting, and LLM generation, returning a direct answer.
This is the final step in getting a response from your RAG system.
from langchain_core.runnables import RunnablePassthrough, RunnableLambda
from langchain_core.output_parsers import StrOutputParser
from langchain_core.prompts import ChatPromptTemplate
from langchain_core.documents import Document
from langchain_community.vectorstores import InMemoryVectorStore
from langchain_core.embeddings import Embeddings
from langchain_core.language_models import BaseChatModel
from langchain_core.messages import BaseMessage, AIMessage
from typing import List, Any
# Mock Embeddings
class MockEmbeddings(Embeddings):
def embed_documents(self, texts: List[str]) -> List[List[float]]:
return [[i * 0.1] * 10 for i in range(len(texts))]
def embed_query(self, text: str) -> List[float]:
return [0.5] * 10
# Mock LLM
class MockChatLLM(BaseChatModel):
def invoke(self, input: Any, config: Any = None) -> BaseMessage:
input_str = str(input)
if "Paris" in input_str and "capital of France" in input_str:
return AIMessage(content="Based on context, Paris is the capital of France.")
elif "Eiffel Tower" in input_str and "Paris" in input_str:
return AIMessage(content="The Eiffel Tower is in Paris, France.")
else:
return AIMessage(content="I don't have enough info in the context.")
async def ainvoke(self, input: Any, config: Any = None) -> BaseMessage:
return self.invoke(input, config)
@property
def _llm_type(self) -> str:
return "mock-chat-llm"
def main():
# Setup Retriever
embeddings = MockEmbeddings()
docs = [
Document(page_content="The capital of France is Paris.", metadata={"source": "wiki"}),
Document(page_content="Eiffel Tower is in Paris, France.", metadata={"source": "travel"})
]
vectorstore = InMemoryVectorStore.from_documents(docs, embeddings)
retriever = vectorstore.as_retriever(search_kwargs={"k": 1})
# Setup Prompt
rag_prompt = ChatPromptTemplate.from_messages([
("system", "You are an AI assistant. Use the following context to answer: {context}. If you don't know, say 'I don't know.'"),
("human", "Question: {question}")
])
# Setup LLM
llm = MockChatLLM()
def format_docs(docs: List[Document]) -> str:
return "\n\n".join(doc.page_content for doc in docs)
# Build the RAG chain
rag_chain = (
{"context": retriever | RunnableLambda(format_docs),
"question": RunnablePassthrough()}
| rag_prompt
| llm
| StrOutputParser()
)
# Invoke the RAG chain with a query
query = "Where is the Eiffel Tower?"
result = rag_chain.invoke(query)
print(f"Query: {query}")
print(f"Answer: {result}")
query_no_context = "What is the capital of Japan?"
result_no_context = rag_chain.invoke(query_no_context)
print(f"\nQuery: {query_no_context}")
print(f"Answer: {result_no_context}")
if __name__ == "__main__":
main()Test Your RAG Flow Knowledge
Consider a LangChain RAG application designed to answer questions from a knowledge base.
Recap: Querying RAG
In this lesson, you learned how to bring all the RAG components together to process user queries and generate answers:
- We transformed a user's question into a query for the retriever.
- We instantiated a retriever from a vector store to fetch relevant documents.
- We crafted a prompt template to guide the LLM.
- We integrated an LLM to synthesize the final answer.
- Finally, we assembled and invoked an end-to-end RAG chain using LangChain Expression Language (LCEL).
You can now build a functional RAG system that delivers grounded, factual answers!
用 AI 导师学习 LangChain / RAG / Vector DBs — 免费
在浏览器中编写并运行真实代码,获得全天候 AI 导师的即时帮助,并在网页或应用中继续学习。
- 课程
- 12
- 课程
- 48
常见问题解答
「查询与答案生成」课时是免费的吗?
是的 — 「查询与答案生成」的完整文本可在网页上免费阅读。要进行交互式练习(内置代码编辑器和全天候 AI 导师)并解锁 LangChain / RAG / Vector DBs 课程的其余内容,请升级到 CoddyKit PRO。 LangChain / RAG / Vector DBs 课程共包含 4 节课。
「查询与答案生成」这节课中我会学到什么?
开发处理用户查询、检索相关上下文并使用 LLM 综合生成答案的逻辑 你通过在浏览器中直接运行的动手代码来练习 LangChain / RAG / Vector DBs,全天候 AI 导师会在你学习这节课的过程中回答你的问题。
学习 LangChain / RAG / Vector DBs 需要有经验吗?
无需任何先前经验。CoddyKit 上的 LangChain / RAG / Vector DBs 课程适合初学者到高级学习者,你可以从这里开始或从头开始,按照自己的节奏学习。 这是第 2 节课,共 4 节。
「查询与答案生成」课时需要多长时间?
大多数 CoddyKit 课程大约需要 5–10 分钟。每节课都很精短且互动,所以你能稳步进步,并在网页和应用中从离开的地方继续。
我能在这节 LangChain / RAG / Vector DBs 课中编写并运行代码吗?
能。每节 LangChain / RAG / Vector DBs 课都包含内置代码编辑器,你可以在浏览器中直接编写并运行真实代码,并获得即时 AI 反馈 — 无需本地设置。