0Pricing
LangChain / RAG / Vector DBs · 강의

쿼리 처리와 답변 생성

사용자 쿼리를 처리하고 관련 컨텍스트를 검색하며 LLM을 사용해 답변을 종합하는 로직을 개발합니다.

쿼리 처리와 답변 생성은(는) CoddyKit의 무료 LangChain / RAG / Vector DBs 강의입니다. 이것은 4개 중 2번째 강의입니다. 아래에서 전체 강의를 무료로 읽을 수 있으며, 내장 코드 에디터와 24/7 AI 튜터와 함께 브라우저에서 직접 실습할 수 있습니다. 이 강의는 LangChain / RAG / Vector DBs 학습 경로의 일부이며, 진행 상황이 웹과 CoddyKit 앱에 동기화됩니다. LangChain / RAG / Vector DBs 강의에는 총 4개의 강의가 포함되어 있습니다.

이 강의의 일부는 아직 번역되지 않았으며 영어로 표시됩니다.

Querying RAG: The Answer Flow

After integrating RAG components, the next step is to use them to answer user questions. This lesson covers the full process from a user's query to a generated answer.

We'll focus on the 'query-time' logic: how your system takes a question, finds relevant context, and synthesizes a coherent response using an LLM.

Understanding the User Query

A RAG system starts with a user's question, just like a search engine. This raw input is the trigger for the entire process.

  • It defines what information needs to be retrieved.
  • It guides the LLM on what kind of answer to generate.

No special formatting is typically needed at this initial stage; it's just plain text.

Setting Up Your Retriever

To find relevant documents, you need a retriever. This component knows how to query your vector store. You typically obtain it from your VectorStore instance.

The as_retriever() method creates this component, and you can configure parameters like k (number of top documents to fetch).

from langchain_community.vectorstores import InMemoryVectorStore
from langchain_core.documents import Document
from langchain_core.embeddings import Embeddings
from typing import List

# A simple mock for embeddings
class MockEmbeddings(Embeddings):
    def embed_documents(self, texts: List[str]) -> List[List[float]]:
        return [[i * 0.1] * 10 for i in range(len(texts))]
    def embed_query(self, text: str) -> List[float]:
        return [0.5] * 10

def main():
    # Create a dummy vector store with some content
    embeddings = MockEmbeddings()
    docs = [
        Document(page_content="The capital of France is Paris.", metadata={"source": "wiki"}),
        Document(page_content="Eiffel Tower is in Paris, France.", metadata={"source": "travel"})
    ]
    vectorstore = InMemoryVectorStore.from_documents(docs, embeddings)

    # Create a retriever from the vector store
    retriever = vectorstore.as_retriever(search_kwargs={"k": 1})
    print("Retriever created successfully!")

if __name__ == "__main__":
    main()

Fetching Contextual Documents

Once you have a retriever, you can invoke it with the user's query. It will perform a similarity search in your vector store and return the most relevant Document objects.

These documents form the context that will be passed to the LLM.

from langchain_community.vectorstores import InMemoryVectorStore
from langchain_core.documents import Document
from langchain_core.embeddings import Embeddings
from typing import List

class MockEmbeddings(Embeddings):
    def embed_documents(self, texts: List[str]) -> List[List[float]]:
        return [[i * 0.1] * 10 for i in range(len(texts))]
    def embed_query(self, text: str) -> List[float]:
        return [0.5] * 10

def main():
    embeddings = MockEmbeddings()
    docs = [
        Document(page_content="The capital of France is Paris.", metadata={"source": "wiki"}),
        Document(page_content="Eiffel Tower is in Paris, France.", metadata={"source": "travel"}),
        Document(page_content="London is the capital of the UK.", metadata={"source": "wiki"})
    ]
    vectorstore = InMemoryVectorStore.from_documents(docs, embeddings)
    retriever = vectorstore.as_retriever(search_kwargs={"k": 2})

    user_query = "What is the capital of France?"
    retrieved_docs = retriever.invoke(user_query)

    print(f"Retrieved {len(retrieved_docs)} documents:")
    for doc in retrieved_docs:
        print(f"- {doc.page_content[:50]}...")

if __name__ == "__main__":
    main()

Preparing Context for the LLM

LLMs usually prefer a single string of text as context. The retrieved Document objects need to be combined into a coherent format.

A common approach is to concatenate their page_content fields, perhaps with separators, and include source metadata if desired.

  • Ensures all context fits within the LLM's token window.
  • Presents a clean input for the LLM to reason over.

Crafting the RAG Prompt

The prompt template is crucial. It instructs the LLM on how to use the provided context to answer the user's question. It typically includes placeholders for both the context and the question.

A well-designed prompt guides the LLM to be factual and avoid hallucination.

from langchain_core.prompts import ChatPromptTemplate

def main():
    # Define a RAG-specific prompt template
    rag_prompt = ChatPromptTemplate.from_messages([
        ("system", "You are an AI assistant for Q&A. Use the context to answer. If you don't know, say that you don't know."),
        ("human", "Context: {context}\nQuestion: {question}")
    ])

    print("RAG Prompt Template created!")
    # Example of how it formats:
    # print(rag_prompt.format(context="some info", question="a query"))

if __name__ == "__main__":
    main()

Connecting the Generation Engine

The final step in generating an answer is to pass the prepared context and the user's question to a Large Language Model. LangChain allows you to easily plug in various LLM providers.

For this example, we'll use a mock LLM to demonstrate the integration without needing an API key.

from langchain_core.language_models import BaseChatModel
from langchain_core.messages import BaseMessage, AIMessage
from typing import List, Any

# A simple mock LLM
class MockChatLLM(BaseChatModel):
    def invoke(self, input: Any, config: Any = None) -> BaseMessage:
        # Simulate LLM response based on input
        if "Paris" in str(input):
            return AIMessage(content="Paris is the capital of France.")
        elif "London" in str(input):
            return AIMessage(content="London is the capital of the UK.")
        else:
            return AIMessage(content="I don't have enough info to answer.")

    async def ainvoke(self, input: Any, config: Any = None) -> BaseMessage:
        return self.invoke(input, config) # Simple async pass-through

    @property
    def _llm_type(self) -> str:
        return "mock-chat-llm"

def main():
    llm = MockChatLLM()
    print("Mock LLM initialized!")

    # Example invocation (not part of the RAG chain yet)
    response = llm.invoke("Tell me about Paris.")
    print(f"LLM Response: {response.content}")

if __name__ == "__main__":
    main()

Assembling the End-to-End RAG Chain

Now, we combine the retriever, prompt template, and LLM using LangChain Expression Language (LCEL) to create a powerful, flexible RAG chain. This chain handles the entire flow.

We'll use RunnablePassthrough to manage inputs and StrOutputParser to extract the final text answer.

from langchain_core.runnables import RunnablePassthrough, RunnableLambda
from langchain_core.output_parsers import StrOutputParser
from langchain_core.prompts import ChatPromptTemplate
from langchain_core.documents import Document
from langchain_community.vectorstores import InMemoryVectorStore
from langchain_core.embeddings import Embeddings
from langchain_core.language_models import BaseChatModel
from langchain_core.messages import BaseMessage, AIMessage
from typing import List, Any

# Mock Embeddings
class MockEmbeddings(Embeddings):
    def embed_documents(self, texts: List[str]) -> List[List[float]]:
        return [[i * 0.1] * 10 for i in range(len(texts))]
    def embed_query(self, text: str) -> List[float]:
        return [0.5] * 10

# Mock LLM
class MockChatLLM(BaseChatModel):
    def invoke(self, input: Any, config: Any = None) -> BaseMessage:
        input_str = str(input)
        if "Paris" in input_str and "capital of France" in input_str:
            return AIMessage(content="Based on context, Paris is the capital of France.")
        elif "Eiffel Tower" in input_str and "Paris" in input_str:
            return AIMessage(content="The Eiffel Tower is in Paris, France.")
        else:
            return AIMessage(content="I don't have enough info in the context.")
    async def ainvoke(self, input: Any, config: Any = None) -> BaseMessage:
        return self.invoke(input, config)
    @property
    def _llm_type(self) -> str:
        return "mock-chat-llm"

def main():
    # 1. Setup Retriever
    embeddings = MockEmbeddings()
    docs = [
        Document(page_content="The capital of France is Paris.", metadata={"source": "wiki"}),
        Document(page_content="Eiffel Tower is in Paris, France.", metadata={"source": "travel"})
    ]
    vectorstore = InMemoryVectorStore.from_documents(docs, embeddings)
    retriever = vectorstore.as_retriever(search_kwargs={"k": 1})

    # 2. Setup Prompt
    rag_prompt = ChatPromptTemplate.from_messages([
        ("system", "You are an AI assistant. Use the following context to answer: {context}. If you don't know, say 'I don't know.'"),
        ("human", "Question: {question}")
    ])

    # 3. Setup LLM
    llm = MockChatLLM()

    # 4. Define how to format retrieved documents
    def format_docs(docs: List[Document]) -> str:
        return "\n\n".join(doc.page_content for doc in docs)

    # 5. Build the RAG chain
    rag_chain = (
        {"context": retriever | RunnableLambda(format_docs),
         "question": RunnablePassthrough()}
        | rag_prompt
        | llm
        | StrOutputParser()
    )

    print("RAG chain assembled!")

if __name__ == "__main__":
    main()

Querying Your RAG Application

With the RAG chain fully constructed, you can now invoke it with a user's question. The chain will internally handle retrieval, context formatting, prompting, and LLM generation, returning a direct answer.

This is the final step in getting a response from your RAG system.

from langchain_core.runnables import RunnablePassthrough, RunnableLambda
from langchain_core.output_parsers import StrOutputParser
from langchain_core.prompts import ChatPromptTemplate
from langchain_core.documents import Document
from langchain_community.vectorstores import InMemoryVectorStore
from langchain_core.embeddings import Embeddings
from langchain_core.language_models import BaseChatModel
from langchain_core.messages import BaseMessage, AIMessage
from typing import List, Any

# Mock Embeddings
class MockEmbeddings(Embeddings):
    def embed_documents(self, texts: List[str]) -> List[List[float]]:
        return [[i * 0.1] * 10 for i in range(len(texts))]
    def embed_query(self, text: str) -> List[float]:
        return [0.5] * 10

# Mock LLM
class MockChatLLM(BaseChatModel):
    def invoke(self, input: Any, config: Any = None) -> BaseMessage:
        input_str = str(input)
        if "Paris" in input_str and "capital of France" in input_str:
            return AIMessage(content="Based on context, Paris is the capital of France.")
        elif "Eiffel Tower" in input_str and "Paris" in input_str:
            return AIMessage(content="The Eiffel Tower is in Paris, France.")
        else:
            return AIMessage(content="I don't have enough info in the context.")
    async def ainvoke(self, input: Any, config: Any = None) -> BaseMessage:
        return self.invoke(input, config)
    @property
    def _llm_type(self) -> str:
        return "mock-chat-llm"

def main():
    # Setup Retriever
    embeddings = MockEmbeddings()
    docs = [
        Document(page_content="The capital of France is Paris.", metadata={"source": "wiki"}),
        Document(page_content="Eiffel Tower is in Paris, France.", metadata={"source": "travel"})
    ]
    vectorstore = InMemoryVectorStore.from_documents(docs, embeddings)
    retriever = vectorstore.as_retriever(search_kwargs={"k": 1})

    # Setup Prompt
    rag_prompt = ChatPromptTemplate.from_messages([
        ("system", "You are an AI assistant. Use the following context to answer: {context}. If you don't know, say 'I don't know.'"),
        ("human", "Question: {question}")
    ])

    # Setup LLM
    llm = MockChatLLM()

    def format_docs(docs: List[Document]) -> str:
        return "\n\n".join(doc.page_content for doc in docs)

    # Build the RAG chain
    rag_chain = (
        {"context": retriever | RunnableLambda(format_docs),
         "question": RunnablePassthrough()}
        | rag_prompt
        | llm
        | StrOutputParser()
    )

    # Invoke the RAG chain with a query
    query = "Where is the Eiffel Tower?"
    result = rag_chain.invoke(query)
    print(f"Query: {query}")
    print(f"Answer: {result}")

    query_no_context = "What is the capital of Japan?"
    result_no_context = rag_chain.invoke(query_no_context)
    print(f"\nQuery: {query_no_context}")
    print(f"Answer: {result_no_context}")

if __name__ == "__main__":
    main()

Test Your RAG Flow Knowledge

Consider a LangChain RAG application designed to answer questions from a knowledge base.

Recap: Querying RAG

In this lesson, you learned how to bring all the RAG components together to process user queries and generate answers:

  • We transformed a user's question into a query for the retriever.
  • We instantiated a retriever from a vector store to fetch relevant documents.
  • We crafted a prompt template to guide the LLM.
  • We integrated an LLM to synthesize the final answer.
  • Finally, we assembled and invoked an end-to-end RAG chain using LangChain Expression Language (LCEL).

You can now build a functional RAG system that delivers grounded, factual answers!

자주 묻는 질문

“쿼리 처리와 답변 생성” 강의는 무료인가요?

네 — “쿼리 처리와 답변 생성” 전체 내용을 이 웹사이트에서 무료로 읽을 수 있습니다. 인터랙티브하게 실습하려면(내장 코드 에디터와 24/7 AI 튜터), CoddyKit PRO로 업그레이드하면 LangChain / RAG / Vector DBs 강의 전체를 잠금 해제할 수 있습니다. LangChain / RAG / Vector DBs 강의에는 총 4개의 강의가 포함되어 있습니다.

“쿼리 처리와 답변 생성”에서 뭘 배우나요?

사용자 쿼리를 처리하고 관련 컨텍스트를 검색하며 LLM을 사용해 답변을 종합하는 로직을 개발합니다. 브라우저에서 직접 실행하는 실습 코드로 LangChain / RAG / Vector DBs을(를) 배우며, 24/7 AI 튜터가 강의를 진행하면서 질문에 답변해줍니다.

LangChain / RAG / Vector DBs을(를) 시작하는 데 경험이 필요한가요?

사전 경험은 필요하지 않습니다. CoddyKit의 LangChain / RAG / Vector DBs은(는) 초급자부터 고급 학습자까지를 위해 구성되어 있으므로, 여기서 시작하거나 처음부터 시작할 수 있으며 자신의 속도대로 진행할 수 있습니다. 이것은 4개 중 2번째 강의입니다.

“쿼리 처리와 답변 생성” 강의는 얼마나 걸리나요?

대부분의 CoddyKit 강의는 약 5~10분이 소요됩니다. 각 강의는 간결하고 인터랙티브하여 꾸준한 진행이 가능하며, 웹과 앱에서 중단한 부분부터 바로 시작할 수 있습니다.

이 LangChain / RAG / Vector DBs 강의에서 코드를 작성하고 실행할 수 있나요?

네. 모든 LangChain / RAG / Vector DBs 강의에는 내장 코드 에디터가 포함되어 있으므로, 브라우저에서 바로 실제 코드를 작성하고 실행한 후 즉시 AI 피드백을 받을 수 있습니다 — 로컬 설정이 필요 없습니다.

이 강의의 모든 강의

  1. 모든 RAG 구성 요소 통합
  2. 쿼리 처리와 답변 생성
  3. RAG 시스템 성능 평가
  4. RAG용 골든 테스트 세트 만들기
← LangChain / RAG / Vector DBs(으)로 돌아가기