RAGが解決する問題
LLMが古い情報や誤った情報をハルシネーションする実際の失敗例を調べ、取得したドキュメントに回答をgroundingすることで問題を解決する仕組みを理解します。
「RAGが解決する問題」はCoddyKit上の無料AI Engineering Academyレッスンです。 これはレッスン1/4です。 下記で完全なレッスンを無料で読むことができます。その後、ブラウザ内の組み込みコードエディタと24時間対応のAIチューターでハンズオン演習できます。 これはAI Engineering Academy学習パスの一部であり、ウェブとCoddyKitアプリ全体で進捗が同期されます。 AI Engineering Academyコースには全4レッスンが含まれています。
このレッスンの一部はまだ翻訳されておらず、英語で表示されています。
LLMs Have a Knowledge Cutoff
Every large language model is trained on a snapshot of the internet up to a specific date called the knowledge cutoff. GPT-4o has a cutoff in early 2024. Ask it about events that happened after that date and it will either confess ignorance or, worse, confidently fabricate plausible-sounding but wrong information. For applications that need current or proprietary knowledge, this is a fundamental problem.
The Hallucination Problem
Hallucination occurs when an LLM generates text that sounds authoritative but is factually wrong. Models are trained to produce fluent, coherent text — they are not explicitly trained to refuse when they do not know something. As a result, they fill gaps in knowledge with plausible guesses. Studies show that even the best models hallucinate on knowledge-intensive tasks 10-40% of the time without external grounding.
A Concrete Hallucination Example
Consider asking an LLM: What are the terms of our company's Q3 2025 vendor contract? The model has never seen your internal document. Rather than saying it does not know, it may generate a plausible-sounding contract summary using generic legal language. If an employee acts on that fabricated information, the consequences can be serious. This is the exact failure mode RAG was designed to prevent.
# Without RAG — LLM guesses from parametric memory
response = client.chat.completions.create(
model='gpt-4o',
messages=[
{'role': 'user',
'content': 'What are our Q3 2025 vendor contract terms?'}
]
)
# Model has no access to your documents — may hallucinate
print(response.choices[0].message.content)Grounding Solves Hallucination
The insight behind RAG is simple: if you give the model the relevant information inside the prompt, it does not need to rely on memorized knowledge. The model shifts from generating from memory to reading from context. This is how humans work too — when you need exact details, you look them up rather than rely on recall. RAG operationalizes that same workflow for LLMs.
# With RAG — answer grounded in retrieved documents
context = retrieve_relevant_chunks(query='Q3 2025 vendor contract terms')
response = client.chat.completions.create(
model='gpt-4o',
messages=[
{'role': 'system', 'content': 'Answer using only the provided context.'},
{'role': 'user', 'content': f'Context: {context}\n\nQuestion: What are our Q3 2025 vendor contract terms?'}
]
)Static Fine-Tuning Does Not Help Here
A common misconception is that fine-tuning the model on your documents fixes hallucination. Fine-tuning updates the model's weights to improve its style, format, and task adherence, but it does not reliably inject factual knowledge. Studies show fine-tuned models still hallucinate on the training data itself. Knowledge must be provided at inference time via the context window to be reliably recalled.
RAG Enables Real-Time Knowledge
Because RAG retrieves from a live document store, it handles knowledge that changes over time naturally. When your policy document is updated, you re-index it, and every subsequent query instantly uses the new version — no model retraining required. This makes RAG far more practical than periodic fine-tuning for applications like internal knowledge bases, customer support systems, and financial research tools.
RAG Works on Private Proprietary Data
Most enterprise data cannot be sent to OpenAI for training due to privacy and compliance requirements. RAG sidesteps this: your sensitive documents stay in your own vector database, and only the relevant chunks are sent to the LLM per query. You can even run a local LLM like Llama 3 to keep all data on-premises. RAG is the primary pattern for building AI on confidential enterprise content.
RAG Enables Source Attribution
When an LLM generates from memory, there is no source to cite. When it generates from retrieved documents, it can cite the exact sources. You can instruct the model to include document names and page numbers in its response, and users can click through to verify the original. Source attribution dramatically increases trust in AI-generated answers, which is critical in legal, medical, and financial applications.
system_prompt = '''You are a helpful assistant.
Answer questions using ONLY the provided context.
At the end of your answer, list the sources you used
in this format: [Source: document_name, page X]
If the context does not contain the answer, say:
"I don't have that information in the provided documents."'''The Retrieve-Then-Generate Pattern
RAG follows a two-step pattern at inference time: Retrieve — convert the user question to an embedding, search the vector store for the most relevant document chunks, and collect the top-K results. Generate — construct a prompt that includes the retrieved chunks as context and ask the LLM to answer the question based only on that context. The LLM reads, synthesizes, and responds.
def answer_with_rag(user_question, vector_store, llm_client):
# Step 1: Retrieve
query_embedding = embed(user_question)
chunks = vector_store.search(query_embedding, top_k=5)
context = '\n\n'.join([c['text'] for c in chunks])
# Step 2: Generate
response = llm_client.chat.completions.create(
model='gpt-4o',
messages=[
{'role': 'system', 'content': f'Answer using only:\n{context}'},
{'role': 'user', 'content': user_question}
]
)
return response.choices[0].message.contentWhat RAG Does Not Solve
RAG is powerful but not a silver bullet. It still fails when: the answer requires synthesizing across hundreds of documents (retrieval only returns a few chunks), the question is inherently multi-hop and the model needs to reason through intermediate steps, or the retrieved chunks are misleading or contradictory. Understanding these limitations helps you design hybrid systems that combine RAG with reasoning agents.
RAG vs the Alternatives
You have three main options for giving an LLM domain knowledge: prompt stuffing (put everything in the prompt — works only for very small corpora), fine-tuning (trains style and format well but not reliable for factual recall), and RAG (dynamically retrieves relevant facts at inference time, scales to millions of documents). For most production use cases requiring current, private, or large-scale knowledge, RAG is the right choice.
Quick Check
Test your understanding of AI Engineering concepts from this lesson.
Lesson Recap
In this lesson you learned: why LLMs hallucinate due to knowledge cutoffs and the inability to say 'I don't know', why fine-tuning does not solve hallucination for factual recall, and how RAG grounds answers in retrieved documents enabling source attribution, real-time knowledge, and safe use of private data. Next up we explore the full RAG architecture including its indexing and retrieval phases.
よくある質問
「RAGが解決する問題」レッスンは無料ですか?
はい。「RAGが解決する問題」の完全なテキストはこのウェブで無料で読めます。インタラクティブに演習し(組み込みコードエディタと24時間対応のAIチューター)、AI Engineering Academyコースの残りをアンロックするには、CoddyKit PROにアップグレードしてください。 AI Engineering Academyコースには全4レッスンが含まれています。
「RAGが解決する問題」で何を学びますか?
LLMが古い情報や誤った情報をハルシネーションする実際の失敗例を調べ、取得したドキュメントに回答をgroundingすることで問題を解決する仕組みを理解します。 ブラウザで直接実行するハンズオンコードでAI Engineering Academyを演習し、24時間対応のAIチューターがレッスンを進める中での質問に答えます。
AI Engineering Academyを始めるのに経験は必要ですか?
事前経験は必要ありません。CoddyKitのAI Engineering Academyは初級者から上級者向けに構成されているため、ここから始めるか最初から始めて、自分のペースで進むことができます。 これはレッスン1/4です。
「RAGが解決する問題」レッスンにはどのくらい時間がかかりますか?
ほとんどのCoddyKitレッスンは約5~10分かかります。各レッスンはコンパクトでインタラクティブなので、着実に進歩し、ウェブとアプリ全体で正確に前回の場所から再開できます。
このAI Engineering Academyレッスンでコードを書いて実行できますか?
はい。すべてのAI Engineering Academyレッスンに組み込みコードエディタが含まれているため、ブラウザでリアルコードを書いて実行し、即座のAIフィードバックを取得できます。ローカル設定は不要です。