Le problème résolu par RAG
Examinez des cas réels d’échec où les LLM produisent des informations obsolètes ou erronées, puis comprenez comment l’ancrage des réponses dans des documents récupérés permet de résoudre ces problèmes.
Le problème résolu par RAG est une leçon AI Engineering Academy gratuite sur CoddyKit. Ceci est la leçon 1 sur 4. Tu peux lire la leçon complète ci-dessous gratuitement — puis la pratiquer en direct dans le navigateur avec un éditeur de code intégré et un tuteur IA 24/7. Elle fait partie du parcours d'apprentissage AI Engineering Academy, et ta progression se synchronise sur le web et l'application CoddyKit. Le cours AI Engineering Academy comprend 4 leçons au total.
Certaines parties de cette leçon n'ont pas encore été traduites et s'affichent en anglais.
LLMs Have a Knowledge Cutoff
Every large language model is trained on a snapshot of the internet up to a specific date called the knowledge cutoff. GPT-4o has a cutoff in early 2024. Ask it about events that happened after that date and it will either confess ignorance or, worse, confidently fabricate plausible-sounding but wrong information. For applications that need current or proprietary knowledge, this is a fundamental problem.
The Hallucination Problem
Hallucination occurs when an LLM generates text that sounds authoritative but is factually wrong. Models are trained to produce fluent, coherent text — they are not explicitly trained to refuse when they do not know something. As a result, they fill gaps in knowledge with plausible guesses. Studies show that even the best models hallucinate on knowledge-intensive tasks 10-40% of the time without external grounding.
A Concrete Hallucination Example
Consider asking an LLM: What are the terms of our company's Q3 2025 vendor contract? The model has never seen your internal document. Rather than saying it does not know, it may generate a plausible-sounding contract summary using generic legal language. If an employee acts on that fabricated information, the consequences can be serious. This is the exact failure mode RAG was designed to prevent.
# Without RAG — LLM guesses from parametric memory
response = client.chat.completions.create(
model='gpt-4o',
messages=[
{'role': 'user',
'content': 'What are our Q3 2025 vendor contract terms?'}
]
)
# Model has no access to your documents — may hallucinate
print(response.choices[0].message.content)Grounding Solves Hallucination
The insight behind RAG is simple: if you give the model the relevant information inside the prompt, it does not need to rely on memorized knowledge. The model shifts from generating from memory to reading from context. This is how humans work too — when you need exact details, you look them up rather than rely on recall. RAG operationalizes that same workflow for LLMs.
# With RAG — answer grounded in retrieved documents
context = retrieve_relevant_chunks(query='Q3 2025 vendor contract terms')
response = client.chat.completions.create(
model='gpt-4o',
messages=[
{'role': 'system', 'content': 'Answer using only the provided context.'},
{'role': 'user', 'content': f'Context: {context}\n\nQuestion: What are our Q3 2025 vendor contract terms?'}
]
)Static Fine-Tuning Does Not Help Here
A common misconception is that fine-tuning the model on your documents fixes hallucination. Fine-tuning updates the model's weights to improve its style, format, and task adherence, but it does not reliably inject factual knowledge. Studies show fine-tuned models still hallucinate on the training data itself. Knowledge must be provided at inference time via the context window to be reliably recalled.
RAG Enables Real-Time Knowledge
Because RAG retrieves from a live document store, it handles knowledge that changes over time naturally. When your policy document is updated, you re-index it, and every subsequent query instantly uses the new version — no model retraining required. This makes RAG far more practical than periodic fine-tuning for applications like internal knowledge bases, customer support systems, and financial research tools.
RAG Works on Private Proprietary Data
Most enterprise data cannot be sent to OpenAI for training due to privacy and compliance requirements. RAG sidesteps this: your sensitive documents stay in your own vector database, and only the relevant chunks are sent to the LLM per query. You can even run a local LLM like Llama 3 to keep all data on-premises. RAG is the primary pattern for building AI on confidential enterprise content.
RAG Enables Source Attribution
When an LLM generates from memory, there is no source to cite. When it generates from retrieved documents, it can cite the exact sources. You can instruct the model to include document names and page numbers in its response, and users can click through to verify the original. Source attribution dramatically increases trust in AI-generated answers, which is critical in legal, medical, and financial applications.
system_prompt = '''You are a helpful assistant.
Answer questions using ONLY the provided context.
At the end of your answer, list the sources you used
in this format: [Source: document_name, page X]
If the context does not contain the answer, say:
"I don't have that information in the provided documents."'''The Retrieve-Then-Generate Pattern
RAG follows a two-step pattern at inference time: Retrieve — convert the user question to an embedding, search the vector store for the most relevant document chunks, and collect the top-K results. Generate — construct a prompt that includes the retrieved chunks as context and ask the LLM to answer the question based only on that context. The LLM reads, synthesizes, and responds.
def answer_with_rag(user_question, vector_store, llm_client):
# Step 1: Retrieve
query_embedding = embed(user_question)
chunks = vector_store.search(query_embedding, top_k=5)
context = '\n\n'.join([c['text'] for c in chunks])
# Step 2: Generate
response = llm_client.chat.completions.create(
model='gpt-4o',
messages=[
{'role': 'system', 'content': f'Answer using only:\n{context}'},
{'role': 'user', 'content': user_question}
]
)
return response.choices[0].message.contentWhat RAG Does Not Solve
RAG is powerful but not a silver bullet. It still fails when: the answer requires synthesizing across hundreds of documents (retrieval only returns a few chunks), the question is inherently multi-hop and the model needs to reason through intermediate steps, or the retrieved chunks are misleading or contradictory. Understanding these limitations helps you design hybrid systems that combine RAG with reasoning agents.
RAG vs the Alternatives
You have three main options for giving an LLM domain knowledge: prompt stuffing (put everything in the prompt — works only for very small corpora), fine-tuning (trains style and format well but not reliable for factual recall), and RAG (dynamically retrieves relevant facts at inference time, scales to millions of documents). For most production use cases requiring current, private, or large-scale knowledge, RAG is the right choice.
Quick Check
Test your understanding of AI Engineering concepts from this lesson.
Lesson Recap
In this lesson you learned: why LLMs hallucinate due to knowledge cutoffs and the inability to say 'I don't know', why fine-tuning does not solve hallucination for factual recall, and how RAG grounds answers in retrieved documents enabling source attribution, real-time knowledge, and safe use of private data. Next up we explore the full RAG architecture including its indexing and retrieval phases.
Questions Fréquemment Posées
La leçon « Le problème résolu par RAG » est-elle gratuite ?
Oui — le texte complet de « Le problème résolu par RAG » est gratuit à lire ici sur le web. Pour la pratiquer de manière interactive (un éditeur de code intégré et un tuteur IA 24/7) et déverrouiller le reste du cours AI Engineering Academy, passe à CoddyKit PRO. Le cours AI Engineering Academy comprend 4 leçons au total.
Qu'est-ce que j'apprendrai dans « Le problème résolu par RAG » ?
Examinez des cas réels d’échec où les LLM produisent des informations obsolètes ou erronées, puis comprenez comment l’ancrage des réponses dans des documents récupérés permet de résoudre ces problème… Tu pratiques AI Engineering Academy avec du code pratique que tu exécutes directement dans le navigateur, et un tuteur IA 24/7 répond à tes questions au fur et à mesure que tu avances dans la leçon.
Dois-je avoir de l'expérience pour commencer AI Engineering Academy ?
Aucune expérience préalable n'est requise. AI Engineering Academy sur CoddyKit est structuré pour les débutants jusqu'aux apprenants avancés, donc tu peux commencer ici ou depuis le début et avancer à ton rythme. Ceci est la leçon 1 sur 4.
Combien de temps prend la leçon « Le problème résolu par RAG » ?
La plupart des leçons CoddyKit prennent environ 5–10 minutes. Chacune est courte et interactive, tu progresses régulièrement et tu repiques exactement où tu t'es arrêté sur le web et l'app.
Peux-tu écrire et exécuter du code dans cette leçon AI Engineering Academy ?
Oui. Chaque leçon AI Engineering Academy inclut un éditeur de code intégré, tu écris et exécutes du vrai code directement dans ton navigateur et tu reçois des retours IA instantanés — aucune configuration locale requise.
Toutes les leçons de ce cours
- Le problème résolu par RAG
- Architecture RAG : indexation et récupération
- Rédiger le prompt enrichi
- RAG ou affinage : quand utiliser l’un ou l’autre