El problema que resuelve RAG
Examinará casos reales de fallo en los que los LLM alucinan información desactualizada o incorrecta, y comprenderá cómo fundamentar las respuestas en documentos recuperados soluciona estos problemas.
El problema que resuelve RAG es una lección gratuita de AI Engineering Academy en CoddyKit. Esta es la lección 1 de 4. Puedes leer la lección completa abajo gratuitamente — luego la practicas en el navegador con un editor de código integrado y un tutor de IA 24/7. Forma parte de la ruta de aprendizaje de AI Engineering Academy, y tu progreso se sincroniza en la web y la app de CoddyKit. El curso de AI Engineering Academy incluye 4 lecciones en total.
Partes de esta lección aún no han sido traducidas y se muestran en inglés.
LLMs Have a Knowledge Cutoff
Every large language model is trained on a snapshot of the internet up to a specific date called the knowledge cutoff. GPT-4o has a cutoff in early 2024. Ask it about events that happened after that date and it will either confess ignorance or, worse, confidently fabricate plausible-sounding but wrong information. For applications that need current or proprietary knowledge, this is a fundamental problem.
The Hallucination Problem
Hallucination occurs when an LLM generates text that sounds authoritative but is factually wrong. Models are trained to produce fluent, coherent text — they are not explicitly trained to refuse when they do not know something. As a result, they fill gaps in knowledge with plausible guesses. Studies show that even the best models hallucinate on knowledge-intensive tasks 10-40% of the time without external grounding.
A Concrete Hallucination Example
Consider asking an LLM: What are the terms of our company's Q3 2025 vendor contract? The model has never seen your internal document. Rather than saying it does not know, it may generate a plausible-sounding contract summary using generic legal language. If an employee acts on that fabricated information, the consequences can be serious. This is the exact failure mode RAG was designed to prevent.
# Without RAG — LLM guesses from parametric memory
response = client.chat.completions.create(
model='gpt-4o',
messages=[
{'role': 'user',
'content': 'What are our Q3 2025 vendor contract terms?'}
]
)
# Model has no access to your documents — may hallucinate
print(response.choices[0].message.content)Grounding Solves Hallucination
The insight behind RAG is simple: if you give the model the relevant information inside the prompt, it does not need to rely on memorized knowledge. The model shifts from generating from memory to reading from context. This is how humans work too — when you need exact details, you look them up rather than rely on recall. RAG operationalizes that same workflow for LLMs.
# With RAG — answer grounded in retrieved documents
context = retrieve_relevant_chunks(query='Q3 2025 vendor contract terms')
response = client.chat.completions.create(
model='gpt-4o',
messages=[
{'role': 'system', 'content': 'Answer using only the provided context.'},
{'role': 'user', 'content': f'Context: {context}\n\nQuestion: What are our Q3 2025 vendor contract terms?'}
]
)Static Fine-Tuning Does Not Help Here
A common misconception is that fine-tuning the model on your documents fixes hallucination. Fine-tuning updates the model's weights to improve its style, format, and task adherence, but it does not reliably inject factual knowledge. Studies show fine-tuned models still hallucinate on the training data itself. Knowledge must be provided at inference time via the context window to be reliably recalled.
RAG Enables Real-Time Knowledge
Because RAG retrieves from a live document store, it handles knowledge that changes over time naturally. When your policy document is updated, you re-index it, and every subsequent query instantly uses the new version — no model retraining required. This makes RAG far more practical than periodic fine-tuning for applications like internal knowledge bases, customer support systems, and financial research tools.
RAG Works on Private Proprietary Data
Most enterprise data cannot be sent to OpenAI for training due to privacy and compliance requirements. RAG sidesteps this: your sensitive documents stay in your own vector database, and only the relevant chunks are sent to the LLM per query. You can even run a local LLM like Llama 3 to keep all data on-premises. RAG is the primary pattern for building AI on confidential enterprise content.
RAG Enables Source Attribution
When an LLM generates from memory, there is no source to cite. When it generates from retrieved documents, it can cite the exact sources. You can instruct the model to include document names and page numbers in its response, and users can click through to verify the original. Source attribution dramatically increases trust in AI-generated answers, which is critical in legal, medical, and financial applications.
system_prompt = '''You are a helpful assistant.
Answer questions using ONLY the provided context.
At the end of your answer, list the sources you used
in this format: [Source: document_name, page X]
If the context does not contain the answer, say:
"I don't have that information in the provided documents."'''The Retrieve-Then-Generate Pattern
RAG follows a two-step pattern at inference time: Retrieve — convert the user question to an embedding, search the vector store for the most relevant document chunks, and collect the top-K results. Generate — construct a prompt that includes the retrieved chunks as context and ask the LLM to answer the question based only on that context. The LLM reads, synthesizes, and responds.
def answer_with_rag(user_question, vector_store, llm_client):
# Step 1: Retrieve
query_embedding = embed(user_question)
chunks = vector_store.search(query_embedding, top_k=5)
context = '\n\n'.join([c['text'] for c in chunks])
# Step 2: Generate
response = llm_client.chat.completions.create(
model='gpt-4o',
messages=[
{'role': 'system', 'content': f'Answer using only:\n{context}'},
{'role': 'user', 'content': user_question}
]
)
return response.choices[0].message.contentWhat RAG Does Not Solve
RAG is powerful but not a silver bullet. It still fails when: the answer requires synthesizing across hundreds of documents (retrieval only returns a few chunks), the question is inherently multi-hop and the model needs to reason through intermediate steps, or the retrieved chunks are misleading or contradictory. Understanding these limitations helps you design hybrid systems that combine RAG with reasoning agents.
RAG vs the Alternatives
You have three main options for giving an LLM domain knowledge: prompt stuffing (put everything in the prompt — works only for very small corpora), fine-tuning (trains style and format well but not reliable for factual recall), and RAG (dynamically retrieves relevant facts at inference time, scales to millions of documents). For most production use cases requiring current, private, or large-scale knowledge, RAG is the right choice.
Quick Check
Test your understanding of AI Engineering concepts from this lesson.
Lesson Recap
In this lesson you learned: why LLMs hallucinate due to knowledge cutoffs and the inability to say 'I don't know', why fine-tuning does not solve hallucination for factual recall, and how RAG grounds answers in retrieved documents enabling source attribution, real-time knowledge, and safe use of private data. Next up we explore the full RAG architecture including its indexing and retrieval phases.
Preguntas frecuentes
¿La lección «El problema que resuelve RAG» es gratis?
Sí — el texto completo de «El problema que resuelve RAG» es gratis para leer aquí en la web. Para practicarla de forma interactiva (editor de código integrado y tutor de IA 24/7) y desbloquear el resto del curso de AI Engineering Academy, actualiza a CoddyKit PRO. El curso de AI Engineering Academy incluye 4 lecciones en total.
¿Qué aprenderé en «El problema que resuelve RAG»?
Examinará casos reales de fallo en los que los LLM alucinan información desactualizada o incorrecta, y comprenderá cómo fundamentar las respuestas en documentos recuperados soluciona estos problemas. Practicas AI Engineering Academy con código real que ejecutas directamente en el navegador, y un tutor de IA 24/7 responde tus preguntas mientras trabajas en la lección.
¿Necesito experiencia previa para empezar AI Engineering Academy?
No se requiere experiencia previa. AI Engineering Academy en CoddyKit está estructurado para principiantes hasta estudiantes avanzados, así que puedes empezar aquí o desde el inicio y avanzar a tu ritmo. Esta es la lección 1 de 4.
¿Cuánto tiempo toma la lección «El problema que resuelve RAG»?
La mayoría de las lecciones de CoddyKit toman alrededor de 5–10 minutos. Cada una es compacta e interactiva, así que avanzas constantemente y retomas exactamente por donde dejaste en la web y la app.
¿Puedo escribir y ejecutar código en esta lección de AI Engineering Academy?
Sí. Cada lección de AI Engineering Academy incluye un editor de código integrado, así que escribes y ejecutas código real directamente en tu navegador y obtienes retroalimentación instantánea de IA — sin configuración local necesaria.
Todas las lecciones de este curso
- El problema que resuelve RAG
- La arquitectura RAG: indexación y recuperación
- Creación del prompt aumentado
- RAG frente a fine-tuning: cuándo usar cada uno