Testing & Evaluating Your RAG App
Build confidence in your first RAG application by creating a test set and measuring retrieval and answer quality with practical metrics before you ship.
Testing & Evaluating Your RAG App is a free LLM Apps in Production (RAG + Vector DB + Caching) lesson on CoddyKit — lesson 4 of 4. You can read the complete lesson below for free — then practise it hands-on in the browser with a built-in code editor and a 24/7 AI tutor. It is part of the LLM Apps in Production (RAG + Vector DB + Caching) learning path, one of 4 lessons in the course, and your progress syncs across the web and the CoddyKit app.
Why Evaluate RAG
A RAG app can look fine on a few queries and fail badly on others. Without measurement you cannot tell if a change helped or hurt.
Evaluation gives you a repeatable score to guide improvements.
Two Things to Measure
RAG quality has two parts:
- Retrieval: did we fetch the right documents?
- Generation: did the answer use them correctly?
A bad answer can come from either, so measure both.
Building a Test Set
Create a small set of questions with known correct answers and the documents that contain them. Even 20 to 50 examples are enough to start.
testset = [
{'q': 'What is the refund window?',
'answer': '30 days',
'source': 'policy.md'}
]Retrieval Metric: Hit Rate
Hit rate (or recall@k) checks whether the correct source appears in the top-k retrieved chunks. High hit rate means retrieval is doing its job.
def hit(retrieved, expected_source):
return any(d.metadata['source'] == expected_source
for d in retrieved)Faithfulness
Faithfulness asks: is the answer supported by the retrieved context, or did the model make things up? An LLM judge can score this automatically.
Answer Relevance
Answer relevance measures whether the response actually addresses the question, regardless of sources. A faithful answer can still be off-topic.
LLM as a Judge
You can use a strong model to grade outputs against the expected answer, returning a pass or score with a reason.
judge_prompt = (
'Question: {q}\nExpected: {gold}\n'
'Got: {pred}\nIs it correct? Answer yes or no.'
)Running the Evaluation
Loop over the test set, run your pipeline, and aggregate scores into a single report you can compare across versions.
scores = []
for case in testset:
pred = rag.invoke(case['q'])
scores.append(grade(case, pred))
print(sum(scores) / len(scores))Comparing Configurations
Change one variable — chunk size, k, prompt, model — rerun the same test set, and compare scores. This turns guesswork into evidence-based tuning.
Watching for Regressions
Keep the test suite in CI. When a change drops a metric, you catch the regression before users do. Treat evaluation like unit tests for AI quality.
Improving From Results
Use failures to guide fixes:
- Low hit rate? Adjust chunking or retrieval
- Low faithfulness? Strengthen grounding instructions
- Low relevance? Improve the prompt
Quick Check
Test your evaluation knowledge.
Recap
You learned to evaluate your RAG app:
- Measure both retrieval and generation
- Build a small test set with known answers
- Use hit rate, faithfulness, and answer relevance
- Let an LLM judge grade outputs
- Compare configs and guard against regressions in CI
Evaluation turns RAG improvement into a measurable, repeatable process.
Frequently asked questions
Is the “Testing & Evaluating Your RAG App” lesson free?
Yes — the full text of “Testing & Evaluating Your RAG App” is free to read here on the web, and the LLM Apps in Production (RAG + Vector DB + Caching) course includes 4 lessons in total. To practise it interactively (a built-in code editor and a 24/7 AI tutor) and unlock the rest of the LLM Apps in Production (RAG + Vector DB + Caching) course, upgrade to CoddyKit PRO.
What will I learn in “Testing & Evaluating Your RAG App”?
Build confidence in your first RAG application by creating a test set and measuring retrieval and answer quality with practical metrics before you ship. You practise LLM Apps in Production (RAG + Vector DB + Caching) with hands-on code you run directly in the browser, and a 24/7 AI tutor answers your questions as you work through the lesson.
Do I need any experience to start LLM Apps in Production (RAG + Vector DB + Caching)?
No prior experience is required. LLM Apps in Production (RAG + Vector DB + Caching) on CoddyKit is structured for beginners through advanced learners; this is — lesson 4 of 4, so you can start here or from the beginning and move at your own pace.
How long does the “Testing & Evaluating Your RAG App” lesson take?
Most CoddyKit lessons take about 5–10 minutes. Each one is bite-sized and interactive, so you make steady progress and pick up exactly where you left off across the web and the app.
Can I write and run code in this LLM Apps in Production (RAG + Vector DB + Caching) lesson?
Yes. Every LLM Apps in Production (RAG + Vector DB + Caching) lesson includes a built-in code editor, so you write and run real code right in your browser and get instant AI feedback — no local setup required.
All lessons in this course
- Choosing an LLM Provider
- Data Loading and Text Chunking Basics
- Building a Simple RAG Pipeline
- Testing & Evaluating Your RAG App