0Pricing
LLM Apps in Production (RAG + Vector DB + Caching) · Lektion

Ihre RAG-App testen und bewerten

Schaffen Sie Vertrauen in Ihre erste RAG-Anwendung, indem Sie einen Testsatz erstellen und die Qualität von Retrieval und Antworten vor dem Release mit praxisnahen Metriken messen.

Ihre RAG-App testen und bewerten ist eine kostenlose LLM Apps in Production (RAG + Vector DB + Caching)-Lektion auf CoddyKit. Dies ist Lektion 4 von 4. Du kannst die komplette Lektion unten kostenlos lesen – dann übst du sie direkt im Browser mit einem integrierten Code-Editor und einem KI-Tutor rund um die Uhr. Sie ist Teil des LLM Apps in Production (RAG + Vector DB + Caching)-Lernpfads, und dein Fortschritt wird über Web und CoddyKit-App synchronisiert. Der LLM Apps in Production (RAG + Vector DB + Caching)-Kurs umfasst insgesamt 4 Lektionen.

Teile dieser Lektion wurden noch nicht übersetzt und werden auf Englisch angezeigt.

Why Evaluate RAG

A RAG app can look fine on a few queries and fail badly on others. Without measurement you cannot tell if a change helped or hurt.

Evaluation gives you a repeatable score to guide improvements.

Two Things to Measure

RAG quality has two parts:

  • Retrieval: did we fetch the right documents?
  • Generation: did the answer use them correctly?

A bad answer can come from either, so measure both.

Building a Test Set

Create a small set of questions with known correct answers and the documents that contain them. Even 20 to 50 examples are enough to start.

testset = [
    {'q': 'What is the refund window?',
     'answer': '30 days',
     'source': 'policy.md'}
]

Retrieval Metric: Hit Rate

Hit rate (or recall@k) checks whether the correct source appears in the top-k retrieved chunks. High hit rate means retrieval is doing its job.

def hit(retrieved, expected_source):
    return any(d.metadata['source'] == expected_source
               for d in retrieved)

Faithfulness

Faithfulness asks: is the answer supported by the retrieved context, or did the model make things up? An LLM judge can score this automatically.

Answer Relevance

Answer relevance measures whether the response actually addresses the question, regardless of sources. A faithful answer can still be off-topic.

LLM as a Judge

You can use a strong model to grade outputs against the expected answer, returning a pass or score with a reason.

judge_prompt = (
    'Question: {q}\nExpected: {gold}\n'
    'Got: {pred}\nIs it correct? Answer yes or no.'
)

Running the Evaluation

Loop over the test set, run your pipeline, and aggregate scores into a single report you can compare across versions.

scores = []
for case in testset:
    pred = rag.invoke(case['q'])
    scores.append(grade(case, pred))
print(sum(scores) / len(scores))

Comparing Configurations

Change one variable — chunk size, k, prompt, model — rerun the same test set, and compare scores. This turns guesswork into evidence-based tuning.

Watching for Regressions

Keep the test suite in CI. When a change drops a metric, you catch the regression before users do. Treat evaluation like unit tests for AI quality.

Improving From Results

Use failures to guide fixes:

  • Low hit rate? Adjust chunking or retrieval
  • Low faithfulness? Strengthen grounding instructions
  • Low relevance? Improve the prompt

Quick Check

Test your evaluation knowledge.

Recap

You learned to evaluate your RAG app:

  • Measure both retrieval and generation
  • Build a small test set with known answers
  • Use hit rate, faithfulness, and answer relevance
  • Let an LLM judge grade outputs
  • Compare configs and guard against regressions in CI

Evaluation turns RAG improvement into a measurable, repeatable process.

Häufig gestellte Fragen

Ist die Lektion „Ihre RAG-App testen und bewerten“ kostenlos?

Ja — der vollständige Text von „Ihre RAG-App testen und bewerten“ ist hier im Web kostenlos zu lesen. Um sie interaktiv zu üben (integrierter Code-Editor und 24/7 KI-Tutor) und den Rest des LLM Apps in Production (RAG + Vector DB + Caching)-Kurses freizuschalten, upgrade auf CoddyKit PRO. Der LLM Apps in Production (RAG + Vector DB + Caching)-Kurs umfasst insgesamt 4 Lektionen.

Was lerne ich in „Ihre RAG-App testen und bewerten“?

Schaffen Sie Vertrauen in Ihre erste RAG-Anwendung, indem Sie einen Testsatz erstellen und die Qualität von Retrieval und Antworten vor dem Release mit praxisnahen Metriken messen. Du übst LLM Apps in Production (RAG + Vector DB + Caching) mit praktischem Code, den du direkt im Browser ausführst, und ein 24/7 KI-Tutor beantwortet deine Fragen während du die Lektion bearbeitest.

Brauche ich Erfahrung, um LLM Apps in Production (RAG + Vector DB + Caching) zu starten?

Keine Vorkenntnisse erforderlich. LLM Apps in Production (RAG + Vector DB + Caching) auf CoddyKit ist für Anfänger bis fortgeschrittene Lernende strukturiert, sodass du hier starten oder von Anfang an beginnen und in deinem eigenen Tempo voranschreiten kannst. Dies ist Lektion 4 von 4.

Wie lange dauert die Lektion „Ihre RAG-App testen und bewerten“?

Die meisten CoddyKit-Lektionen dauern etwa 5–10 Minuten. Jede ist kompakt und interaktiv, sodass du stetig Fortschritte machst und genau dort weitermachst, wo du aufgehört hast – im Web und in der App.

Kann ich in dieser LLM Apps in Production (RAG + Vector DB + Caching)-Lektion Code schreiben und ausführen?

Ja. Jede LLM Apps in Production (RAG + Vector DB + Caching)-Lektion enthält einen integrierten Code-Editor, sodass du echten Code direkt in deinem Browser schreibst und ausführst und sofort KI-Feedback erhältst — ohne lokale Einrichtung erforderlich.

Alle Lektionen in diesem Kurs

  1. Auswahl eines LLM-Anbieters
  2. Grundlagen des Ladens und Aufteilens von Textdaten
  3. Eine einfache RAG-Pipeline erstellen
  4. Ihre RAG-App testen und bewerten
← Zurück zu LLM Apps in Production (RAG + Vector DB + Caching)