0Pricing
AI Engineering Academy · Lektion

Die Auswirkungen des Re-Rankings messen

Führen Sie einen Vorher-Nachher-Benchmark durch, der einstufigen Abruf mit zweistufigem Abruf und Re-Ranking vergleicht, und messen Sie NDCG, MRR sowie die Qualität der Antworten von Anfang bis Ende.

Die Auswirkungen des Re-Rankings messen ist eine kostenlose AI Engineering Academy-Lektion auf CoddyKit. Dies ist Lektion 4 von 4. Du kannst die komplette Lektion unten kostenlos lesen – dann übst du sie direkt im Browser mit einem integrierten Code-Editor und einem KI-Tutor rund um die Uhr. Sie ist Teil des AI Engineering Academy-Lernpfads, und dein Fortschritt wird über Web und CoddyKit-App synchronisiert. Der AI Engineering Academy-Kurs umfasst insgesamt 4 Lektionen.

Teile dieser Lektion wurden noch nicht übersetzt und werden auf Englisch angezeigt.

Why Measure Re-ranking Impact?

Re-ranking adds latency and cost to your pipeline. Without measurement, you cannot answer whether the added complexity is worth it. Benchmarking quantifies the improvement in retrieval quality and end-to-end answer quality so you can make an informed decision. It also reveals which query types benefit most, enabling you to apply re-ranking selectively rather than on every request.

Building a Golden Test Set

A reliable benchmark requires a golden test set: a collection of queries paired with the document IDs that are known to be relevant. Create it by sampling real user queries from your application logs, identifying the relevant documents manually or with expert annotation, and organizing them into a structured format. A test set of 50-200 queries is sufficient for most RAG evaluation purposes.

golden_test_set = [
    {
        'query': 'How does pgvector HNSW indexing improve search speed?',
        'relevant_doc_ids': ['doc_042', 'doc_107'],
    },
    {
        'query': 'What is the difference between BM25 and dense retrieval?',
        'relevant_doc_ids': ['doc_015'],
    },
    {
        'query': 'How to implement reciprocal rank fusion in Python?',
        'relevant_doc_ids': ['doc_093', 'doc_094'],
    },
    # ... 47 more entries
]

print(f'Test set size: {len(golden_test_set)} queries')
print(f'Avg relevant docs per query: {sum(len(e["relevant_doc_ids"]) for e in golden_test_set) / len(golden_test_set):.1f}')

Retrieval Metrics: NDCG, MRR, Hit Rate

Use three complementary metrics to evaluate retrieval quality. Hit Rate at K measures whether at least one relevant document appears in the top K results. MRR (Mean Reciprocal Rank) measures the average reciprocal of the rank of the first relevant document. NDCG at K (Normalized Discounted Cumulative Gain) measures ranking quality with higher positions weighted more than lower ones.

def compute_retrieval_metrics(results: list[str], relevant_ids: set, k: int = 5):
    results_at_k = results[:k]
    relevant_found = [r for r in results_at_k if r in relevant_ids]

    # Hit rate
    hit = 1 if relevant_found else 0

    # MRR
    rr = 0
    for i, doc_id in enumerate(results_at_k, start=1):
        if doc_id in relevant_ids:
            rr = 1.0 / i
            break

    # NDCG (binary relevance)
    import math
    dcg = sum(
        1.0 / math.log2(i + 1)
        for i, doc_id in enumerate(results_at_k, start=1)
        if doc_id in relevant_ids
    )
    ideal = sum(1.0 / math.log2(i + 1) for i in range(1, min(len(relevant_ids), k) + 1))
    ndcg = dcg / ideal if ideal > 0 else 0

    return {'hit': hit, 'rr': rr, 'ndcg': ndcg}

Baseline: Single-Stage Dense Retrieval

Before measuring the impact of re-ranking, establish a baseline using single-stage dense retrieval. Run every query in your test set through the bi-encoder retriever, collect the ranked document IDs, and compute average NDCG, MRR, and hit rate. This baseline tells you how much improvement you are starting from — if your baseline is already 0.95 NDCG@5, re-ranking has little room to improve.

def evaluate_pipeline(retriever_fn, test_set: list[dict], k: int = 5) -> dict:
    all_metrics = []

    for entry in test_set:
        query = entry['query']
        relevant = set(entry['relevant_doc_ids'])

        results = retriever_fn(query, top_k=k)
        result_ids = [r['id'] for r in results]

        metrics = compute_retrieval_metrics(result_ids, relevant, k)
        all_metrics.append(metrics)

    n = len(all_metrics)
    return {
        f'hit_rate@{k}': sum(m['hit'] for m in all_metrics) / n,
        f'mrr@{k}': sum(m['rr'] for m in all_metrics) / n,
        f'ndcg@{k}': sum(m['ndcg'] for m in all_metrics) / n,
    }

Running the Before-and-After Benchmark

Run the same evaluation function against both your single-stage retriever and your two-stage retriever with re-ranking. Print results side by side to make the improvement (or lack thereof) immediately visible. Track latency per query alongside quality metrics — a retrieval quality gain of 5 percent may not be worth a latency increase of 400ms depending on your application's SLA requirements.

import time

def evaluate_with_latency(retriever_fn, test_set, k=5):
    metrics_list = []
    latencies = []

    for entry in test_set:
        t0 = time.perf_counter()
        results = retriever_fn(entry['query'], top_k=k)
        latencies.append((time.perf_counter() - t0) * 1000)

        result_ids = [r['id'] for r in results]
        metrics_list.append(compute_retrieval_metrics(
            result_ids, set(entry['relevant_doc_ids']), k
        ))

    n = len(metrics_list)
    return {
        f'hit_rate@{k}': sum(m['hit'] for m in metrics_list) / n,
        f'ndcg@{k}': sum(m['ndcg'] for m in metrics_list) / n,
        'p50_latency_ms': sorted(latencies)[n // 2],
        'p99_latency_ms': sorted(latencies)[int(n * 0.99)],
    }

baseline = evaluate_with_latency(dense_retrieval_fn, golden_test_set)
two_stage = evaluate_with_latency(two_stage_fn, golden_test_set)
print('Baseline:', baseline)
print('Two-stage:', two_stage)

Interpreting NDCG Improvements

Typical improvements from adding cross-encoder re-ranking to dense retrieval range from 0.05 to 0.15 absolute NDCG@5, which translates to roughly 5-15 percentage points. The improvement is larger when: (1) your queries are diverse with many paraphrases, (2) your corpus has many near-relevant chunks, or (3) your first-stage retriever is weak. If you see less than 0.02 NDCG improvement, the benefit may not justify the added complexity.

# Interpreting benchmark results

example_results = {
    'baseline': {'hit_rate@5': 0.78, 'ndcg@5': 0.64, 'p99_latency_ms': 35},
    'two_stage': {'hit_rate@5': 0.89, 'ndcg@5': 0.77, 'p99_latency_ms': 287},
}

delta_ndcg = example_results['two_stage']['ndcg@5'] - example_results['baseline']['ndcg@5']
delta_latency = example_results['two_stage']['p99_latency_ms'] - example_results['baseline']['p99_latency_ms']

print(f'NDCG improvement: +{delta_ndcg:.2f} (+{delta_ndcg/example_results["baseline"]["ndcg@5"]*100:.0f}%)')
print(f'Latency increase: +{delta_latency}ms')
# NDCG improvement: +0.13 (+20%) — clearly worth the 252ms latency cost

End-to-End Answer Quality Measurement

Retrieval metrics measure whether the right documents were retrieved, but the ultimate measure is end-to-end answer quality. Use an LLM judge to evaluate whether answers generated from re-ranked context are more correct and faithful than answers from single-stage context. Score on a 1-5 scale for correctness, faithfulness, and relevance, then average across your test set.

from openai import OpenAI

client = OpenAI()

JUDGE_PROMPT = '''
Rate the following answer on a scale of 1-5 for correctness and faithfulness to the context.

Question: {question}
Context: {context}
Answer: {answer}
Ground truth: {ground_truth}

Return a JSON with fields: {"correctness": int, "faithfulness": int, "explanation": str}
'''

def judge_answer(question, context, answer, ground_truth):
    prompt = JUDGE_PROMPT.format(
        question=question, context=context,
        answer=answer, ground_truth=ground_truth,
    )
    response = client.chat.completions.create(
        model='gpt-4o',
        messages=[{'role': 'user', 'content': prompt}],
        response_format={'type': 'json_object'},
    )
    import json
    return json.loads(response.choices[0].message.content)

Stratified Analysis by Query Type

Average metrics hide important differences across query types. Segment your test set into categories — factual lookups (who, what, when), procedural queries (how to), conceptual queries (why, explain), and technical queries (error codes, API names) — and compute metrics separately for each group. Re-ranking often helps most on conceptual and procedural queries where semantic understanding matters more than keyword matching.

def stratified_eval(retriever_fn, test_set, k=5):
    groups = {'factual': [], 'procedural': [], 'conceptual': [], 'technical': []}

    for entry in test_set:
        q = entry['query'].lower()
        if any(w in q for w in ['how to', 'how do', 'steps to']):
            groups['procedural'].append(entry)
        elif any(w in q for w in ['why', 'explain', 'what is the reason']):
            groups['conceptual'].append(entry)
        elif any(c.isupper() for c in q.split()) or 'error' in q:
            groups['technical'].append(entry)
        else:
            groups['factual'].append(entry)

    for group_name, group_entries in groups.items():
        if group_entries:
            metrics = evaluate_pipeline(retriever_fn, group_entries, k)
            print(f'{group_name} ({len(group_entries)} queries): ndcg@{k}={metrics[f"ndcg@{k}"]:.3f}')

Regression Testing with CI Integration

Run your retrieval benchmark as a regression test in CI. Set minimum acceptable thresholds for NDCG@5, MRR, and hit rate. Any pipeline change that causes metrics to drop below the threshold fails the CI build, preventing retrieval quality regressions from shipping to production. This is especially important after changing chunk sizes, embedding models, or re-ranking models.

# pytest integration for retrieval quality gates
import pytest

MIN_NDCG_5 = 0.70
MIN_HIT_RATE_5 = 0.85

def test_retrieval_quality_meets_threshold():
    metrics = evaluate_pipeline(production_retriever_fn, golden_test_set, k=5)
    assert metrics['ndcg@5'] >= MIN_NDCG_5, (
        f'NDCG@5 {metrics["ndcg@5"]:.3f} below threshold {MIN_NDCG_5}'
    )
    assert metrics['hit_rate@5'] >= MIN_HIT_RATE_5, (
        f'Hit rate {metrics["hit_rate@5"]:.3f} below threshold {MIN_HIT_RATE_5}'
    )

# Run with: pytest tests/test_retrieval.py -v

Visualizing Retrieval Metrics

Raw numbers are hard to interpret across multiple experiments. Create a simple comparison table or bar chart that shows NDCG, MRR, hit rate, and latency side by side for baseline, hybrid-only, and hybrid-with-reranking pipelines. Tracking these metrics over time as you make improvements creates a retrieval improvement history that guides future optimization decisions.

def print_comparison_table(results: dict[str, dict]):
    headers = ['Pipeline', 'NDCG@5', 'MRR@5', 'Hit@5', 'P99 ms']
    print('|'.join(f'{h:20}' for h in headers))
    print('-' * (len(headers) * 21))
    for pipeline_name, metrics in results.items():
        row = [
            pipeline_name,
            f'{metrics.get("ndcg@5", 0):.3f}',
            f'{metrics.get("mrr@5", 0):.3f}',
            f'{metrics.get("hit_rate@5", 0):.3f}',
            f'{metrics.get("p99_latency_ms", 0):.0f}',
        ]
        print('|'.join(f'{v:20}' for v in row))

results = {
    'Dense only': {'ndcg@5': 0.64, 'mrr@5': 0.68, 'hit_rate@5': 0.78, 'p99_latency_ms': 35},
    'Hybrid RRF': {'ndcg@5': 0.71, 'mrr@5': 0.74, 'hit_rate@5': 0.84, 'p99_latency_ms': 55},
    'Hybrid + Rerank': {'ndcg@5': 0.77, 'mrr@5': 0.81, 'hit_rate@5': 0.89, 'p99_latency_ms': 287},
}
print_comparison_table(results)

Acting on Benchmark Results

After running benchmarks, use the results to make concrete decisions. If re-ranking improves NDCG by less than 0.03, skip it and focus on improving the first stage. If hit rate is low, your first stage is missing relevant documents — increase the candidate set size or switch to hybrid retrieval. If end-to-end answer quality improves significantly despite modest retrieval gains, the re-ranker may be surfacing highly relevant sentences the LLM uses effectively even if ranked position is unchanged.

Quick Check

Test your understanding of measuring retrieval and re-ranking impact from this lesson.

Lesson Recap

In this lesson you learned: building a golden test set with known relevant documents is essential before measuring retrieval quality, NDCG, MRR, and hit rate are the three core retrieval metrics that together measure ranking quality comprehensively, and end-to-end answer quality using an LLM judge provides the ultimate measure of pipeline improvement. Always run retrieval benchmarks as regression tests in CI. Next up we explore LLM streaming to display tokens as they are generated.

Häufig gestellte Fragen

Ist die Lektion „Die Auswirkungen des Re-Rankings messen“ kostenlos?

Ja — der vollständige Text von „Die Auswirkungen des Re-Rankings messen“ ist hier im Web kostenlos zu lesen. Um sie interaktiv zu üben (integrierter Code-Editor und 24/7 KI-Tutor) und den Rest des AI Engineering Academy-Kurses freizuschalten, upgrade auf CoddyKit PRO. Der AI Engineering Academy-Kurs umfasst insgesamt 4 Lektionen.

Was lerne ich in „Die Auswirkungen des Re-Rankings messen“?

Führen Sie einen Vorher-Nachher-Benchmark durch, der einstufigen Abruf mit zweistufigem Abruf und Re-Ranking vergleicht, und messen Sie NDCG, MRR sowie die Qualität der Antworten von Anfang bis Ende. Du übst AI Engineering Academy mit praktischem Code, den du direkt im Browser ausführst, und ein 24/7 KI-Tutor beantwortet deine Fragen während du die Lektion bearbeitest.

Brauche ich Erfahrung, um AI Engineering Academy zu starten?

Keine Vorkenntnisse erforderlich. AI Engineering Academy auf CoddyKit ist für Anfänger bis fortgeschrittene Lernende strukturiert, sodass du hier starten oder von Anfang an beginnen und in deinem eigenen Tempo voranschreiten kannst. Dies ist Lektion 4 von 4.

Wie lange dauert die Lektion „Die Auswirkungen des Re-Rankings messen“?

Die meisten CoddyKit-Lektionen dauern etwa 5–10 Minuten. Jede ist kompakt und interaktiv, sodass du stetig Fortschritte machst und genau dort weitermachst, wo du aufgehört hast – im Web und in der App.

Kann ich in dieser AI Engineering Academy-Lektion Code schreiben und ausführen?

Ja. Jede AI Engineering Academy-Lektion enthält einen integrierten Code-Editor, sodass du echten Code direkt in deinem Browser schreibst und ausführst und sofort KI-Feedback erhältst — ohne lokale Einrichtung erforderlich.

Alle Lektionen in diesem Kurs

  1. Warum zweistufiger Abruf funktioniert
  2. Cross-Encoder-Re-Ranking mit Cohere und BGE
  3. Kontextuelle Komprimierung und Relevanzfilterung
  4. Die Auswirkungen des Re-Rankings messen
← Zurück zu AI Engineering Academy