0Pricing
Prompt Engineering & LLM Optimization for Developers · Lektion

Grundlagen von Retrieval-Augmented Generation (RAG)

Lernen Sie, wie RAG LLM-Antworten auf Ihre eigenen Dokumente stützt, indem es zum Anfragezeitpunkt relevanten Kontext abruft und in den Prompt einfügt

Grundlagen von Retrieval-Augmented Generation (RAG) ist eine kostenlose Prompt Engineering & LLM Optimization for Developers-Lektion auf CoddyKit. Dies ist Lektion 4 von 4. Du kannst die komplette Lektion unten kostenlos lesen – dann übst du sie direkt im Browser mit einem integrierten Code-Editor und einem KI-Tutor rund um die Uhr. Sie ist Teil des Prompt Engineering & LLM Optimization for Developers-Lernpfads, und dein Fortschritt wird über Web und CoddyKit-App synchronisiert. Der Prompt Engineering & LLM Optimization for Developers-Kurs umfasst insgesamt 4 Lektionen.

Teile dieser Lektion wurden noch nicht übersetzt und werden auf Englisch angezeigt.

The Knowledge Gap

An LLM only knows what it was trained on. It cannot answer questions about your private docs or recent events. RAG (Retrieval-Augmented Generation) closes this gap by fetching relevant text and putting it in the prompt.

The Core Idea

Instead of fine-tuning the model on your data, you retrieve the most relevant snippets at query time and let the model answer using them as context. Cheaper, faster to update, and easy to cite.

Step 1: Chunking

Documents are split into small chunks (a few hundred tokens each). Chunks small enough to be precise, large enough to keep meaning.

chunks = split(document, size=500, overlap=50)

Step 2: Embeddings

Each chunk is converted to a vector with an embedding model. Similar meanings produce nearby vectors, enabling semantic search.

vector = embed("Refunds are processed in 5 days.")
// -> [0.012, -0.43, 0.88, ...]

Step 3: The Vector Store

Vectors are saved in a vector database like Pinecone, Weaviate, or pgvector. It supports fast nearest-neighbor search over millions of chunks.

Step 4: Retrieve at Query Time

When a user asks a question, you embed the question and find the top-k most similar chunks.

q = embed(userQuestion);
results = store.search(q, topK=4);

Step 5: Augment the Prompt

The retrieved chunks are inserted into the prompt as context, and the model is told to answer using only that context.

Use only the context below to answer.
Context:
{retrieved_chunks}
Question: {user_question}

Grounding and Citations

Because the answer is built from real chunks, you can show citations back to the source documents, and you can instruct the model to say I do not know when the context lacks the answer.

Why Not Just Fine-Tune?

  • RAG updates instantly: change a doc, re-index, done.
  • Fine-tuning is slow and bakes knowledge in.
  • RAG gives traceable sources; fine-tuning does not.

Common RAG Problems

Poor chunking, weak embeddings, or retrieving too few chunks all hurt quality. If answers are wrong, inspect what was retrieved first; the issue is usually retrieval, not the model.

Improving Retrieval

  • Add overlap between chunks.
  • Use hybrid keyword + vector search.
  • Re-rank results with a cross-encoder.
  • Tune top-k for your context window.

Quick Check

Test your understanding of RAG.

Recap

RAG chunks documents, embeds them into a vector store, retrieves the most relevant chunks for each query, and augments the prompt. It grounds answers, enables citations, and stays current without retraining.

Häufig gestellte Fragen

Ist die Lektion „Grundlagen von Retrieval-Augmented Generation (RAG)“ kostenlos?

Ja — der vollständige Text von „Grundlagen von Retrieval-Augmented Generation (RAG)“ ist hier im Web kostenlos zu lesen. Um sie interaktiv zu üben (integrierter Code-Editor und 24/7 KI-Tutor) und den Rest des Prompt Engineering & LLM Optimization for Developers-Kurses freizuschalten, upgrade auf CoddyKit PRO. Der Prompt Engineering & LLM Optimization for Developers-Kurs umfasst insgesamt 4 Lektionen.

Was lerne ich in „Grundlagen von Retrieval-Augmented Generation (RAG)“?

Lernen Sie, wie RAG LLM-Antworten auf Ihre eigenen Dokumente stützt, indem es zum Anfragezeitpunkt relevanten Kontext abruft und in den Prompt einfügt Du übst Prompt Engineering & LLM Optimization for Developers mit praktischem Code, den du direkt im Browser ausführst, und ein 24/7 KI-Tutor beantwortet deine Fragen während du die Lektion bearbeitest.

Brauche ich Erfahrung, um Prompt Engineering & LLM Optimization for Developers zu starten?

Keine Vorkenntnisse erforderlich. Prompt Engineering & LLM Optimization for Developers auf CoddyKit ist für Anfänger bis fortgeschrittene Lernende strukturiert, sodass du hier starten oder von Anfang an beginnen und in deinem eigenen Tempo voranschreiten kannst. Dies ist Lektion 4 von 4.

Wie lange dauert die Lektion „Grundlagen von Retrieval-Augmented Generation (RAG)“?

Die meisten CoddyKit-Lektionen dauern etwa 5–10 Minuten. Jede ist kompakt und interaktiv, sodass du stetig Fortschritte machst und genau dort weitermachst, wo du aufgehört hast – im Web und in der App.

Kann ich in dieser Prompt Engineering & LLM Optimization for Developers-Lektion Code schreiben und ausführen?

Ja. Jede Prompt Engineering & LLM Optimization for Developers-Lektion enthält einen integrierten Code-Editor, sodass du echten Code direkt in deinem Browser schreibst und ausführst und sofort KI-Feedback erhältst — ohne lokale Einrichtung erforderlich.

Alle Lektionen in diesem Kurs

  1. Mit LLM-APIs interagieren (OpenAI, Anthropic)
  2. Grundlagen von LangChain und LlamaIndex
  3. Prompt-Verwaltung und Versionierung
  4. Grundlagen von Retrieval-Augmented Generation (RAG)
← Zurück zu Prompt Engineering & LLM Optimization for Developers