检索增强生成流水线
组装 RAG 流程,让模型响应以检索到的文档上下文为依据。
检索增强生成流水线 是 CoddyKit 上的免费 Spring Boot 4 Complete Guide 课时。 这是第 3 节课,共 4 节。 你可以在下方免费阅读本课时的完整内容 — 然后在浏览器中使用内置代码编辑器和全天候 AI 导师进行实践。 这是 Spring Boot 4 Complete Guide 学习路径的一部分,你的进度在网页和 CoddyKit 应用中同步。 Spring Boot 4 Complete Guide 课程共包含 4 节课。
本课时的部分内容尚未翻译,以英文显示。
Why RAG?
Retrieval-Augmented Generation (RAG) grounds an LLM's answers in your own documents instead of relying solely on what the model memorized during training.
- Fresh & private data — answer questions about internal docs the model never saw.
- Less hallucination — the model cites retrieved context rather than inventing facts.
- Cheaper than fine-tuning — you update a vector store, not model weights.
A Spring AI RAG pipeline has two phases: an ingestion phase (read → split → embed → store) and a query phase (embed question → retrieve → augment prompt → generate).
The Pipeline at a Glance
Spring AI gives you composable building blocks for both phases. The core types you will assemble are:
DocumentReader— loads raw sources (PDF, Markdown, JSON, web pages).DocumentTransformer— splits documents into chunks (e.g.TokenTextSplitter).EmbeddingModel— turns text into vectors.VectorStore— stores and similarity-searches those vectors.ChatClientwith a RAG advisor — wires retrieval into the prompt automatically.
The first three feed ingestion; the last two power querying.
Ingestion: Read and Split
During ingestion you read source files and split them into chunks small enough to fit the model's context window while staying semantically coherent. TokenTextSplitter chunks by token count with overlap so meaning isn't cut mid-sentence.
Below, a Markdown file is read and split into ~800-token chunks before storage.
@Component
class DocumentIngestor {
private final VectorStore vectorStore;
DocumentIngestor(VectorStore vectorStore) {
this.vectorStore = vectorStore;
}
void ingest(Resource markdown) {
var reader = new TextReader(markdown);
List<Document> raw = reader.get();
var splitter = new TokenTextSplitter(800, 350, 5, 10000, true);
List<Document> chunks = splitter.apply(raw);
vectorStore.add(chunks);
}
}Embeddings: Text to Vectors
An embedding is a dense vector that captures the meaning of text. Two passages about the same topic land close together in vector space, which is what makes similarity search work.
Spring AI auto-configures an EmbeddingModel based on your starter (OpenAI, Azure, Ollama, etc.). You rarely call it directly — the VectorStore uses it internally — but you can:
@Service
class EmbeddingDemo {
private final EmbeddingModel embeddingModel;
EmbeddingDemo(EmbeddingModel embeddingModel) {
this.embeddingModel = embeddingModel;
}
float[] embed(String text) {
return embeddingModel.embed(text);
}
int dimensions() {
return embeddingModel.dimensions();
}
}Configuring a VectorStore
The VectorStore persists embeddings and runs similarity searches. Spring AI ships adapters for PgVector, Redis, Chroma, Qdrant, Milvus, and a simple in-memory store.
With the spring-ai-starter-vector-store-pgvector dependency, a bean is auto-configured from properties — no manual wiring needed:
spring:
ai:
vectorstore:
pgvector:
initialize-schema: true
index-type: HNSW
distance-type: COSINE_DISTANCE
dimensions: 1536
datasource:
url: jdbc:postgresql://localhost:5432/ragdb
username: rag
password: secretQuerying: Similarity Search
At query time you turn the user's question into a vector and ask the store for the nearest chunks. A SearchRequest controls how many results (topK) and a minimum similarityThreshold to filter out weak matches.
List<Document> retrieve(VectorStore store, String question) {
var request = SearchRequest.builder()
.query(question)
.topK(4)
.similarityThreshold(0.7)
.build();
List<Document> hits = store.similaritySearch(request);
return hits;
}Augmenting the Prompt Manually
Before reaching for advisors, it helps to see the mechanic. RAG simply stuffs the retrieved text into the prompt and instructs the model to answer only from it. This is the "augment" step:
String answer(ChatClient chat, VectorStore store, String question) {
List<Document> docs = store.similaritySearch(
SearchRequest.builder().query(question).topK(4).build());
String context = docs.stream()
.map(Document::getText)
.collect(Collectors.joining("\n---\n"));
return chat.prompt()
.system("Answer using ONLY the context. If unknown, say you don't know.")
.user(u -> u.text("Context:\n{ctx}\n\nQuestion: {q}")
.param("ctx", context)
.param("q", question))
.call()
.content();
}The QuestionAnswerAdvisor
Spring AI packages that manual pattern as the QuestionAnswerAdvisor. You attach it to a ChatClient and it transparently runs the similarity search and injects context for every call.
This is the idiomatic, declarative way to do RAG in Spring Boot 4:
@Service
class RagAssistant {
private final ChatClient chatClient;
RagAssistant(ChatClient.Builder builder, VectorStore vectorStore) {
this.chatClient = builder
.defaultAdvisors(new QuestionAnswerAdvisor(vectorStore))
.build();
}
String ask(String question) {
return chatClient.prompt()
.user(question)
.call()
.content();
}
}Per-Request Retrieval Tuning
Defaults are convenient, but real apps tune retrieval per request — a focused FAQ may need topK=2, a broad research query topK=8. Pass an advisor at call time and override its SearchRequest, plus metadata filters to scope to a tenant or document set:
String ask(ChatClient chatClient, VectorStore store, String q, String tenant) {
var advisor = QuestionAnswerAdvisor.builder(store)
.searchRequest(SearchRequest.builder()
.topK(6)
.similarityThreshold(0.75)
.filterExpression("tenant == '" + tenant + "'")
.build())
.build();
return chatClient.prompt()
.advisors(advisor)
.user(q)
.call()
.content();
}Modular RAG with RetrievalAugmentationAdvisor
For advanced pipelines, Spring AI 1.0 offers the Modular RAG API via RetrievalAugmentationAdvisor. It exposes each stage as a swappable component:
QueryTransformer— rewrite, compress, or translate the query.DocumentRetriever— the source of chunks (e.g.VectorStoreDocumentRetriever).QueryAugmenter— control how context is merged and how empty-context is handled.
var retriever = VectorStoreDocumentRetriever.builder()
.vectorStore(vectorStore)
.similarityThreshold(0.72)
.topK(5)
.build();
var ragAdvisor = RetrievalAugmentationAdvisor.builder()
.queryTransformers(RewriteQueryTransformer.builder()
.chatClientBuilder(chatClientBuilder)
.build())
.documentRetriever(retriever)
.build();
String answer = chatClient.prompt()
.advisors(ragAdvisor)
.user(question)
.call()
.content();Guarding Against Empty Context
A subtle RAG failure: when retrieval finds nothing relevant, a naive prompt lets the model fall back to its training data and hallucinate. The ContextualQueryAugmenter lets you decide that policy explicitly.
Set allowEmptyContext(false) to force a graceful "I don't have information on that" instead of a confident guess:
var augmenter = ContextualQueryAugmenter.builder()
.allowEmptyContext(false)
.build();
var ragAdvisor = RetrievalAugmentationAdvisor.builder()
.documentRetriever(retriever)
.queryAugmenter(augmenter)
.build();
// With no matching documents, the model returns a safe
// "no answer available" response instead of fabricating one.Quick Check
You build a RAG assistant. When a user asks something outside your knowledge base, similarity search returns no chunks above the threshold, yet the model still answers confidently with made-up facts. Which change best fixes this?
Recap
You assembled a complete RAG pipeline in Spring AI:
- Ingestion —
DocumentReader→TokenTextSplitter→EmbeddingModel→VectorStore.add(). - Query — embed the question,
similaritySearchwithtopKandsimilarityThreshold, then augment the prompt. - Declarative RAG —
QuestionAnswerAdvisorwires retrieval into aChatClientautomatically; tune it per request with customSearchRequestand metadata filters. - Modular RAG —
RetrievalAugmentationAdvisorwith query transformers, retrievers, and augmenters for advanced control. - Safety —
ContextualQueryAugmenter.allowEmptyContext(false)prevents hallucination when nothing relevant is retrieved.
Grounding model output in retrieved context is the single most effective way to make LLM features trustworthy.
常见问题解答
「检索增强生成流水线」课时是免费的吗?
是的 — 「检索增强生成流水线」的完整文本可在网页上免费阅读。要进行交互式练习(内置代码编辑器和全天候 AI 导师)并解锁 Spring Boot 4 Complete Guide 课程的其余内容,请升级到 CoddyKit PRO。 Spring Boot 4 Complete Guide 课程共包含 4 节课。
「检索增强生成流水线」这节课中我会学到什么?
组装 RAG 流程,让模型响应以检索到的文档上下文为依据。 你通过在浏览器中直接运行的动手代码来练习 Spring Boot 4 Complete Guide,全天候 AI 导师会在你学习这节课的过程中回答你的问题。
学习 Spring Boot 4 Complete Guide 需要有经验吗?
无需任何先前经验。CoddyKit 上的 Spring Boot 4 Complete Guide 课程适合初学者到高级学习者,你可以从这里开始或从头开始,按照自己的节奏学习。 这是第 3 节课,共 4 节。
「检索增强生成流水线」课时需要多长时间?
大多数 CoddyKit 课程大约需要 5–10 分钟。每节课都很精短且互动,所以你能稳步进步,并在网页和应用中从离开的地方继续。
我能在这节 Spring Boot 4 Complete Guide 课中编写并运行代码吗?
能。每节 Spring Boot 4 Complete Guide 课都包含内置代码编辑器,你可以在浏览器中直接编写并运行真实代码,并获得即时 AI 反馈 — 无需本地设置。