0Pricing
Spring Boot 4 Complete Guide · 강의

검색 증강 생성 파이프라인

검색한 문서의 맥락을 바탕으로 모델 응답을 생성하는 RAG 흐름을 구성합니다.

검색 증강 생성 파이프라인은(는) CoddyKit의 무료 Spring Boot 4 Complete Guide 강의입니다. 이것은 4개 중 3번째 강의입니다. 아래에서 전체 강의를 무료로 읽을 수 있으며, 내장 코드 에디터와 24/7 AI 튜터와 함께 브라우저에서 직접 실습할 수 있습니다. 이 강의는 Spring Boot 4 Complete Guide 학습 경로의 일부이며, 진행 상황이 웹과 CoddyKit 앱에 동기화됩니다. Spring Boot 4 Complete Guide 강의에는 총 4개의 강의가 포함되어 있습니다.

이 강의의 일부는 아직 번역되지 않았으며 영어로 표시됩니다.

Why RAG?

Retrieval-Augmented Generation (RAG) grounds an LLM's answers in your own documents instead of relying solely on what the model memorized during training.

  • Fresh & private data — answer questions about internal docs the model never saw.
  • Less hallucination — the model cites retrieved context rather than inventing facts.
  • Cheaper than fine-tuning — you update a vector store, not model weights.

A Spring AI RAG pipeline has two phases: an ingestion phase (read → split → embed → store) and a query phase (embed question → retrieve → augment prompt → generate).

The Pipeline at a Glance

Spring AI gives you composable building blocks for both phases. The core types you will assemble are:

  • DocumentReader — loads raw sources (PDF, Markdown, JSON, web pages).
  • DocumentTransformer — splits documents into chunks (e.g. TokenTextSplitter).
  • EmbeddingModel — turns text into vectors.
  • VectorStore — stores and similarity-searches those vectors.
  • ChatClient with a RAG advisor — wires retrieval into the prompt automatically.

The first three feed ingestion; the last two power querying.

Ingestion: Read and Split

During ingestion you read source files and split them into chunks small enough to fit the model's context window while staying semantically coherent. TokenTextSplitter chunks by token count with overlap so meaning isn't cut mid-sentence.

Below, a Markdown file is read and split into ~800-token chunks before storage.

@Component
class DocumentIngestor {

    private final VectorStore vectorStore;

    DocumentIngestor(VectorStore vectorStore) {
        this.vectorStore = vectorStore;
    }

    void ingest(Resource markdown) {
        var reader = new TextReader(markdown);
        List<Document> raw = reader.get();

        var splitter = new TokenTextSplitter(800, 350, 5, 10000, true);
        List<Document> chunks = splitter.apply(raw);

        vectorStore.add(chunks);
    }
}

Embeddings: Text to Vectors

An embedding is a dense vector that captures the meaning of text. Two passages about the same topic land close together in vector space, which is what makes similarity search work.

Spring AI auto-configures an EmbeddingModel based on your starter (OpenAI, Azure, Ollama, etc.). You rarely call it directly — the VectorStore uses it internally — but you can:

@Service
class EmbeddingDemo {

    private final EmbeddingModel embeddingModel;

    EmbeddingDemo(EmbeddingModel embeddingModel) {
        this.embeddingModel = embeddingModel;
    }

    float[] embed(String text) {
        return embeddingModel.embed(text);
    }

    int dimensions() {
        return embeddingModel.dimensions();
    }
}

Configuring a VectorStore

The VectorStore persists embeddings and runs similarity searches. Spring AI ships adapters for PgVector, Redis, Chroma, Qdrant, Milvus, and a simple in-memory store.

With the spring-ai-starter-vector-store-pgvector dependency, a bean is auto-configured from properties — no manual wiring needed:

spring:
  ai:
    vectorstore:
      pgvector:
        initialize-schema: true
        index-type: HNSW
        distance-type: COSINE_DISTANCE
        dimensions: 1536
  datasource:
    url: jdbc:postgresql://localhost:5432/ragdb
    username: rag
    password: secret

Querying: Similarity Search

At query time you turn the user's question into a vector and ask the store for the nearest chunks. A SearchRequest controls how many results (topK) and a minimum similarityThreshold to filter out weak matches.

List<Document> retrieve(VectorStore store, String question) {
    var request = SearchRequest.builder()
            .query(question)
            .topK(4)
            .similarityThreshold(0.7)
            .build();

    List<Document> hits = store.similaritySearch(request);
    return hits;
}

Augmenting the Prompt Manually

Before reaching for advisors, it helps to see the mechanic. RAG simply stuffs the retrieved text into the prompt and instructs the model to answer only from it. This is the "augment" step:

String answer(ChatClient chat, VectorStore store, String question) {
    List<Document> docs = store.similaritySearch(
            SearchRequest.builder().query(question).topK(4).build());

    String context = docs.stream()
            .map(Document::getText)
            .collect(Collectors.joining("\n---\n"));

    return chat.prompt()
            .system("Answer using ONLY the context. If unknown, say you don't know.")
            .user(u -> u.text("Context:\n{ctx}\n\nQuestion: {q}")
                        .param("ctx", context)
                        .param("q", question))
            .call()
            .content();
}

The QuestionAnswerAdvisor

Spring AI packages that manual pattern as the QuestionAnswerAdvisor. You attach it to a ChatClient and it transparently runs the similarity search and injects context for every call.

This is the idiomatic, declarative way to do RAG in Spring Boot 4:

@Service
class RagAssistant {

    private final ChatClient chatClient;

    RagAssistant(ChatClient.Builder builder, VectorStore vectorStore) {
        this.chatClient = builder
                .defaultAdvisors(new QuestionAnswerAdvisor(vectorStore))
                .build();
    }

    String ask(String question) {
        return chatClient.prompt()
                .user(question)
                .call()
                .content();
    }
}

Per-Request Retrieval Tuning

Defaults are convenient, but real apps tune retrieval per request — a focused FAQ may need topK=2, a broad research query topK=8. Pass an advisor at call time and override its SearchRequest, plus metadata filters to scope to a tenant or document set:

String ask(ChatClient chatClient, VectorStore store, String q, String tenant) {
    var advisor = QuestionAnswerAdvisor.builder(store)
            .searchRequest(SearchRequest.builder()
                    .topK(6)
                    .similarityThreshold(0.75)
                    .filterExpression("tenant == '" + tenant + "'")
                    .build())
            .build();

    return chatClient.prompt()
            .advisors(advisor)
            .user(q)
            .call()
            .content();
}

Modular RAG with RetrievalAugmentationAdvisor

For advanced pipelines, Spring AI 1.0 offers the Modular RAG API via RetrievalAugmentationAdvisor. It exposes each stage as a swappable component:

  • QueryTransformer — rewrite, compress, or translate the query.
  • DocumentRetriever — the source of chunks (e.g. VectorStoreDocumentRetriever).
  • QueryAugmenter — control how context is merged and how empty-context is handled.
var retriever = VectorStoreDocumentRetriever.builder()
        .vectorStore(vectorStore)
        .similarityThreshold(0.72)
        .topK(5)
        .build();

var ragAdvisor = RetrievalAugmentationAdvisor.builder()
        .queryTransformers(RewriteQueryTransformer.builder()
                .chatClientBuilder(chatClientBuilder)
                .build())
        .documentRetriever(retriever)
        .build();

String answer = chatClient.prompt()
        .advisors(ragAdvisor)
        .user(question)
        .call()
        .content();

Guarding Against Empty Context

A subtle RAG failure: when retrieval finds nothing relevant, a naive prompt lets the model fall back to its training data and hallucinate. The ContextualQueryAugmenter lets you decide that policy explicitly.

Set allowEmptyContext(false) to force a graceful "I don't have information on that" instead of a confident guess:

var augmenter = ContextualQueryAugmenter.builder()
        .allowEmptyContext(false)
        .build();

var ragAdvisor = RetrievalAugmentationAdvisor.builder()
        .documentRetriever(retriever)
        .queryAugmenter(augmenter)
        .build();

// With no matching documents, the model returns a safe
// "no answer available" response instead of fabricating one.

Quick Check

You build a RAG assistant. When a user asks something outside your knowledge base, similarity search returns no chunks above the threshold, yet the model still answers confidently with made-up facts. Which change best fixes this?

Recap

You assembled a complete RAG pipeline in Spring AI:

  • Ingestion — DocumentReader → TokenTextSplitter → EmbeddingModel → VectorStore.add().
  • Query — embed the question, similaritySearch with topK and similarityThreshold, then augment the prompt.
  • Declarative RAG — QuestionAnswerAdvisor wires retrieval into a ChatClient automatically; tune it per request with custom SearchRequest and metadata filters.
  • Modular RAG — RetrievalAugmentationAdvisor with query transformers, retrievers, and augmenters for advanced control.
  • Safety — ContextualQueryAugmenter.allowEmptyContext(false) prevents hallucination when nothing relevant is retrieved.

Grounding model output in retrieved context is the single most effective way to make LLM features trustworthy.

자주 묻는 질문

“검색 증강 생성 파이프라인” 강의는 무료인가요?

네 — “검색 증강 생성 파이프라인” 전체 내용을 이 웹사이트에서 무료로 읽을 수 있습니다. 인터랙티브하게 실습하려면(내장 코드 에디터와 24/7 AI 튜터), CoddyKit PRO로 업그레이드하면 Spring Boot 4 Complete Guide 강의 전체를 잠금 해제할 수 있습니다. Spring Boot 4 Complete Guide 강의에는 총 4개의 강의가 포함되어 있습니다.

“검색 증강 생성 파이프라인”에서 뭘 배우나요?

검색한 문서의 맥락을 바탕으로 모델 응답을 생성하는 RAG 흐름을 구성합니다. 브라우저에서 직접 실행하는 실습 코드로 Spring Boot 4 Complete Guide을(를) 배우며, 24/7 AI 튜터가 강의를 진행하면서 질문에 답변해줍니다.

Spring Boot 4 Complete Guide을(를) 시작하는 데 경험이 필요한가요?

사전 경험은 필요하지 않습니다. CoddyKit의 Spring Boot 4 Complete Guide은(는) 초급자부터 고급 학습자까지를 위해 구성되어 있으므로, 여기서 시작하거나 처음부터 시작할 수 있으며 자신의 속도대로 진행할 수 있습니다. 이것은 4개 중 3번째 강의입니다.

“검색 증강 생성 파이프라인” 강의는 얼마나 걸리나요?

대부분의 CoddyKit 강의는 약 5~10분이 소요됩니다. 각 강의는 간결하고 인터랙티브하여 꾸준한 진행이 가능하며, 웹과 앱에서 중단한 부분부터 바로 시작할 수 있습니다.

이 Spring Boot 4 Complete Guide 강의에서 코드를 작성하고 실행할 수 있나요?

네. 모든 Spring Boot 4 Complete Guide 강의에는 내장 코드 에디터가 포함되어 있으므로, 브라우저에서 바로 실제 코드를 작성하고 실행한 후 즉시 AI 피드백을 받을 수 있습니다 — 로컬 설정이 필요 없습니다.

이 강의의 모든 강의

  1. ChatClient, 프롬프트 및 구조화된 출력
  2. 임베딩 및 벡터 저장소 검색
  3. 검색 증강 생성 파이프라인
  4. 도구 호출 및 에이전트 조언자
← Spring Boot 4 Complete Guide(으)로 돌아가기