嵌入与向量存储检索
生成嵌入并查询向量存储,为数据提供语义搜索能力。
嵌入与向量存储检索 是 CoddyKit 上的免费 Spring Boot 4 Complete Guide 课时。 这是第 2 节课,共 4 节。 你可以在下方免费阅读本课时的完整内容 — 然后在浏览器中使用内置代码编辑器和全天候 AI 导师进行实践。 这是 Spring Boot 4 Complete Guide 学习路径的一部分,你的进度在网页和 CoddyKit 应用中同步。 Spring Boot 4 Complete Guide 课程共包含 4 节课。
本课时的部分内容尚未翻译,以英文显示。
Why Embeddings Power Semantic Search
Keyword search matches literal tokens. Semantic search matches meaning. The bridge between the two is an embedding: a fixed-length vector of floats that captures the semantic content of a piece of text.
- Texts with similar meaning produce vectors that are close together in vector space.
- Closeness is measured by cosine similarity or Euclidean distance.
- A query like
"how do I reset my password"can match a document titled"account recovery steps"even with zero shared keywords.
In Spring AI, you generate embeddings with an EmbeddingModel and store/query them with a VectorStore. This lesson wires both together into a Retrieval-Augmented pipeline.
The EmbeddingModel Abstraction
Spring AI exposes EmbeddingModel as a provider-agnostic interface. Add the starter (for example spring-ai-starter-model-openai) and Spring Boot auto-configures a bean you can inject.
embed(String)returns a singlefloat[].embed(List<String>)batches multiple texts efficiently.dimensions()tells you the vector size (for example 1536 fortext-embedding-3-small).
Configure the model in application.yml so the bean is ready for injection.
spring:
ai:
openai:
api-key: ${OPENAI_API_KEY}
embedding:
options:
model: text-embedding-3-smallGenerating Your First Embedding
Inject EmbeddingModel and call embed. The result is a dense vector you can inspect or persist. The richer embedForResponse call also returns token usage metadata.
- Always log
dimensions()once at startup so a model swap that changes vector size is caught early. - A vector store column must match this dimension exactly, or inserts fail.
@RestController
class EmbeddingController {
private final EmbeddingModel embeddingModel;
EmbeddingController(EmbeddingModel embeddingModel) {
this.embeddingModel = embeddingModel;
}
@GetMapping("/embed")
Map<String, Object> embed(@RequestParam String text) {
float[] vector = embeddingModel.embed(text);
return Map.of(
"dimensions", vector.length,
"preview", List.of(vector[0], vector[1], vector[2])
);
}
}Cosine Similarity By Hand
To build intuition, here is the math a vector store does for you. Cosine similarity is the dot product of two vectors divided by the product of their magnitudes. It ranges from -1 (opposite) to 1 (identical direction).
- A value near 1.0 means the texts are semantically very close.
- Vector stores usually expose a similarity score derived from this metric.
This standalone program computes cosine similarity for two toy vectors.
public class CosineSimilarity {
static double cosine(double[] a, double[] b) {
double dot = 0, normA = 0, normB = 0;
for (int i = 0; i < a.length; i++) {
dot += a[i] * b[i];
normA += a[i] * a[i];
normB += b[i] * b[i];
}
return dot / (Math.sqrt(normA) * Math.sqrt(normB));
}
public static void main(String[] args) {
double[] query = {0.9, 0.1, 0.2};
double[] docA = {0.8, 0.2, 0.1};
double[] docB = {-0.5, 0.9, -0.4};
System.out.printf("query vs A: %.4f%n", cosine(query, docA));
System.out.printf("query vs B: %.4f%n", cosine(query, docB));
}
}The VectorStore Abstraction
A VectorStore persists Document objects together with their embeddings and supports similarity queries. Spring AI ships implementations for PGVector, Redis, Chroma, Milvus, Qdrant, and a SimpleVectorStore for tests.
add(List<Document>)embeds and stores documents.similaritySearch(SearchRequest)embeds the query and returns the nearest documents.delete(...)removes documents by id or filter.
Critically, add calls the EmbeddingModel for you, so you rarely embed manually when using a store.
Configuring PGVector
For production, PGVector (the Postgres extension) is a common choice because it co-locates vectors with your relational data. Add spring-ai-starter-vector-store-pgvector and configure the schema initialization.
index-type: HNSWgives fast approximate nearest-neighbor search.dimensionsmust equal your embedding model's output size.initialize-schema: truelets Spring create thevector_storetable on startup.
spring:
ai:
vectorstore:
pgvector:
initialize-schema: true
index-type: HNSW
distance-type: COSINE_DISTANCE
dimensions: 1536
datasource:
url: jdbc:postgresql://localhost:5432/appdb
username: app
password: ${DB_PASSWORD}Ingesting Documents
Wrap each chunk of source text in a Document. You can attach arbitrary metadata (source, author, tenant id) which becomes filterable at query time. Calling vectorStore.add(...) embeds and persists in one step.
- Use a stable id when you want to upsert rather than duplicate.
- Keep chunks small (a few hundred tokens) so retrieved context is focused.
@Service
class IngestionService {
private final VectorStore vectorStore;
IngestionService(VectorStore vectorStore) {
this.vectorStore = vectorStore;
}
void ingest() {
List<Document> docs = List.of(
new Document("Spring Boot 4 requires Java 17 or later.",
Map.of("source", "release-notes", "version", "4.0")),
new Document("Reset your password from the account recovery page.",
Map.of("source", "help-center", "topic", "auth"))
);
vectorStore.add(docs);
}
}Chunking With ETL Readers
Real corpora arrive as PDFs, Markdown, or JSON. Spring AI's ETL pipeline provides DocumentReader sources and DocumentTransformer splitters. The TokenTextSplitter breaks large documents into embedding-sized chunks while preserving metadata.
TikaDocumentReaderreads PDFs, Word, HTML.TokenTextSplittersplits by token count with configurable overlap.- The output feeds straight into
vectorStore.add(...).
@Service
class PdfIngestion {
private final VectorStore vectorStore;
PdfIngestion(VectorStore vectorStore) {
this.vectorStore = vectorStore;
}
void load(Resource pdf) {
var reader = new TikaDocumentReader(pdf);
var splitter = new TokenTextSplitter();
List<Document> chunks = splitter.apply(reader.get());
vectorStore.add(chunks);
}
}Similarity Search With SearchRequest
SearchRequest is a builder that controls retrieval. The key knobs are topK (how many results) and similarityThreshold (drop weak matches below a cutoff).
- A higher
similarityThreshold(for example 0.75) reduces noise but may return fewer hits. topKcaps how much context you feed to the LLM, controlling token cost.
@Service
class SearchService {
private final VectorStore vectorStore;
SearchService(VectorStore vectorStore) {
this.vectorStore = vectorStore;
}
List<Document> search(String query) {
return vectorStore.similaritySearch(
SearchRequest.builder()
.query(query)
.topK(4)
.similarityThreshold(0.7)
.build()
);
}
}Metadata Filtering
Semantic relevance alone is not enough for multi-tenant or scoped data. SearchRequest accepts a filter expression evaluated against document metadata, combining vector similarity with structured constraints.
- Use
FilterExpressionBuilderor a portable string expression. - The store applies the filter and the similarity ranking together.
- This is how you enforce tenant isolation:
tenantId == 'acme'.
List<Document> scoped = vectorStore.similaritySearch(
SearchRequest.builder()
.query("how do I reset my password")
.topK(3)
.similarityThreshold(0.7)
.filterExpression("source == 'help-center' && topic == 'auth'")
.build()
);From Retrieval to RAG
Retrieval is the first half of RAG (Retrieval-Augmented Generation). The second half stuffs retrieved chunks into the prompt so the LLM answers grounded in your data. Spring AI's QuestionAnswerAdvisor wires the vector store directly into a ChatClient call.
- The advisor runs a similarity search per request and injects results as context.
- You pass the same
SearchRequesttuning (topK, threshold) into the advisor. - This keeps the answer faithful to your indexed documents and reduces hallucination.
@Service
class RagService {
private final ChatClient chatClient;
RagService(ChatClient.Builder builder, VectorStore vectorStore) {
this.chatClient = builder
.defaultAdvisors(QuestionAnswerAdvisor.builder(vectorStore)
.searchRequest(SearchRequest.builder().topK(4).similarityThreshold(0.7).build())
.build())
.build();
}
String ask(String question) {
return chatClient.prompt().user(question).call().content();
}
}Quick Check: Tuning Retrieval
A teammate reports that their RAG endpoint frequently injects irrelevant documents into the prompt, bloating token cost and confusing the model. They want to keep only strongly relevant matches without changing the embedding model. Which single SearchRequest adjustment most directly addresses this?
Recap and Key Takeaways
You built a complete semantic retrieval pipeline in Spring AI:
- EmbeddingModel turns text into dense vectors;
dimensions()must match your store's column size. - VectorStore (PGVector, Redis, Qdrant, etc.) stores
Documents and embeds them automatically onadd(). - Use the ETL pipeline (
TikaDocumentReader+TokenTextSplitter) to chunk real-world files before ingestion. - SearchRequest controls retrieval via
topK,similarityThreshold, and metadatafilterExpressionfor scoping and tenant isolation. - QuestionAnswerAdvisor turns retrieval into full RAG by injecting matched context into a
ChatClientprompt.
Tune topK for cost and similarityThreshold for precision to balance grounding against token budget.
常见问题解答
「嵌入与向量存储检索」课时是免费的吗?
是的 — 「嵌入与向量存储检索」的完整文本可在网页上免费阅读。要进行交互式练习(内置代码编辑器和全天候 AI 导师)并解锁 Spring Boot 4 Complete Guide 课程的其余内容,请升级到 CoddyKit PRO。 Spring Boot 4 Complete Guide 课程共包含 4 节课。
「嵌入与向量存储检索」这节课中我会学到什么?
生成嵌入并查询向量存储,为数据提供语义搜索能力。 你通过在浏览器中直接运行的动手代码来练习 Spring Boot 4 Complete Guide,全天候 AI 导师会在你学习这节课的过程中回答你的问题。
学习 Spring Boot 4 Complete Guide 需要有经验吗?
无需任何先前经验。CoddyKit 上的 Spring Boot 4 Complete Guide 课程适合初学者到高级学习者,你可以从这里开始或从头开始,按照自己的节奏学习。 这是第 2 节课,共 4 节。
「嵌入与向量存储检索」课时需要多长时间?
大多数 CoddyKit 课程大约需要 5–10 分钟。每节课都很精短且互动,所以你能稳步进步,并在网页和应用中从离开的地方继续。
我能在这节 Spring Boot 4 Complete Guide 课中编写并运行代码吗?
能。每节 Spring Boot 4 Complete Guide 课都包含内置代码编辑器,你可以在浏览器中直接编写并运行真实代码,并获得即时 AI 反馈 — 无需本地设置。