0Pricing
Vector Databases: Pinecone, Weaviate & pgvector · 课时

上下文信息检索

实现从向量存储中检索最相关上下文的策略,用于增强 LLM 提示词。

上下文信息检索 是 CoddyKit 上的免费 Vector Databases: Pinecone, Weaviate & pgvector 课时。 这是第 3 节课,共 4 节。 你可以在下方免费阅读本课时的完整内容 — 然后在浏览器中使用内置代码编辑器和全天候 AI 导师进行实践。 这是 Vector Databases: Pinecone, Weaviate & pgvector 学习路径的一部分,你的进度在网页和 CoddyKit 应用中同步。 Vector Databases: Pinecone, Weaviate & pgvector 课程共包含 4 节课。

本课时的部分内容尚未翻译,以英文显示。

What is Context Retrieval?

In a Retrieval-Augmented Generation (RAG) system, the Large Language Model (LLM) needs relevant information to generate accurate responses.

  • Contextual Information Retrieval is the process of finding and fetching this relevant data from your vector database.
  • It's the bridge that connects the user's query to the knowledge stored in your specialized data.

The Retrieval Workflow

When a user asks a question, several steps happen to get the right context:

  1. The user's question (query) is converted into a vector embedding.
  2. This query embedding is sent to the vector database.
  3. The vector database searches for stored document embeddings that are most similar to the query embedding.
  4. The text chunks associated with these similar embeddings are retrieved and sent to the LLM.

Vector Similarity Basics

The core of retrieval is vector similarity. Your vector database calculates how 'close' your query vector is to all the stored document vectors.

  • Closer vectors mean higher semantic similarity.
  • Common similarity metrics include cosine similarity or Euclidean distance.
  • The database efficiently finds the closest vectors, usually using specialized indexing techniques.

Simple Top-K Retrieval

The most straightforward retrieval strategy is Top-K Retrieval.

  • You simply ask the vector database to return the K most similar document chunks to your query.
  • K is a number you choose (e.g., 3, 5, or 10), representing how many pieces of context you want to provide to the LLM.
  • While simple, choosing the right K is crucial for balancing relevance and LLM token limits.

Python Top-K Retrieval Demo

This simple Python code simulates a vector store and demonstrates how top_k retrieval works. Try changing the top_k value!

import math

class SimpleVectorStore:
    def __init__(self):
        self.vectors = {}

    def add_document(self, doc_id, vector, text):
        self.vectors[doc_id] = {"vector": vector, "text": text}

    def _cosine_similarity(self, vec1, vec2):
        dot_product = sum(v1 * v2 for v1, v2 in zip(vec1, vec2))
        magnitude1 = math.sqrt(sum(v**2 for v in vec1))
        magnitude2 = math.sqrt(sum(v**2 for v in vec2))
        if magnitude1 == 0 or magnitude2 == 0:
            return 0.0
        return dot_product / (magnitude1 * magnitude2)

    def query(self, query_vector, top_k=3):
        similarities = []
        for doc_id, data in self.vectors.items():
            sim = self._cosine_similarity(query_vector, data["vector"])
            similarities.append((sim, doc_id, data["text"]))

        similarities.sort(key=lambda x: x[0], reverse=True)
        return [{"id": s[1], "text": s[2], "similarity": s[0]} for s in similarities[:top_k]]

if __name__ == "__main__":
    store = SimpleVectorStore()

    store.add_document("doc1", [0.1, 0.2, 0.3], "The quick brown fox jumps over the lazy dog.")
    store.add_document("doc2", [0.15, 0.25, 0.35], "A fast fox leaps over a sleepy canine.")
    store.add_document("doc3", [0.8, 0.7, 0.9], "Artificial intelligence is transforming industries.")
    store.add_document("doc4", [0.75, 0.85, 0.95], "Machine learning algorithms are key to AI.")

    query_vec = [0.12, 0.22, 0.32] # Simulating an embedding for "fast animal"

    print("--- Top 2 Relevant Chunks ---")
    results = store.query(query_vec, top_k=2)
    for res in results:
        print(f"ID: {res['id']}, Sim: {res['similarity']:.2f}, Text: {res['text']}")

    print("\n--- Top 1 Relevant Chunk ---")
    results_single = store.query(query_vec, top_k=1)
    for res in results_single:
        print(f"ID: {res['id']}, Sim: {res['similarity']:.2f}, Text: {res['text']}")

Chunking for Effective Retrieval

The quality of your retrieval heavily depends on how your original documents were broken down into chunks before being embedded.

  • Chunk Size: Too small, and context might be lost. Too large, and irrelevant information might be included.
  • Overlap: Adding overlap between chunks helps ensure that important information isn't split across boundaries.
  • Good chunking ensures each retrieved piece of context is meaningful and self-contained.

Enhancing Retrieval with Re-ranking

Sometimes, simple Top-K retrieval isn't enough. The most 'similar' vectors aren't always the most 'relevant' in context.

  • Re-ranking is an optional but powerful step performed after initial retrieval.
  • It takes the top K chunks from the vector database and uses a smaller, more specialized model to score their relevance more deeply.
  • This secondary scoring helps filter out less useful chunks and prioritize truly pertinent information.

The Need for Re-ranking

Why do we need re-ranking?

  • Vector similarity can sometimes be fooled by superficial semantic closeness.
  • A re-ranker, often a smaller transformer model, can better understand the nuanced relationship between the query and the retrieved document chunks.
  • It helps ensure the context provided to the LLM is not only similar but also highly relevant and useful for answering the user's specific question.

Advanced Retrieval Concepts

Beyond basic Top-K and re-ranking, advanced strategies can further improve retrieval:

  • Query Expansion: Rewriting or adding terms to the user's original query to improve search results.
  • Hybrid Search: Combining traditional keyword search (like full-text search) with vector similarity for a more comprehensive retrieval.
  • These methods aim to make the initial retrieval even more robust before context is sent to the LLM.

Check Your Knowledge

Consider a RAG system that initially retrieves 10 document chunks using vector similarity. What are key benefits of adding a re-ranking step?

Contextual Retrieval Summary

You've learned how to retrieve relevant context for RAG systems!

  • Contextual retrieval bridges user queries and stored knowledge.
  • It involves converting queries to embeddings, querying the vector DB for similar vectors, and fetching associated text.
  • Top-K retrieval is the basic method, but good chunking is vital.
  • Re-ranking can further refine results by applying a secondary relevance filter.
  • Advanced techniques like query expansion and hybrid search offer even more sophisticated retrieval.

常见问题解答

「上下文信息检索」课时是免费的吗?

是的 — 「上下文信息检索」的完整文本可在网页上免费阅读。要进行交互式练习(内置代码编辑器和全天候 AI 导师)并解锁 Vector Databases: Pinecone, Weaviate & pgvector 课程的其余内容,请升级到 CoddyKit PRO。 Vector Databases: Pinecone, Weaviate & pgvector 课程共包含 4 节课。

「上下文信息检索」这节课中我会学到什么?

实现从向量存储中检索最相关上下文的策略,用于增强 LLM 提示词。 你通过在浏览器中直接运行的动手代码来练习 Vector Databases: Pinecone, Weaviate & pgvector,全天候 AI 导师会在你学习这节课的过程中回答你的问题。

学习 Vector Databases: Pinecone, Weaviate & pgvector 需要有经验吗?

无需任何先前经验。CoddyKit 上的 Vector Databases: Pinecone, Weaviate & pgvector 课程适合初学者到高级学习者,你可以从这里开始或从头开始,按照自己的节奏学习。 这是第 3 节课,共 4 节。

「上下文信息检索」课时需要多长时间?

大多数 CoddyKit 课程大约需要 5–10 分钟。每节课都很精短且互动,所以你能稳步进步,并在网页和应用中从离开的地方继续。

我能在这节 Vector Databases: Pinecone, Weaviate & pgvector 课中编写并运行代码吗?

能。每节 Vector Databases: Pinecone, Weaviate & pgvector 课都包含内置代码编辑器,你可以在浏览器中直接编写并运行真实代码,并获得即时 AI 反馈 — 无需本地设置。

此课程中的所有课时

  1. RAG 系统架构概览
  2. 与 LLM 框架集成
  3. 上下文信息检索
  4. RAG 的文本块拆分策略
← 返回 Vector Databases: Pinecone, Weaviate & pgvector