0Pricing
Vector Databases: Pinecone, Weaviate & pgvector · 강의

컨텍스트 기반 정보 검색

LLM 프롬프트를 보강할 수 있도록 벡터 저장소에서 가장 관련성 높은 컨텍스트를 검색하는 전략을 구현합니다.

컨텍스트 기반 정보 검색은(는) CoddyKit의 무료 Vector Databases: Pinecone, Weaviate & pgvector 강의입니다. 이것은 4개 중 3번째 강의입니다. 아래에서 전체 강의를 무료로 읽을 수 있으며, 내장 코드 에디터와 24/7 AI 튜터와 함께 브라우저에서 직접 실습할 수 있습니다. 이 강의는 Vector Databases: Pinecone, Weaviate & pgvector 학습 경로의 일부이며, 진행 상황이 웹과 CoddyKit 앱에 동기화됩니다. Vector Databases: Pinecone, Weaviate & pgvector 강의에는 총 4개의 강의가 포함되어 있습니다.

이 강의의 일부는 아직 번역되지 않았으며 영어로 표시됩니다.

What is Context Retrieval?

In a Retrieval-Augmented Generation (RAG) system, the Large Language Model (LLM) needs relevant information to generate accurate responses.

  • Contextual Information Retrieval is the process of finding and fetching this relevant data from your vector database.
  • It's the bridge that connects the user's query to the knowledge stored in your specialized data.

The Retrieval Workflow

When a user asks a question, several steps happen to get the right context:

  1. The user's question (query) is converted into a vector embedding.
  2. This query embedding is sent to the vector database.
  3. The vector database searches for stored document embeddings that are most similar to the query embedding.
  4. The text chunks associated with these similar embeddings are retrieved and sent to the LLM.

Vector Similarity Basics

The core of retrieval is vector similarity. Your vector database calculates how 'close' your query vector is to all the stored document vectors.

  • Closer vectors mean higher semantic similarity.
  • Common similarity metrics include cosine similarity or Euclidean distance.
  • The database efficiently finds the closest vectors, usually using specialized indexing techniques.

Simple Top-K Retrieval

The most straightforward retrieval strategy is Top-K Retrieval.

  • You simply ask the vector database to return the K most similar document chunks to your query.
  • K is a number you choose (e.g., 3, 5, or 10), representing how many pieces of context you want to provide to the LLM.
  • While simple, choosing the right K is crucial for balancing relevance and LLM token limits.

Python Top-K Retrieval Demo

This simple Python code simulates a vector store and demonstrates how top_k retrieval works. Try changing the top_k value!

import math

class SimpleVectorStore:
    def __init__(self):
        self.vectors = {}

    def add_document(self, doc_id, vector, text):
        self.vectors[doc_id] = {"vector": vector, "text": text}

    def _cosine_similarity(self, vec1, vec2):
        dot_product = sum(v1 * v2 for v1, v2 in zip(vec1, vec2))
        magnitude1 = math.sqrt(sum(v**2 for v in vec1))
        magnitude2 = math.sqrt(sum(v**2 for v in vec2))
        if magnitude1 == 0 or magnitude2 == 0:
            return 0.0
        return dot_product / (magnitude1 * magnitude2)

    def query(self, query_vector, top_k=3):
        similarities = []
        for doc_id, data in self.vectors.items():
            sim = self._cosine_similarity(query_vector, data["vector"])
            similarities.append((sim, doc_id, data["text"]))

        similarities.sort(key=lambda x: x[0], reverse=True)
        return [{"id": s[1], "text": s[2], "similarity": s[0]} for s in similarities[:top_k]]

if __name__ == "__main__":
    store = SimpleVectorStore()

    store.add_document("doc1", [0.1, 0.2, 0.3], "The quick brown fox jumps over the lazy dog.")
    store.add_document("doc2", [0.15, 0.25, 0.35], "A fast fox leaps over a sleepy canine.")
    store.add_document("doc3", [0.8, 0.7, 0.9], "Artificial intelligence is transforming industries.")
    store.add_document("doc4", [0.75, 0.85, 0.95], "Machine learning algorithms are key to AI.")

    query_vec = [0.12, 0.22, 0.32] # Simulating an embedding for "fast animal"

    print("--- Top 2 Relevant Chunks ---")
    results = store.query(query_vec, top_k=2)
    for res in results:
        print(f"ID: {res['id']}, Sim: {res['similarity']:.2f}, Text: {res['text']}")

    print("\n--- Top 1 Relevant Chunk ---")
    results_single = store.query(query_vec, top_k=1)
    for res in results_single:
        print(f"ID: {res['id']}, Sim: {res['similarity']:.2f}, Text: {res['text']}")

Chunking for Effective Retrieval

The quality of your retrieval heavily depends on how your original documents were broken down into chunks before being embedded.

  • Chunk Size: Too small, and context might be lost. Too large, and irrelevant information might be included.
  • Overlap: Adding overlap between chunks helps ensure that important information isn't split across boundaries.
  • Good chunking ensures each retrieved piece of context is meaningful and self-contained.

Enhancing Retrieval with Re-ranking

Sometimes, simple Top-K retrieval isn't enough. The most 'similar' vectors aren't always the most 'relevant' in context.

  • Re-ranking is an optional but powerful step performed after initial retrieval.
  • It takes the top K chunks from the vector database and uses a smaller, more specialized model to score their relevance more deeply.
  • This secondary scoring helps filter out less useful chunks and prioritize truly pertinent information.

The Need for Re-ranking

Why do we need re-ranking?

  • Vector similarity can sometimes be fooled by superficial semantic closeness.
  • A re-ranker, often a smaller transformer model, can better understand the nuanced relationship between the query and the retrieved document chunks.
  • It helps ensure the context provided to the LLM is not only similar but also highly relevant and useful for answering the user's specific question.

Advanced Retrieval Concepts

Beyond basic Top-K and re-ranking, advanced strategies can further improve retrieval:

  • Query Expansion: Rewriting or adding terms to the user's original query to improve search results.
  • Hybrid Search: Combining traditional keyword search (like full-text search) with vector similarity for a more comprehensive retrieval.
  • These methods aim to make the initial retrieval even more robust before context is sent to the LLM.

Check Your Knowledge

Consider a RAG system that initially retrieves 10 document chunks using vector similarity. What are key benefits of adding a re-ranking step?

Contextual Retrieval Summary

You've learned how to retrieve relevant context for RAG systems!

  • Contextual retrieval bridges user queries and stored knowledge.
  • It involves converting queries to embeddings, querying the vector DB for similar vectors, and fetching associated text.
  • Top-K retrieval is the basic method, but good chunking is vital.
  • Re-ranking can further refine results by applying a secondary relevance filter.
  • Advanced techniques like query expansion and hybrid search offer even more sophisticated retrieval.

자주 묻는 질문

“컨텍스트 기반 정보 검색” 강의는 무료인가요?

네 — “컨텍스트 기반 정보 검색” 전체 내용을 이 웹사이트에서 무료로 읽을 수 있습니다. 인터랙티브하게 실습하려면(내장 코드 에디터와 24/7 AI 튜터), CoddyKit PRO로 업그레이드하면 Vector Databases: Pinecone, Weaviate & pgvector 강의 전체를 잠금 해제할 수 있습니다. Vector Databases: Pinecone, Weaviate & pgvector 강의에는 총 4개의 강의가 포함되어 있습니다.

“컨텍스트 기반 정보 검색”에서 뭘 배우나요?

LLM 프롬프트를 보강할 수 있도록 벡터 저장소에서 가장 관련성 높은 컨텍스트를 검색하는 전략을 구현합니다. 브라우저에서 직접 실행하는 실습 코드로 Vector Databases: Pinecone, Weaviate & pgvector을(를) 배우며, 24/7 AI 튜터가 강의를 진행하면서 질문에 답변해줍니다.

Vector Databases: Pinecone, Weaviate & pgvector을(를) 시작하는 데 경험이 필요한가요?

사전 경험은 필요하지 않습니다. CoddyKit의 Vector Databases: Pinecone, Weaviate & pgvector은(는) 초급자부터 고급 학습자까지를 위해 구성되어 있으므로, 여기서 시작하거나 처음부터 시작할 수 있으며 자신의 속도대로 진행할 수 있습니다. 이것은 4개 중 3번째 강의입니다.

“컨텍스트 기반 정보 검색” 강의는 얼마나 걸리나요?

대부분의 CoddyKit 강의는 약 5~10분이 소요됩니다. 각 강의는 간결하고 인터랙티브하여 꾸준한 진행이 가능하며, 웹과 앱에서 중단한 부분부터 바로 시작할 수 있습니다.

이 Vector Databases: Pinecone, Weaviate & pgvector 강의에서 코드를 작성하고 실행할 수 있나요?

네. 모든 Vector Databases: Pinecone, Weaviate & pgvector 강의에는 내장 코드 에디터가 포함되어 있으므로, 브라우저에서 바로 실제 코드를 작성하고 실행한 후 즉시 AI 피드백을 받을 수 있습니다 — 로컬 설정이 필요 없습니다.

이 강의의 모든 강의

  1. RAG 시스템 아키텍처 개요
  2. LLM 프레임워크와 연동하기
  3. 컨텍스트 기반 정보 검색
  4. RAG를 위한 청킹 전략
← Vector Databases: Pinecone, Weaviate & pgvector(으)로 돌아가기