0Pricing
Vector Databases: Pinecone, Weaviate & pgvector · 课时

存储与更新嵌入

了解在数据流水线中存储、建立索引并高效更新嵌入的最佳实践。

存储与更新嵌入 是 CoddyKit 上的免费 Vector Databases: Pinecone, Weaviate & pgvector 课时。 这是第 3 节课,共 4 节。 你可以在下方免费阅读本课时的完整内容 — 然后在浏览器中使用内置代码编辑器和全天候 AI 导师进行实践。 这是 Vector Databases: Pinecone, Weaviate & pgvector 学习路径的一部分,你的进度在网页和 CoddyKit 应用中同步。 Vector Databases: Pinecone, Weaviate & pgvector 课程共包含 4 节课。

本课时的部分内容尚未翻译,以英文显示。

Storing & Updating Embeddings

Welcome to Lesson 3! This lesson covers the essential practices for managing embeddings: how to store them effectively, the role of indexing for search performance, and strategies for updating embeddings to keep your data fresh.

These concepts are crucial for building dynamic and responsive AI applications.

Why Persist Embeddings?

Generating embeddings can be computationally intensive and time-consuming. Storing them after creation offers significant benefits:

  • Reuse: Avoid re-computing the same embedding for multiple queries.
  • Speed: Enable much faster similarity searches.
  • Scale: Support larger applications without constant re-generation.

Think of it as caching the 'meaning' of your data for quick access.

Choosing an Embedding Store

Where should you keep your embeddings? While simple options exist, specialized solutions are usually best:

  • Vector Databases: Designed specifically for storing and searching vector embeddings (e.g., Pinecone, Weaviate). They offer optimal performance.
  • Relational DBs (with extensions): Traditional databases like PostgreSQL can store vectors using extensions like pgvector.
  • File Systems: Simple for very small or static datasets, but not practical for scalable search.

For most AI applications, a vector database is the preferred choice.

Structure of Stored Data

When you store an embedding, you typically store more than just the raw vector. A complete entry usually includes:

  • Vector: The numerical array representing your data (e.g., [0.1, 0.2, ...]).
  • ID: A unique identifier that links the embedding back to its original data source (e.g., a document ID, image hash).
  • Metadata: Additional descriptive information about the original data (e.g., title, author, category). This is invaluable for filtering and enriching search results.

Simulating Embedding Storage

Here's a simple Python example demonstrating how you might conceptually store an embedding with its ID and metadata. In a real-world scenario, a vector database would handle this more robustly.

def main():
    # Simulate an embedding store (a list of dictionaries)
    embedding_store = []

    # Example data for a document
    doc_id = "doc_abc_123"
    doc_embedding = [0.1, 0.2, 0.3, 0.4, 0.5] # Simplified vector
    doc_metadata = {"title": "Intro to Vectors", "author": "Alice"}

    # Create an entry for the embedding
    embedding_entry = {
        "id": doc_id,
        "vector": doc_embedding,
        "metadata": doc_metadata
    }

    # Add the entry to our simulated store
    embedding_store.append(embedding_entry)

    print(f"Stored entry for ID: {embedding_entry['id']}")
    print(f"Vector: {embedding_entry['vector']}")
    print(f"Metadata: {embedding_entry['metadata']}")

if __name__ == "__main__":
    main()

Introduction to Indexing

Once embeddings are stored, they need to be organized in a way that allows for fast similarity searches. This organization process is called indexing.

Unlike traditional database indexes for exact matches, vector indexes are designed to speed up Approximate Nearest Neighbor (ANN) searches. This means finding vectors that are 'close enough' to a query vector very quickly, even if it's not the absolute closest every single time.

Indexing Trade-offs

When designing or choosing an indexing strategy for embeddings, there are important trade-offs:

  • Speed: How quickly can similarity queries be processed?
  • Accuracy (Recall): How many of the true nearest neighbors are actually found by the index?
  • Memory Usage: How much memory or disk space does the index itself consume?

Often, a slight reduction in accuracy is accepted to gain significant improvements in search speed and memory efficiency, especially with very large datasets.

The Challenge of Updates

Data in real-world applications is rarely static. Documents are edited, images are replaced, and user profiles are updated. When the original data changes, its corresponding embedding also needs to be updated to reflect the new content.

Failing to update embeddings can lead to outdated or inaccurate search results, making your AI application less effective.

Strategies for Updating Embeddings

There are two primary approaches to handling embedding updates:

  • Full Re-indexing (Batch Updates): Regenerate all embeddings from scratch and completely rebuild the entire vector index. This is simple but very resource-intensive for large datasets.
  • Partial Updates (Upserts): Modify or insert specific vectors without rebuilding the whole index. Most vector databases support this operation, often called 'upsert' (update if exists, otherwise insert). This is much more efficient for dynamic data.

Simulating an Embedding Update

Here's how you might conceptually update an embedding in our simulated store. A real vector database would provide an optimized 'upsert' command to handle this efficiently.

def main():
    # Simulate an embedding store with an existing entry
    embedding_store = [
        {
            "id": "doc_abc_123",
            "vector": [0.1, 0.2, 0.3, 0.4, 0.5],
            "metadata": {"title": "Intro to Vectors", "author": "Alice"}
        }
    ]

    # New embedding data for an existing ID
    updated_doc_id = "doc_abc_123"
    new_embedding_vector = [0.6, 0.7, 0.8, 0.9, 1.0] # The new vector
    new_metadata = {"title": "Intro to Vectors (Revised)", "author": "Alice"}

    # Find and update the entry in our store
    found = False
    for entry in embedding_store:
        if entry["id"] == updated_doc_id:
            entry["vector"] = new_embedding_vector
            entry["metadata"] = new_metadata
            found = True
            break

    if found:
        print(f"Updated entry for ID: {updated_doc_id}")
        print(f"New vector: {embedding_store[0]['vector']}")
        print(f"New metadata: {embedding_store[0]['metadata']}")
    else:
        print(f"ID {updated_doc_id} not found for update.")

if __name__ == "__main__":
    main()

Understanding Updates

Imagine you have a document stored in your vector database. The document's content is updated, meaning its embedding needs to change. Which term describes the most efficient way to replace an existing embedding with a new one in most modern vector databases?

Recap: Storing & Updating

In this lesson, we covered the essential aspects of managing embeddings:

  • Storing: Persisting embeddings with IDs and metadata for reuse and speed.
  • Indexing: Organizing embeddings for efficient Approximate Nearest Neighbor (ANN) search.
  • Updating: Strategies like 'upsert' to keep embeddings fresh when source data changes.

Mastering these practices is key to building robust and performant vector-search applications.

常见问题解答

「存储与更新嵌入」课时是免费的吗?

是的 — 「存储与更新嵌入」的完整文本可在网页上免费阅读。要进行交互式练习(内置代码编辑器和全天候 AI 导师)并解锁 Vector Databases: Pinecone, Weaviate & pgvector 课程的其余内容,请升级到 CoddyKit PRO。 Vector Databases: Pinecone, Weaviate & pgvector 课程共包含 4 节课。

「存储与更新嵌入」这节课中我会学到什么?

了解在数据流水线中存储、建立索引并高效更新嵌入的最佳实践。 你通过在浏览器中直接运行的动手代码来练习 Vector Databases: Pinecone, Weaviate & pgvector,全天候 AI 导师会在你学习这节课的过程中回答你的问题。

学习 Vector Databases: Pinecone, Weaviate & pgvector 需要有经验吗?

无需任何先前经验。CoddyKit 上的 Vector Databases: Pinecone, Weaviate & pgvector 课程适合初学者到高级学习者,你可以从这里开始或从头开始,按照自己的节奏学习。 这是第 3 节课,共 4 节。

「存储与更新嵌入」课时需要多长时间?

大多数 CoddyKit 课程大约需要 5–10 分钟。每节课都很精短且互动,所以你能稳步进步,并在网页和应用中从离开的地方继续。

我能在这节 Vector Databases: Pinecone, Weaviate & pgvector 课中编写并运行代码吗?

能。每节 Vector Databases: Pinecone, Weaviate & pgvector 课都包含内置代码编辑器,你可以在浏览器中直接编写并运行真实代码,并获得即时 AI 反馈 — 无需本地设置。

此课程中的所有课时

  1. 文本嵌入模型
  2. 使用嵌入应用程序接口
  3. 存储与更新嵌入
  4. 为更好的嵌入拆分文本
← 返回 Vector Databases: Pinecone, Weaviate & pgvector