Upsert danych do Pinecone
Opanuje Pan/Pani proces wstawiania i aktualizowania danych wektorowych wraz z powiązanymi metadanymi w indeksie Pinecone.
Upsert danych do Pinecone to bezpłatna lekcja Vector Databases: Pinecone, Weaviate & pgvector na CoddyKit. To lekcja 2 z 4. Możesz przeczytać całą lekcję poniżej za darmo — a potem ćwiczyć ją interaktywnie w przeglądarce z wbudowanym edytorem kodu i tutorem AI dostępnym 24/7. To część ścieżki edukacyjnej Vector Databases: Pinecone, Weaviate & pgvector, a Twój postęp synchronizuje się między webem a aplikacją CoddyKit. Kurs Vector Databases: Pinecone, Weaviate & pgvector zawiera 4 lekcji w sumie.
Części tej lekcji nie zostały jeszcze przetłumaczone i są wyświetlane po angielsku.
What is Upserting Data?
In vector databases like Pinecone, upserting is a key operation. It's a blend of 'update' and 'insert'.
- If a vector with a given ID doesn't exist, it's inserted.
- If it already exists, the existing vector (and its metadata) is updated with the new data.
This single operation simplifies managing dynamic vector data without needing to check for existence first.
Connecting to Your Pinecone Index
Before we upsert, you need to connect to your Pinecone index. This involves initializing the Pinecone client and selecting your target index.
Remember to replace YOUR_API_KEY and YOUR_ENVIRONMENT with your actual credentials, typically loaded from environment variables.
import os
from pinecone import Pinecone
# Initialize Pinecone client
pc = Pinecone(
api_key=os.environ.get("PINECONE_API_KEY", "YOUR_API_KEY"),
environment=os.environ.get("PINECONE_ENVIRONMENT", "YOUR_ENVIRONMENT")
)
# Connect to your index (e.g., 'my-index')
index_name = "my-index"
index = pc.Index(index_name)
print(f"Connected to index: {index_name}")Upserting a Single Vector
To upsert a single vector, you provide a unique id and the vector data (a list of floats). Each vector needs an ID for Pinecone to identify it.
The dimension of the vector must match the dimension configured for your Pinecone index (e.g., 3 for our simple examples).
Demo: Single Vector Upsert
Here's how to insert a single vector into your index. We'll use a simple 3-dimensional vector for demonstration.
import os
from pinecone import Pinecone
pc = Pinecone(
api_key=os.environ.get("PINECONE_API_KEY", "YOUR_API_KEY"),
environment=os.environ.get("PINECONE_ENVIRONMENT", "YOUR_ENVIRONMENT")
)
index = pc.Index("my-index") # Assuming 'my-index' exists
# Upsert a single vector
index.upsert(
vectors=[
{"id": "vec1", "values": [0.1, 0.2, 0.3]}
]
)
print("Single vector 'vec1' upserted!")Efficient Batch Upserts
For better performance, it's highly recommended to upsert multiple vectors in batches rather than one by one. This reduces network overhead.
You can prepare a list of dictionaries, where each dictionary represents a vector with its id and values.
Demo: Batch Upserting Vectors
This example shows how to upsert several vectors at once. Notice the vectors parameter takes a list of vector objects.
import os
from pinecone import Pinecone
pc = Pinecone(
api_key=os.environ.get("PINECONE_API_KEY", "YOUR_API_KEY"),
environment=os.environ.get("PINECONE_ENVIRONMENT", "YOUR_ENVIRONMENT")
)
index = pc.Index("my-index") # Assuming 'my-index' exists
# Prepare multiple vectors for batch upsert
batch_vectors = [
{"id": "vec2", "values": [0.4, 0.5, 0.6]},
{"id": "vec3", "values": [0.7, 0.8, 0.9]},
{"id": "vec4", "values": [0.11, 0.12, 0.13]}
]
# Upsert the batch
index.upsert(vectors=batch_vectors)
print("Batch of vectors upserted!")Adding Metadata to Vectors
Metadata allows you to store additional key-value pairs alongside your vectors. This is incredibly useful for filtering search results later.
Metadata can include properties like a document's title, author, category, or creation date. It's stored as a dictionary within each vector object.
Demo: Upsert with Metadata
Let's add some context to our vectors using metadata. This makes them much more powerful for real-world applications.
import os
from pinecone import Pinecone
pc = Pinecone(
api_key=os.environ.get("PINECONE_API_KEY", "YOUR_API_KEY"),
environment=os.environ.get("PINECONE_ENVIRONMENT", "YOUR_ENVIRONMENT")
)
index = pc.Index("my-index") # Assuming 'my-index' exists
# Upsert a vector with metadata
index.upsert(
vectors=[
{
"id": "doc1",
"values": [0.2, 0.3, 0.4],
"metadata": {"genre": "sci-fi", "year": 2023}
},
{
"id": "doc2",
"values": [0.5, 0.6, 0.7],
"metadata": {"genre": "fantasy", "author": "J. Doe"}
}
]
)
print("Vectors 'doc1' and 'doc2' upserted with metadata!")How Upsert Handles Updates
The 'update' part of 'upsert' means that if you try to upsert a vector with an id that already exists in your index, Pinecone won't create a new entry.
Instead, it will overwrite the existing vector's values and metadata with the new data you provide. This ensures data consistency and avoids duplicates.
Upserting Knowledge Check
You've learned about upserting. Let's check your understanding!
Recap: Mastering Upserts
Great job! You've learned the essentials of upserting data into Pinecone:
- Upsert = Update + Insert: A single operation for adding new or modifying existing vectors.
- IDs are Key: Each vector requires a unique ID for identification and updates.
- Batching for Performance: Always upsert multiple vectors in batches for efficiency.
- Metadata for Context: Add key-value pairs to vectors for powerful filtering and search.
- Updates are Automatic: Upserting with an existing ID overwrites the old vector.
Next, we'll explore how to query these vectors to find similar items!
Ucz się Vector Databases: Pinecone, Weaviate & pgvector dzięki korepetycjom AI — za darmo
Pisz i uruchamiaj kod w przeglądarce, otrzymuj natychmiastową pomoc od korepetytora AI dostępnego 24/7 i kontynuuj naukę w sieci lub w aplikacji.
- Kursy
- 12
- Lekcje
- 48
Często zadawane pytania
Czy lekcja „Upsert danych do Pinecone” jest bezpłatna?
Tak — pełny tekst „Upsert danych do Pinecone” jest dostępny za darmo tutaj w sieci. Aby ćwiczyć ją interaktywnie (wbudowany edytor kodu i tutor AI dostępny 24/7) i odblokować resztę kursu Vector Databases: Pinecone, Weaviate & pgvector, przejdź na CoddyKit PRO. Kurs Vector Databases: Pinecone, Weaviate & pgvector zawiera 4 lekcji w sumie.
Co nauczysz się w „Upsert danych do Pinecone”?
Opanuje Pan/Pani proces wstawiania i aktualizowania danych wektorowych wraz z powiązanymi metadanymi w indeksie Pinecone. Ćwiczysz Vector Databases: Pinecone, Weaviate & pgvector z praktycznym kodem, który uruchamiasz bezpośrednio w przeglądarce, a tutor AI dostępny 24/7 odpowiada na Twoje pytania podczas pracy nad lekcją.
Czy potrzebuję doświadczenia, aby zacząć Vector Databases: Pinecone, Weaviate & pgvector?
Nie wymagamy żadnego doświadczenia. Vector Databases: Pinecone, Weaviate & pgvector w CoddyKit jest strukturyzowany dla początkujących i zaawansowanych użytkowników, więc możesz zacząć tutaj lub od początku i uczyć się w swoim tempie. To lekcja 2 z 4.
Ile czasu zajmuje lekcja „Upsert danych do Pinecone”?
Większość lekcji CoddyKit trwa około 5–10 minut. Każda lekcja to mały, interaktywny krok, dzięki czemu robisz systematyczne postępy i zawsze wracasz dokładnie do tego samego miejsca — na webie i w aplikacji.
Czy mogę pisać i uruchamiać kod w tej lekcji Vector Databases: Pinecone, Weaviate & pgvector?
Tak. Każda lekcja Vector Databases: Pinecone, Weaviate & pgvector zawiera wbudowany edytor kodu, więc piszesz i uruchamiasz prawdziwy kod bezpośrednio w przeglądarce i od razu otrzymujesz sprzężenie zwrotne od AI — bez konfiguracji na komputerze.
Wszystkie lekcje w tym kursie
- Tworzenie indeksu Pinecone
- Upsert danych do Pinecone
- Wykonywanie zapytań o dane wektorowe w Pinecone
- Zrozumienie cen i podów Pinecone