Memasukkan Data ke Pinecone
Kuasai proses menyisipkan dan memperbarui data vektor beserta metadata terkait ke dalam indeks Pinecone Anda.
Memasukkan Data ke Pinecone adalah pelajaran Vector Databases: Pinecone, Weaviate & pgvector gratis di CoddyKit. Ini adalah pelajaran 2 dari 4. Kamu bisa membaca pelajaran lengkapnya di bawah secara gratis — lalu praktikkan langsung di browser dengan editor kode bawaan dan tutor AI 24/7. Ini adalah bagian dari jalur belajar Vector Databases: Pinecone, Weaviate & pgvector, dan progresmu tersinkronisasi di web dan aplikasi CoddyKit. Kursus Vector Databases: Pinecone, Weaviate & pgvector mencakup 4 pelajaran total.
Bagian dari pelajaran ini belum diterjemahkan dan ditampilkan dalam bahasa Inggris.
What is Upserting Data?
In vector databases like Pinecone, upserting is a key operation. It's a blend of 'update' and 'insert'.
- If a vector with a given ID doesn't exist, it's inserted.
- If it already exists, the existing vector (and its metadata) is updated with the new data.
This single operation simplifies managing dynamic vector data without needing to check for existence first.
Connecting to Your Pinecone Index
Before we upsert, you need to connect to your Pinecone index. This involves initializing the Pinecone client and selecting your target index.
Remember to replace YOUR_API_KEY and YOUR_ENVIRONMENT with your actual credentials, typically loaded from environment variables.
import os
from pinecone import Pinecone
# Initialize Pinecone client
pc = Pinecone(
api_key=os.environ.get("PINECONE_API_KEY", "YOUR_API_KEY"),
environment=os.environ.get("PINECONE_ENVIRONMENT", "YOUR_ENVIRONMENT")
)
# Connect to your index (e.g., 'my-index')
index_name = "my-index"
index = pc.Index(index_name)
print(f"Connected to index: {index_name}")Upserting a Single Vector
To upsert a single vector, you provide a unique id and the vector data (a list of floats). Each vector needs an ID for Pinecone to identify it.
The dimension of the vector must match the dimension configured for your Pinecone index (e.g., 3 for our simple examples).
Demo: Single Vector Upsert
Here's how to insert a single vector into your index. We'll use a simple 3-dimensional vector for demonstration.
import os
from pinecone import Pinecone
pc = Pinecone(
api_key=os.environ.get("PINECONE_API_KEY", "YOUR_API_KEY"),
environment=os.environ.get("PINECONE_ENVIRONMENT", "YOUR_ENVIRONMENT")
)
index = pc.Index("my-index") # Assuming 'my-index' exists
# Upsert a single vector
index.upsert(
vectors=[
{"id": "vec1", "values": [0.1, 0.2, 0.3]}
]
)
print("Single vector 'vec1' upserted!")Efficient Batch Upserts
For better performance, it's highly recommended to upsert multiple vectors in batches rather than one by one. This reduces network overhead.
You can prepare a list of dictionaries, where each dictionary represents a vector with its id and values.
Demo: Batch Upserting Vectors
This example shows how to upsert several vectors at once. Notice the vectors parameter takes a list of vector objects.
import os
from pinecone import Pinecone
pc = Pinecone(
api_key=os.environ.get("PINECONE_API_KEY", "YOUR_API_KEY"),
environment=os.environ.get("PINECONE_ENVIRONMENT", "YOUR_ENVIRONMENT")
)
index = pc.Index("my-index") # Assuming 'my-index' exists
# Prepare multiple vectors for batch upsert
batch_vectors = [
{"id": "vec2", "values": [0.4, 0.5, 0.6]},
{"id": "vec3", "values": [0.7, 0.8, 0.9]},
{"id": "vec4", "values": [0.11, 0.12, 0.13]}
]
# Upsert the batch
index.upsert(vectors=batch_vectors)
print("Batch of vectors upserted!")Adding Metadata to Vectors
Metadata allows you to store additional key-value pairs alongside your vectors. This is incredibly useful for filtering search results later.
Metadata can include properties like a document's title, author, category, or creation date. It's stored as a dictionary within each vector object.
Demo: Upsert with Metadata
Let's add some context to our vectors using metadata. This makes them much more powerful for real-world applications.
import os
from pinecone import Pinecone
pc = Pinecone(
api_key=os.environ.get("PINECONE_API_KEY", "YOUR_API_KEY"),
environment=os.environ.get("PINECONE_ENVIRONMENT", "YOUR_ENVIRONMENT")
)
index = pc.Index("my-index") # Assuming 'my-index' exists
# Upsert a vector with metadata
index.upsert(
vectors=[
{
"id": "doc1",
"values": [0.2, 0.3, 0.4],
"metadata": {"genre": "sci-fi", "year": 2023}
},
{
"id": "doc2",
"values": [0.5, 0.6, 0.7],
"metadata": {"genre": "fantasy", "author": "J. Doe"}
}
]
)
print("Vectors 'doc1' and 'doc2' upserted with metadata!")How Upsert Handles Updates
The 'update' part of 'upsert' means that if you try to upsert a vector with an id that already exists in your index, Pinecone won't create a new entry.
Instead, it will overwrite the existing vector's values and metadata with the new data you provide. This ensures data consistency and avoids duplicates.
Upserting Knowledge Check
You've learned about upserting. Let's check your understanding!
Recap: Mastering Upserts
Great job! You've learned the essentials of upserting data into Pinecone:
- Upsert = Update + Insert: A single operation for adding new or modifying existing vectors.
- IDs are Key: Each vector requires a unique ID for identification and updates.
- Batching for Performance: Always upsert multiple vectors in batches for efficiency.
- Metadata for Context: Add key-value pairs to vectors for powerful filtering and search.
- Updates are Automatic: Upserting with an existing ID overwrites the old vector.
Next, we'll explore how to query these vectors to find similar items!
Belajar Vector Databases: Pinecone, Weaviate & pgvector dengan tutor AI — gratis
Tulis dan jalankan kode asli di browser kamu, dapatkan bantuan instan dari tutor AI 24/7, dan lanjutkan di mana kamu tinggalkan di web atau aplikasi.
- Kursus
- 12
- Pelajaran
- 48
Pertanyaan yang Sering Diajukan
Apakah pelajaran “Memasukkan Data ke Pinecone” gratis?
Ya — teks lengkap “Memasukkan Data ke Pinecone” gratis dibaca di sini di web. Untuk praktiknya secara interaktif (editor kode bawaan dan tutor AI 24/7) dan buka sisa kursus Vector Databases: Pinecone, Weaviate & pgvector, upgrade ke CoddyKit PRO. Kursus Vector Databases: Pinecone, Weaviate & pgvector mencakup 4 pelajaran total.
Apa yang akan aku pelajari di “Memasukkan Data ke Pinecone”?
Kuasai proses menyisipkan dan memperbarui data vektor beserta metadata terkait ke dalam indeks Pinecone Anda. Kamu berlatih Vector Databases: Pinecone, Weaviate & pgvector dengan kode praktik yang langsung kamu jalankan di browser, dan tutor AI 24/7 menjawab pertanyaanmu saat kamu mengerjakan pelajaran ini.
Apakah aku perlu pengalaman untuk memulai Vector Databases: Pinecone, Weaviate & pgvector?
Tidak diperlukan pengalaman sebelumnya. Vector Databases: Pinecone, Weaviate & pgvector di CoddyKit dirancang untuk pemula hingga pelajar tingkat lanjut, jadi kamu bisa memulai di sini atau dari awal dan belajar sesuai kecepatan kamu sendiri. Ini adalah pelajaran 2 dari 4.
Berapa lama pelajaran “Memasukkan Data ke Pinecone” memakan waktu?
Sebagian besar pelajaran CoddyKit memakan waktu sekitar 5–10 menit. Setiap pelajaran ringkas dan interaktif, jadi kamu membuat kemajuan stabil dan melanjutkan dari tempat kamu tinggalkan di web dan aplikasi.
Bisakah aku menulis dan menjalankan kode dalam pelajaran Vector Databases: Pinecone, Weaviate & pgvector ini?
Ya. Setiap pelajaran Vector Databases: Pinecone, Weaviate & pgvector menyertakan editor kode bawaan, jadi kamu menulis dan menjalankan kode nyata langsung di browser dan mendapatkan umpan balik AI instan — tidak diperlukan penyiapan lokal.
Semua pelajaran dalam kursus ini
- Membuat Indeks Pinecone
- Memasukkan Data ke Pinecone
- Membuat Kueri Data Vektor di Pinecone
- Memahami Harga dan Pod Pinecone