Menyimpan Data dalam Basis Data NoSQL
Pelajari kapan dan cara menyimpan data hasil scraping secara permanen di penyimpanan NoSQL berorientasi dokumen seperti MongoDB untuk penyimpanan fleksibel dengan skema minimal.
Menyimpan Data dalam Basis Data NoSQL adalah pelajaran Web Scraping & Bots gratis di CoddyKit. Ini adalah pelajaran 4 dari 4. Kamu bisa membaca pelajaran lengkapnya di bawah secara gratis — lalu praktikkan langsung di browser dengan editor kode bawaan dan tutor AI 24/7. Ini adalah bagian dari jalur belajar Web Scraping & Bots, dan progresmu tersinkronisasi di web dan aplikasi CoddyKit. Kursus Web Scraping & Bots mencakup 4 pelajaran total.
Bagian dari pelajaran ini belum diterjemahkan dan ditampilkan dalam bahasa Inggris.
When SQL Is Not Enough
Scraped data is often irregular: different pages yield different fields, nested structures, and evolving shapes. Rigid SQL schemas can be painful here.
NoSQL document databases store flexible JSON-like records, making them a natural fit for messy web data.
Documents and Collections
In a document store like MongoDB:
- A document is one JSON-like record.
- A collection groups related documents (like a table).
- Documents in one collection need not share the same fields.
{
"title": "Widget",
"price": 9.99,
"tags": ["tools", "sale"],
"vendor": { "name": "Acme", "rating": 4.5 }
}Connecting with PyMongo
The pymongo driver connects Python to MongoDB. Grab a database and a collection handle to start writing.
from pymongo import MongoClient
client = MongoClient('mongodb://localhost:27017')
db = client['scraping']
products = db['products']Inserting Documents
Insert a single scraped record with insert_one or a batch with insert_many. MongoDB assigns an _id automatically.
record = {'title': 'Gadget', 'price': 14.5, 'tags': ['new']}
result = products.insert_one(record)
print(result.inserted_id)Avoiding Duplicates with Upsert
Re-running a scraper should not create duplicate rows. An upsert updates the matching document or inserts it if absent, keyed by a stable field like the product URL.
products.update_one(
{'url': record['url']},
{'$set': record},
upsert=True
)Unique Indexes
Enforce uniqueness at the database level with an index. This protects integrity even if your code has a bug.
products.create_index('url', unique=True)Querying Stored Data
Retrieve records with filter documents. Operators like $gt and $in express conditions.
cheap = products.find({'price': {'$lt': 10}})
for doc in cheap:
print(doc['title'], doc['price'])Storing Nested and Array Data
Unlike flat SQL columns, documents keep nested objects and arrays natively. This preserves the original structure of scraped pages without join tables.
review_doc = {
'product': 'Widget',
'reviews': [
{'user': 'a', 'stars': 5},
{'user': 'b', 'stars': 4}
]
}
db['catalog'].insert_one(review_doc)Bulk Writes for Speed
For large scrapes, batch operations dramatically reduce round trips. Collect writes and flush them together.
from pymongo import UpdateOne
ops = [UpdateOne({'url': r['url']}, {'$set': r}, upsert=True) for r in batch]
products.bulk_write(ops)SQL vs NoSQL for Scraping
Choose based on your data:
- NoSQL for variable, nested, fast-changing records.
- SQL when fields are stable and you need joins or strict constraints.
Many pipelines stage raw data in NoSQL, then transform into SQL for analysis.
Adding Timestamps and Metadata
Always stamp each scraped document with when it was captured and its source. This lets you track freshness, debug bad runs, and re-scrape stale records selectively.
from datetime import datetime
record['scraped_at'] = datetime.utcnow()
record['source'] = 'site.com'
products.insert_one(record)Quick Check
Test your understanding of NoSQL storage.
Recap
You learned to persist scraped data in NoSQL: documents and collections, connecting with PyMongo, upserts and unique indexes to prevent duplicates, querying, nested data, and bulk writes.
Document stores give scrapers flexible, scalable persistence.
Pertanyaan yang Sering Diajukan
Apakah pelajaran “Menyimpan Data dalam Basis Data NoSQL” gratis?
Ya — teks lengkap “Menyimpan Data dalam Basis Data NoSQL” gratis dibaca di sini di web. Untuk praktiknya secara interaktif (editor kode bawaan dan tutor AI 24/7) dan buka sisa kursus Web Scraping & Bots, upgrade ke CoddyKit PRO. Kursus Web Scraping & Bots mencakup 4 pelajaran total.
Apa yang akan aku pelajari di “Menyimpan Data dalam Basis Data NoSQL”?
Pelajari kapan dan cara menyimpan data hasil scraping secara permanen di penyimpanan NoSQL berorientasi dokumen seperti MongoDB untuk penyimpanan fleksibel dengan skema minimal. Kamu berlatih Web Scraping & Bots dengan kode praktik yang langsung kamu jalankan di browser, dan tutor AI 24/7 menjawab pertanyaanmu saat kamu mengerjakan pelajaran ini.
Apakah aku perlu pengalaman untuk memulai Web Scraping & Bots?
Tidak diperlukan pengalaman sebelumnya. Web Scraping & Bots di CoddyKit dirancang untuk pemula hingga pelajar tingkat lanjut, jadi kamu bisa memulai di sini atau dari awal dan belajar sesuai kecepatan kamu sendiri. Ini adalah pelajaran 4 dari 4.
Berapa lama pelajaran “Menyimpan Data dalam Basis Data NoSQL” memakan waktu?
Sebagian besar pelajaran CoddyKit memakan waktu sekitar 5–10 menit. Setiap pelajaran ringkas dan interaktif, jadi kamu membuat kemajuan stabil dan melanjutkan dari tempat kamu tinggalkan di web dan aplikasi.
Bisakah aku menulis dan menjalankan kode dalam pelajaran Web Scraping & Bots ini?
Ya. Setiap pelajaran Web Scraping & Bots menyertakan editor kode bawaan, jadi kamu menulis dan menjalankan kode nyata langsung di browser dan mendapatkan umpan balik AI instan — tidak diperlukan penyiapan lokal.
Semua pelajaran dalam kursus ini
- Menyimpan Data dalam CSV/JSON
- Mengintegrasikan dengan Basis Data (SQL)
- Solusi Penyimpanan Cloud
- Menyimpan Data dalam Basis Data NoSQL