NoSQL データベースにデータを保存する
MongoDB のようなドキュメント指向 NoSQL ストアにスクレイピングしたデータを保存する適切な場面と方法を学び、柔軟でスキーマに縛られない保存を実現します。
「NoSQL データベースにデータを保存する」はCoddyKit上の無料Web Scraping & Botsレッスンです。 これはレッスン4/4です。 下記で完全なレッスンを無料で読むことができます。その後、ブラウザ内の組み込みコードエディタと24時間対応のAIチューターでハンズオン演習できます。 これはWeb Scraping & Bots学習パスの一部であり、ウェブとCoddyKitアプリ全体で進捗が同期されます。 Web Scraping & Botsコースには全4レッスンが含まれています。
このレッスンの一部はまだ翻訳されておらず、英語で表示されています。
When SQL Is Not Enough
Scraped data is often irregular: different pages yield different fields, nested structures, and evolving shapes. Rigid SQL schemas can be painful here.
NoSQL document databases store flexible JSON-like records, making them a natural fit for messy web data.
Documents and Collections
In a document store like MongoDB:
- A document is one JSON-like record.
- A collection groups related documents (like a table).
- Documents in one collection need not share the same fields.
{
"title": "Widget",
"price": 9.99,
"tags": ["tools", "sale"],
"vendor": { "name": "Acme", "rating": 4.5 }
}Connecting with PyMongo
The pymongo driver connects Python to MongoDB. Grab a database and a collection handle to start writing.
from pymongo import MongoClient
client = MongoClient('mongodb://localhost:27017')
db = client['scraping']
products = db['products']Inserting Documents
Insert a single scraped record with insert_one or a batch with insert_many. MongoDB assigns an _id automatically.
record = {'title': 'Gadget', 'price': 14.5, 'tags': ['new']}
result = products.insert_one(record)
print(result.inserted_id)Avoiding Duplicates with Upsert
Re-running a scraper should not create duplicate rows. An upsert updates the matching document or inserts it if absent, keyed by a stable field like the product URL.
products.update_one(
{'url': record['url']},
{'$set': record},
upsert=True
)Unique Indexes
Enforce uniqueness at the database level with an index. This protects integrity even if your code has a bug.
products.create_index('url', unique=True)Querying Stored Data
Retrieve records with filter documents. Operators like $gt and $in express conditions.
cheap = products.find({'price': {'$lt': 10}})
for doc in cheap:
print(doc['title'], doc['price'])Storing Nested and Array Data
Unlike flat SQL columns, documents keep nested objects and arrays natively. This preserves the original structure of scraped pages without join tables.
review_doc = {
'product': 'Widget',
'reviews': [
{'user': 'a', 'stars': 5},
{'user': 'b', 'stars': 4}
]
}
db['catalog'].insert_one(review_doc)Bulk Writes for Speed
For large scrapes, batch operations dramatically reduce round trips. Collect writes and flush them together.
from pymongo import UpdateOne
ops = [UpdateOne({'url': r['url']}, {'$set': r}, upsert=True) for r in batch]
products.bulk_write(ops)SQL vs NoSQL for Scraping
Choose based on your data:
- NoSQL for variable, nested, fast-changing records.
- SQL when fields are stable and you need joins or strict constraints.
Many pipelines stage raw data in NoSQL, then transform into SQL for analysis.
Adding Timestamps and Metadata
Always stamp each scraped document with when it was captured and its source. This lets you track freshness, debug bad runs, and re-scrape stale records selectively.
from datetime import datetime
record['scraped_at'] = datetime.utcnow()
record['source'] = 'site.com'
products.insert_one(record)Quick Check
Test your understanding of NoSQL storage.
Recap
You learned to persist scraped data in NoSQL: documents and collections, connecting with PyMongo, upserts and unique indexes to prevent duplicates, querying, nested data, and bulk writes.
Document stores give scrapers flexible, scalable persistence.
よくある質問
「NoSQL データベースにデータを保存する」レッスンは無料ですか?
はい。「NoSQL データベースにデータを保存する」の完全なテキストはこのウェブで無料で読めます。インタラクティブに演習し(組み込みコードエディタと24時間対応のAIチューター)、Web Scraping & Botsコースの残りをアンロックするには、CoddyKit PROにアップグレードしてください。 Web Scraping & Botsコースには全4レッスンが含まれています。
「NoSQL データベースにデータを保存する」で何を学びますか?
MongoDB のようなドキュメント指向 NoSQL ストアにスクレイピングしたデータを保存する適切な場面と方法を学び、柔軟でスキーマに縛られない保存を実現します。 ブラウザで直接実行するハンズオンコードでWeb Scraping & Botsを演習し、24時間対応のAIチューターがレッスンを進める中での質問に答えます。
Web Scraping & Botsを始めるのに経験は必要ですか?
事前経験は必要ありません。CoddyKitのWeb Scraping & Botsは初級者から上級者向けに構成されているため、ここから始めるか最初から始めて、自分のペースで進むことができます。 これはレッスン4/4です。
「NoSQL データベースにデータを保存する」レッスンにはどのくらい時間がかかりますか?
ほとんどのCoddyKitレッスンは約5~10分かかります。各レッスンはコンパクトでインタラクティブなので、着実に進歩し、ウェブとアプリ全体で正確に前回の場所から再開できます。
このWeb Scraping & Botsレッスンでコードを書いて実行できますか?
はい。すべてのWeb Scraping & Botsレッスンに組み込みコードエディタが含まれているため、ブラウザでリアルコードを書いて実行し、即座のAIフィードバックを取得できます。ローカル設定は不要です。
このコースのすべてのレッスン
- CSV/JSONへのデータ保存
- データベース(SQL)との統合
- クラウドストレージソリューション
- NoSQL データベースにデータを保存する