0Pricing
Web Scraping & Bots · درس

تخزين البيانات في قواعد بيانات NoSQL

تعلّم متى وكيف تحفظ البيانات المستخرجة في مخازن NoSQL الموجّهة للمستندات، مثل MongoDB، لتخزين مرن قليل القيود على المخطط.

تخزين البيانات في قواعد بيانات NoSQL درس مجاني في Web Scraping & Bots على CoddyKit. هذا هو الدرس 4 من أصل 4. يمكنك قراءة الدرس كاملاً أدناه مجاناً — ثم تمرن عليه مباشرة في المتصفح باستخدام محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7. هذا الدرس جزء من مسار التعلم في Web Scraping & Bots، وتقدمك يتزامن عبر الويب وتطبيق CoddyKit. تتضمن دورة Web Scraping & Bots 4 دروس في المجموع.

بعض أجزاء هذا الدرس لم تُترجم بعد وتظهر باللغة الإنجليزية.

When SQL Is Not Enough

Scraped data is often irregular: different pages yield different fields, nested structures, and evolving shapes. Rigid SQL schemas can be painful here.

NoSQL document databases store flexible JSON-like records, making them a natural fit for messy web data.

Documents and Collections

In a document store like MongoDB:

  • A document is one JSON-like record.
  • A collection groups related documents (like a table).
  • Documents in one collection need not share the same fields.
{
  "title": "Widget",
  "price": 9.99,
  "tags": ["tools", "sale"],
  "vendor": { "name": "Acme", "rating": 4.5 }
}

Connecting with PyMongo

The pymongo driver connects Python to MongoDB. Grab a database and a collection handle to start writing.

from pymongo import MongoClient

client = MongoClient('mongodb://localhost:27017')
db = client['scraping']
products = db['products']

Inserting Documents

Insert a single scraped record with insert_one or a batch with insert_many. MongoDB assigns an _id automatically.

record = {'title': 'Gadget', 'price': 14.5, 'tags': ['new']}
result = products.insert_one(record)
print(result.inserted_id)

Avoiding Duplicates with Upsert

Re-running a scraper should not create duplicate rows. An upsert updates the matching document or inserts it if absent, keyed by a stable field like the product URL.

products.update_one(
    {'url': record['url']},
    {'$set': record},
    upsert=True
)

Unique Indexes

Enforce uniqueness at the database level with an index. This protects integrity even if your code has a bug.

products.create_index('url', unique=True)

Querying Stored Data

Retrieve records with filter documents. Operators like $gt and $in express conditions.

cheap = products.find({'price': {'$lt': 10}})
for doc in cheap:
    print(doc['title'], doc['price'])

Storing Nested and Array Data

Unlike flat SQL columns, documents keep nested objects and arrays natively. This preserves the original structure of scraped pages without join tables.

review_doc = {
  'product': 'Widget',
  'reviews': [
    {'user': 'a', 'stars': 5},
    {'user': 'b', 'stars': 4}
  ]
}
db['catalog'].insert_one(review_doc)

Bulk Writes for Speed

For large scrapes, batch operations dramatically reduce round trips. Collect writes and flush them together.

from pymongo import UpdateOne

ops = [UpdateOne({'url': r['url']}, {'$set': r}, upsert=True) for r in batch]
products.bulk_write(ops)

SQL vs NoSQL for Scraping

Choose based on your data:

  • NoSQL for variable, nested, fast-changing records.
  • SQL when fields are stable and you need joins or strict constraints.

Many pipelines stage raw data in NoSQL, then transform into SQL for analysis.

Adding Timestamps and Metadata

Always stamp each scraped document with when it was captured and its source. This lets you track freshness, debug bad runs, and re-scrape stale records selectively.

from datetime import datetime

record['scraped_at'] = datetime.utcnow()
record['source'] = 'site.com'
products.insert_one(record)

Quick Check

Test your understanding of NoSQL storage.

Recap

You learned to persist scraped data in NoSQL: documents and collections, connecting with PyMongo, upserts and unique indexes to prevent duplicates, querying, nested data, and bulk writes.

Document stores give scrapers flexible, scalable persistence.

الأسئلة الشائعة

هل درس «تخزين البيانات في قواعد بيانات NoSQL» مجاني؟

نعم — نص درس «تخزين البيانات في قواعد بيانات NoSQL» كامل متاح مجاناً هنا على الويب. لتمرينه بشكل تفاعلي (محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7) وفتح باقي دورة Web Scraping & Bots، انتقل إلى CoddyKit PRO. تتضمن دورة Web Scraping & Bots 4 دروس في المجموع.

ماذا ستتعلم في «تخزين البيانات في قواعد بيانات NoSQL»؟

تعلّم متى وكيف تحفظ البيانات المستخرجة في مخازن NoSQL الموجّهة للمستندات، مثل MongoDB، لتخزين مرن قليل القيود على المخطط. تتمرن على Web Scraping & Bots مع أكواد عملية تشغلها مباشرة في المتصفح، ومدرس ذكاء اصطناعي متاح 24/7 يجيب على أسئلتك أثناء عملك.

هل أحتاج إلى خبرة سابقة لأبدأ Web Scraping & Bots؟

لا تُشترط خبرة سابقة. Web Scraping & Bots على CoddyKit منظم للمبتدئين حتى المتقدمين، لذا يمكنك البدء من هنا أو من البداية والتقدم بسرعتك الخاصة. هذا هو الدرس 4 من أصل 4.

كم من الوقت يستغرق درس «تخزين البيانات في قواعد بيانات NoSQL»؟

معظم دروس CoddyKit تستغرق حوالي 5–10 دقائق. كل منها موجز وتفاعلي، لذا تحرز تقدماً مستمراً وتستأنف من حيث توقفت عبر الويب والتطبيق.

هل يمكنني كتابة وتشغيل أكواد في درس Web Scraping & Bots هذا؟

نعم. كل درس في Web Scraping & Bots يتضمن محرر أكواد مدمج، لذا تكتب وتشغل أكواداً حقيقية مباشرة في متصفحك وتحصل على تعليقات فورية من الذكاء الاصطناعي — بدون إعداد محلي.

جميع الدروس في هذه الدورة

  1. تخزين البيانات بصيغة CSV/JSON
  2. الدمج مع قواعد البيانات (SQL)
  3. حلول التخزين السحابي
  4. تخزين البيانات في قواعد بيانات NoSQL
← العودة إلى Web Scraping & Bots