Przechowywanie danych w bazach NoSQL
Dowiedz się, kiedy i jak utrwalać zescrapowane dane w dokumentowych magazynach NoSQL, takich jak MongoDB, zapewniających elastyczne przechowywanie przy ograniczonym schemacie.
Przechowywanie danych w bazach NoSQL to bezpłatna lekcja Web Scraping & Bots na CoddyKit. To lekcja 4 z 4. Możesz przeczytać całą lekcję poniżej za darmo — a potem ćwiczyć ją interaktywnie w przeglądarce z wbudowanym edytorem kodu i tutorem AI dostępnym 24/7. To część ścieżki edukacyjnej Web Scraping & Bots, a Twój postęp synchronizuje się między webem a aplikacją CoddyKit. Kurs Web Scraping & Bots zawiera 4 lekcji w sumie.
Części tej lekcji nie zostały jeszcze przetłumaczone i są wyświetlane po angielsku.
When SQL Is Not Enough
Scraped data is often irregular: different pages yield different fields, nested structures, and evolving shapes. Rigid SQL schemas can be painful here.
NoSQL document databases store flexible JSON-like records, making them a natural fit for messy web data.
Documents and Collections
In a document store like MongoDB:
- A document is one JSON-like record.
- A collection groups related documents (like a table).
- Documents in one collection need not share the same fields.
{
"title": "Widget",
"price": 9.99,
"tags": ["tools", "sale"],
"vendor": { "name": "Acme", "rating": 4.5 }
}Connecting with PyMongo
The pymongo driver connects Python to MongoDB. Grab a database and a collection handle to start writing.
from pymongo import MongoClient
client = MongoClient('mongodb://localhost:27017')
db = client['scraping']
products = db['products']Inserting Documents
Insert a single scraped record with insert_one or a batch with insert_many. MongoDB assigns an _id automatically.
record = {'title': 'Gadget', 'price': 14.5, 'tags': ['new']}
result = products.insert_one(record)
print(result.inserted_id)Avoiding Duplicates with Upsert
Re-running a scraper should not create duplicate rows. An upsert updates the matching document or inserts it if absent, keyed by a stable field like the product URL.
products.update_one(
{'url': record['url']},
{'$set': record},
upsert=True
)Unique Indexes
Enforce uniqueness at the database level with an index. This protects integrity even if your code has a bug.
products.create_index('url', unique=True)Querying Stored Data
Retrieve records with filter documents. Operators like $gt and $in express conditions.
cheap = products.find({'price': {'$lt': 10}})
for doc in cheap:
print(doc['title'], doc['price'])Storing Nested and Array Data
Unlike flat SQL columns, documents keep nested objects and arrays natively. This preserves the original structure of scraped pages without join tables.
review_doc = {
'product': 'Widget',
'reviews': [
{'user': 'a', 'stars': 5},
{'user': 'b', 'stars': 4}
]
}
db['catalog'].insert_one(review_doc)Bulk Writes for Speed
For large scrapes, batch operations dramatically reduce round trips. Collect writes and flush them together.
from pymongo import UpdateOne
ops = [UpdateOne({'url': r['url']}, {'$set': r}, upsert=True) for r in batch]
products.bulk_write(ops)SQL vs NoSQL for Scraping
Choose based on your data:
- NoSQL for variable, nested, fast-changing records.
- SQL when fields are stable and you need joins or strict constraints.
Many pipelines stage raw data in NoSQL, then transform into SQL for analysis.
Adding Timestamps and Metadata
Always stamp each scraped document with when it was captured and its source. This lets you track freshness, debug bad runs, and re-scrape stale records selectively.
from datetime import datetime
record['scraped_at'] = datetime.utcnow()
record['source'] = 'site.com'
products.insert_one(record)Quick Check
Test your understanding of NoSQL storage.
Recap
You learned to persist scraped data in NoSQL: documents and collections, connecting with PyMongo, upserts and unique indexes to prevent duplicates, querying, nested data, and bulk writes.
Document stores give scrapers flexible, scalable persistence.
Często zadawane pytania
Czy lekcja „Przechowywanie danych w bazach NoSQL” jest bezpłatna?
Tak — pełny tekst „Przechowywanie danych w bazach NoSQL” jest dostępny za darmo tutaj w sieci. Aby ćwiczyć ją interaktywnie (wbudowany edytor kodu i tutor AI dostępny 24/7) i odblokować resztę kursu Web Scraping & Bots, przejdź na CoddyKit PRO. Kurs Web Scraping & Bots zawiera 4 lekcji w sumie.
Co nauczysz się w „Przechowywanie danych w bazach NoSQL”?
Dowiedz się, kiedy i jak utrwalać zescrapowane dane w dokumentowych magazynach NoSQL, takich jak MongoDB, zapewniających elastyczne przechowywanie przy ograniczonym schemacie. Ćwiczysz Web Scraping & Bots z praktycznym kodem, który uruchamiasz bezpośrednio w przeglądarce, a tutor AI dostępny 24/7 odpowiada na Twoje pytania podczas pracy nad lekcją.
Czy potrzebuję doświadczenia, aby zacząć Web Scraping & Bots?
Nie wymagamy żadnego doświadczenia. Web Scraping & Bots w CoddyKit jest strukturyzowany dla początkujących i zaawansowanych użytkowników, więc możesz zacząć tutaj lub od początku i uczyć się w swoim tempie. To lekcja 4 z 4.
Ile czasu zajmuje lekcja „Przechowywanie danych w bazach NoSQL”?
Większość lekcji CoddyKit trwa około 5–10 minut. Każda lekcja to mały, interaktywny krok, dzięki czemu robisz systematyczne postępy i zawsze wracasz dokładnie do tego samego miejsca — na webie i w aplikacji.
Czy mogę pisać i uruchamiać kod w tej lekcji Web Scraping & Bots?
Tak. Każda lekcja Web Scraping & Bots zawiera wbudowany edytor kodu, więc piszesz i uruchamiasz prawdziwy kod bezpośrednio w przeglądarce i od razu otrzymujesz sprzężenie zwrotne od AI — bez konfiguracji na komputerze.
Wszystkie lekcje w tym kursie
- Przechowywanie danych w formatach CSV/JSON
- Integracja z bazami danych (SQL)
- Rozwiązania do przechowywania danych w chmurze
- Przechowywanie danych w bazach NoSQL