0Pricing
Web Scraping & Bots · Pelajaran

Distribusi Tugas Berbasis Antrean

Pelajari cara antrean pesan seperti Redis dan Celery memisahkan penemuan URL dari pengambilan untuk menskalakan scraping di banyak pekerja.

Distribusi Tugas Berbasis Antrean adalah pelajaran Web Scraping & Bots gratis di CoddyKit. Ini adalah pelajaran 4 dari 4. Kamu bisa membaca pelajaran lengkapnya di bawah secara gratis — lalu praktikkan langsung di browser dengan editor kode bawaan dan tutor AI 24/7. Ini adalah bagian dari jalur belajar Web Scraping & Bots, dan progresmu tersinkronisasi di web dan aplikasi CoddyKit. Kursus Web Scraping & Bots mencakup 4 pelajaran total.

Bagian dari pelajaran ini belum diterjemahkan dan ditampilkan dalam bahasa Inggris.

The Scaling Bottleneck

A single-process scraper is limited by one machine's CPU and network. To scale, you split work across many workers running in parallel, possibly on different servers.

A task queue is the glue that distributes work safely.

Producers and Consumers

The queue pattern has two roles:

  • Producers discover URLs and push tasks onto the queue.
  • Consumers (workers) pull tasks and fetch the pages.

Decoupling them lets each side scale independently.

A Redis-Backed Queue

Redis lists make a simple, fast queue. Producers LPUSH URLs; workers BRPOP them, blocking until work is available.

import redis
r = redis.Redis()

# producer
r.lpush('urls', 'https://site.com/page1')

# worker
_, url = r.brpop('urls')
print('processing', url)

Why Not Share a Python List

An in-memory list only works within one process. A Redis queue is shared across processes and machines, persists if a worker crashes, and handles concurrency atomically.

Introducing Celery

Celery is a full task framework built on a broker like Redis. You define tasks as functions and call them asynchronously; workers pick them up automatically.

from celery import Celery
app = Celery('scraper', broker='redis://localhost:6379/0')

@app.task
def scrape(url):
    return fetch_and_parse(url)

Dispatching Tasks

Calling .delay() enqueues the task and returns immediately. Workers running celery worker consume and execute them in parallel.

for url in discovered_urls:
    scrape.delay(url)

Retries and Failures

Celery can automatically retry failed tasks with backoff, so a transient network error does not lose a URL.

@app.task(bind=True, max_retries=3, default_retry_delay=10)
def scrape(self, url):
    try:
        return fetch_and_parse(url)
    except ConnectionError as e:
        raise self.retry(exc=e)

Deduplicating URLs

In distributed crawling the same URL can be discovered twice. Use a Redis set as a 'seen' filter so each page is fetched once.

if r.sadd('seen', url):
    scrape.delay(url)  # sadd returns 1 only if newly added

Backpressure and Concurrency

Tune worker concurrency to match target-site politeness and your bandwidth. Too many workers overwhelm the site; too few leave the queue backed up. Monitor queue length to find balance.

# start 4 worker processes
// celery -A scraper worker --concurrency=4

Results and Storage

Workers should write parsed data to a shared store (a database or object storage), not return it through the queue. The queue carries tasks; the datastore holds results.

Priority Queues

Not all URLs are equal. Route urgent tasks (a category index that unlocks many child pages) to a high-priority queue so workers handle them before low-value pages.

scrape.apply_async(args=[url], priority=9)  # higher runs sooner

Quick Check

Test your understanding of queue-based distribution.

Recap

You learned to scale scraping with task queues: the producer/consumer pattern, Redis-backed queues, Celery tasks with .delay(), automatic retries, URL deduplication with Redis sets, tuning concurrency, and writing results to shared storage.

Pertanyaan yang Sering Diajukan

Apakah pelajaran “Distribusi Tugas Berbasis Antrean” gratis?

Ya — teks lengkap “Distribusi Tugas Berbasis Antrean” gratis dibaca di sini di web. Untuk praktiknya secara interaktif (editor kode bawaan dan tutor AI 24/7) dan buka sisa kursus Web Scraping & Bots, upgrade ke CoddyKit PRO. Kursus Web Scraping & Bots mencakup 4 pelajaran total.

Apa yang akan aku pelajari di “Distribusi Tugas Berbasis Antrean”?

Pelajari cara antrean pesan seperti Redis dan Celery memisahkan penemuan URL dari pengambilan untuk menskalakan scraping di banyak pekerja. Kamu berlatih Web Scraping & Bots dengan kode praktik yang langsung kamu jalankan di browser, dan tutor AI 24/7 menjawab pertanyaanmu saat kamu mengerjakan pelajaran ini.

Apakah aku perlu pengalaman untuk memulai Web Scraping & Bots?

Tidak diperlukan pengalaman sebelumnya. Web Scraping & Bots di CoddyKit dirancang untuk pemula hingga pelajar tingkat lanjut, jadi kamu bisa memulai di sini atau dari awal dan belajar sesuai kecepatan kamu sendiri. Ini adalah pelajaran 4 dari 4.

Berapa lama pelajaran “Distribusi Tugas Berbasis Antrean” memakan waktu?

Sebagian besar pelajaran CoddyKit memakan waktu sekitar 5–10 menit. Setiap pelajaran ringkas dan interaktif, jadi kamu membuat kemajuan stabil dan melanjutkan dari tempat kamu tinggalkan di web dan aplikasi.

Bisakah aku menulis dan menjalankan kode dalam pelajaran Web Scraping & Bots ini?

Ya. Setiap pelajaran Web Scraping & Bots menyertakan editor kode bawaan, jadi kamu menulis dan menjalankan kode nyata langsung di browser dan mendapatkan umpan balik AI instan — tidak diperlukan penyiapan lokal.

Semua pelajaran dalam kursus ini

  1. Scraping Terdistribusi dengan Scrapy
  2. Cloud Functions untuk Scraping
  3. Pemantauan dan Pencatatan Aktivitas
  4. Distribusi Tugas Berbasis Antrean
← Kembali ke Web Scraping & Bots