Robots.txt وأخلاقيات الكشط
افهم الأسس القانونية والأخلاقية لكشط الويب: robots.txt وشروط الخدمة وتحديد معدّل الطلبات والسلوك المسؤول.
Robots.txt وأخلاقيات الكشط درس مجاني في Web Scraping & Bots على CoddyKit. هذا هو الدرس 4 من أصل 4. يمكنك قراءة الدرس كاملاً أدناه مجاناً — ثم تمرن عليه مباشرة في المتصفح باستخدام محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7. هذا الدرس جزء من مسار التعلم في Web Scraping & Bots، وتقدمك يتزامن عبر الويب وتطبيق CoddyKit. تتضمن دورة Web Scraping & Bots 4 دروس في المجموع.
بعض أجزاء هذا الدرس لم تُترجم بعد وتظهر باللغة الإنجليزية.
Scraping Responsibly
Scraping touches other people’s servers and data. Before your first scraper, learn the rules and ethics that keep you safe and respectful.
What Is robots.txt?
robots.txt lives at a site’s root and tells automated clients which paths they may or may not access.
# https://example.com/robots.txt
User-agent: *
Disallow: /admin/
Allow: /public/Reading the Directives
Read the directives: User-agent targets a crawler, Disallow blocks paths, Allow adds exceptions, Crawl-delay sets a wait.
Checking robots.txt in Python
Python’s built-in urllib.robotparser reads and evaluates robots.txt for you — just ask it whether you can fetch a URL.
from urllib.robotparser import RobotFileParser
rp = RobotFileParser()
rp.set_url('https://example.com/robots.txt')
rp.read()
print(rp.can_fetch('*', 'https://example.com/public/page'))robots.txt Is Not Law
Remember: robots.txt is a request, not a lock. Ignoring it isn’t a crime by itself, but it can breach terms of service. Respect it anyway.
Terms of Service
A site’s Terms of Service may flat-out forbid automated access. Breaking them can mean account bans or legal trouble — always check first.
Personal & Copyrighted Data
Tread carefully with personal data (privacy laws like GDPR) and copyrighted content. Collecting or republishing them can carry real consequences.
Rate Limiting
Rapid-fire requests can overload a server — and get you blocked. Add a delay between requests to stay polite.
import time
for url in urls:
fetch(url)
time.sleep(2) # be politeIdentify Yourself
Set an honest User-Agent header, ideally with contact info, so site owners can reach you instead of just blocking you.
headers = {
'User-Agent': 'MyResearchBot/1.0 (contact@example.com)'
}Prefer Official APIs
If a site offers an API, use it. APIs are stable, sanctioned, and far gentler than scraping raw HTML.
Cache to Reduce Load
Cache responses locally so you don’t re-fetch the same page over and over while developing.
import os
if not os.path.exists('page.html'):
save(fetch(url), 'page.html')
html = open('page.html').read()Quick Check
What is the correct way to think about robots.txt?
Recap
You’ve got scraping ethics: respecting robots.txt, terms, privacy and copyright, rate limiting, honest User-Agents, and preferring APIs and caching.
الأسئلة الشائعة
هل درس «Robots.txt وأخلاقيات الكشط» مجاني؟
نعم — نص درس «Robots.txt وأخلاقيات الكشط» كامل متاح مجاناً هنا على الويب. لتمرينه بشكل تفاعلي (محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7) وفتح باقي دورة Web Scraping & Bots، انتقل إلى CoddyKit PRO. تتضمن دورة Web Scraping & Bots 4 دروس في المجموع.
ماذا ستتعلم في «Robots.txt وأخلاقيات الكشط»؟
افهم الأسس القانونية والأخلاقية لكشط الويب: robots.txt وشروط الخدمة وتحديد معدّل الطلبات والسلوك المسؤول. تتمرن على Web Scraping & Bots مع أكواد عملية تشغلها مباشرة في المتصفح، ومدرس ذكاء اصطناعي متاح 24/7 يجيب على أسئلتك أثناء عملك.
هل أحتاج إلى خبرة سابقة لأبدأ Web Scraping & Bots؟
لا تُشترط خبرة سابقة. Web Scraping & Bots على CoddyKit منظم للمبتدئين حتى المتقدمين، لذا يمكنك البدء من هنا أو من البداية والتقدم بسرعتك الخاصة. هذا هو الدرس 4 من أصل 4.
كم من الوقت يستغرق درس «Robots.txt وأخلاقيات الكشط»؟
معظم دروس CoddyKit تستغرق حوالي 5–10 دقائق. كل منها موجز وتفاعلي، لذا تحرز تقدماً مستمراً وتستأنف من حيث توقفت عبر الويب والتطبيق.
هل يمكنني كتابة وتشغيل أكواد في درس Web Scraping & Bots هذا؟
نعم. كل درس في Web Scraping & Bots يتضمن محرر أكواد مدمج، لذا تكتب وتشغل أكواداً حقيقية مباشرة في متصفحك وتحصل على تعليقات فورية من الذكاء الاصطناعي — بدون إعداد محلي.
جميع الدروس في هذه الدورة
- ما هو كشط الويب؟
- طلبات HTTP واستجاباته
- فحص صفحات الويب
- Robots.txt وأخلاقيات الكشط