Web Scraping & Bots · Lezione

Comprendere robots.txt

Interpreti e rispetti il file `robots.txt` per comprendere le politiche e le limitazioni di scraping di un sito web.

Lezione 1 di 411 passaggi

Comprendere robots.txt è una lezione Web Scraping & Bots gratuita su CoddyKit. Questa è la lezione 1 di 4. Puoi leggere la lezione completa qui gratuitamente — poi esercitati direttamente nel browser con un editor di codice integrato e un tutor IA disponibile 24/7. Fa parte del percorso di apprendimento Web Scraping & Bots, e i tuoi progressi si sincronizzano tra il web e l'app CoddyKit. Il corso Web Scraping & Bots include 4 lezioni in totale.

Parti di questa lezione non sono ancora state tradotte e vengono mostrate in inglese.

What is Robots.txt?

When building a web scraper or bot, it's crucial to be a good internet citizen. The robots.txt file is a key part of this.

It's a text file that websites use to communicate with web crawlers and other bots. It tells them which parts of the site they are allowed to access and which parts they should avoid.

Finding Robots.txt

Every website that uses a robots.txt file places it in a standard location: the root directory of its domain.

This means you can always find it by adding /robots.txt to the end of the website's main URL. For example:

  • https://www.example.com/robots.txt
  • https://www.google.com/robots.txt

You can simply type this into your browser to view a site's rules.

The User-agent Directive

The User-agent directive specifies which bot the following rules apply to. Think of it as addressing a specific bot or all bots.

  • User-agent: *: This applies to ALL web crawlers and bots.
  • User-agent: Googlebot: This applies only to Google's specific web crawler.
  • User-agent: MyCustomBot: You can even specify rules for your own bot if the website owner knows its name.

Each set of rules starts with a User-agent line.

Blocking Access: Disallow

The Disallow directive is used to tell bots which URLs or directories they should NOT access. It's the primary way to restrict crawling.

Here are some examples:

  • Disallow: /: Disallows access to the entire website (except for robots.txt itself).
  • Disallow: /private/: Disallows access to the /private/ directory and everything within it.
  • Disallow: /search?: Disallows URLs starting with /search?, often used for search results pages.

Always respect these rules!

Allowing Exceptions: Allow

Sometimes, a website might want to disallow a whole directory but allow access to a specific file or sub-directory within it. This is where the Allow directive comes in.

Allow rules override Disallow rules for more specific paths.

For example:

User-agent: *
Disallow: /images/
Allow: /images/public/

This means all bots should avoid the /images/ folder, but they ARE allowed to access content within /images/public/.

Guiding with Sitemap

The Sitemap directive isn't about restricting access; it's about helping bots discover content.

It points to the XML Sitemap file(s) for the website. A sitemap lists all the pages and files a website owner wants search engines to crawl and index.

Example:

Sitemap: https://www.example.com/sitemap.xml

This helps well-behaved bots find your content more efficiently.

Fetching Robots.txt with Python

You can easily fetch a website's robots.txt file using Python's requests library. This allows your script to programmatically read and interpret the rules.

Try running this example to see the robots.txt for Wikipedia:

import requests

def get_robots_txt(domain):
    try:
        response = requests.get(f"https://{domain}/robots.txt")
        response.raise_for_status() # Raise HTTPError for bad responses
        print(f"--- {domain}/robots.txt ---")
        print(response.text)
        print("--------------------------")
    except requests.exceptions.RequestException as e:
        print(f"Error fetching robots.txt for {domain}: {e}")

if __name__ == "__main__":
    get_robots_txt("www.wikipedia.org")
    # You can try other domains too!
    # get_robots_txt("www.google.com")

Interpreting Complex Rules

Let's look at a combined example to understand how rules interact:

User-agent: *
Disallow: /temp/
Disallow: /admin/
Allow: /admin/public/

User-agent: MyBot
Disallow: /
  • A general bot (*) cannot access /temp/ or /admin/, but it CAN access /admin/public/.
  • A bot named MyBot cannot access ANYTHING on the site.

The most specific rule usually wins, especially Allow over Disallow for sub-paths.

Robots.txt is a Guideline, Not Security

It's crucial to understand that robots.txt is a voluntary agreement for well-behaved bots. It's not a security mechanism!

  • Malicious bots can (and often will) ignore these rules.
  • The content of robots.txt itself is public. Don't put sensitive information there.
  • It's for managing server load and respecting content preferences, not hiding data.

Always scrape ethically and respect website policies.

Quick Check: Robots.txt Rules

Consider the following robots.txt content:

User-agent: *
Disallow: /private/
Allow: /private/data.html
Disallow: /temp/

According to these rules, which path is a general bot (User-agent: *) explicitly allowed to access?

Recap: Respecting Robots.txt

In this lesson, we explored the robots.txt file, a fundamental component of ethical web scraping.

  • You learned how to locate it and its core directives: User-agent, Disallow, Allow, and Sitemap.
  • We saw how to fetch and interpret these rules using Python.
  • Crucially, we emphasized that robots.txt is a guideline for respectful bots, not a security measure.

Always check and respect a website's robots.txt before scraping!

Gratis per iniziare

Impara Python con un tutor IA — gratis

Scrivi ed esegui vero codice nel tuo browser, ricevi aiuto istantaneo da un tutor IA disponibile 24/7, e riprendi da dove hai lasciato sul web o nell'app.

Corsi
12
Lezioni
48

Domande Frequenti

La lezione «Comprendere robots.txt» è gratuita?

Sì — il testo completo di «Comprendere robots.txt» è gratuito qui sul web. Per esercitarvi in modo interattivo (un editor di codice integrato e un tutor IA 24/7) e sbloccare il resto del corso Web Scraping & Bots, passa a CoddyKit PRO. Il corso Web Scraping & Bots include 4 lezioni in totale.

Cosa imparerò in «Comprendere robots.txt»?

Interpreti e rispetti il file `robots.txt` per comprendere le politiche e le limitazioni di scraping di un sito web. Eserciti Web Scraping & Bots con codice pratico che esegui direttamente nel browser, e un tutor IA 24/7 risponde alle tue domande mentre lavori sulla lezione.

Ho bisogno di esperienza per iniziare Web Scraping & Bots?

Non è richiesta alcuna esperienza precedente. Web Scraping & Bots su CoddyKit è strutturato per principianti e studenti avanzati, quindi puoi iniziare da qui o dall'inizio e procedere al tuo ritmo. Questa è la lezione 1 di 4.

Quanto tempo richiede la lezione «Comprendere robots.txt»?

La maggior parte delle lezioni CoddyKit richiede circa 5–10 minuti. Ogni lezione è breve e interattiva, quindi fai progressi costanti e riprendi esattamente da dove hai lasciato su web e app.

Posso scrivere ed eseguire codice in questa lezione Web Scraping & Bots?

Sì. Ogni lezione Web Scraping & Bots include un editor di codice integrato, quindi scrivi ed esegui codice reale direttamente nel tuo browser e ricevi feedback istantaneo dall'IA — nessuna configurazione locale necessaria.

Tutte le lezioni di questo corso

  1. Comprendere robots.txt
  2. Termini di servizio e copyright
  3. Pratiche etiche di scraping
  4. Limitazione della frequenza e crawling rispettoso
← Torna a Web Scraping & Bots