Comprensión de Robots.txt
Interprete y respete el archivo `robots.txt` para comprender las políticas y restricciones de scraping de un sitio web.
Comprensión de Robots.txt es una lección gratuita de Web Scraping & Bots en CoddyKit. Esta es la lección 1 de 4. Puedes leer la lección completa abajo gratuitamente — luego la practicas en el navegador con un editor de código integrado y un tutor de IA 24/7. Forma parte de la ruta de aprendizaje de Web Scraping & Bots, y tu progreso se sincroniza en la web y la app de CoddyKit. El curso de Web Scraping & Bots incluye 4 lecciones en total.
Partes de esta lección aún no han sido traducidas y se muestran en inglés.
What is Robots.txt?
When building a web scraper or bot, it's crucial to be a good internet citizen. The robots.txt file is a key part of this.
It's a text file that websites use to communicate with web crawlers and other bots. It tells them which parts of the site they are allowed to access and which parts they should avoid.
Finding Robots.txt
Every website that uses a robots.txt file places it in a standard location: the root directory of its domain.
This means you can always find it by adding /robots.txt to the end of the website's main URL. For example:
https://www.example.com/robots.txthttps://www.google.com/robots.txt
You can simply type this into your browser to view a site's rules.
The User-agent Directive
The User-agent directive specifies which bot the following rules apply to. Think of it as addressing a specific bot or all bots.
User-agent: *: This applies to ALL web crawlers and bots.User-agent: Googlebot: This applies only to Google's specific web crawler.User-agent: MyCustomBot: You can even specify rules for your own bot if the website owner knows its name.
Each set of rules starts with a User-agent line.
Blocking Access: Disallow
The Disallow directive is used to tell bots which URLs or directories they should NOT access. It's the primary way to restrict crawling.
Here are some examples:
Disallow: /: Disallows access to the entire website (except forrobots.txtitself).Disallow: /private/: Disallows access to the/private/directory and everything within it.Disallow: /search?: Disallows URLs starting with/search?, often used for search results pages.
Always respect these rules!
Allowing Exceptions: Allow
Sometimes, a website might want to disallow a whole directory but allow access to a specific file or sub-directory within it. This is where the Allow directive comes in.
Allow rules override Disallow rules for more specific paths.
For example:
User-agent: *
Disallow: /images/
Allow: /images/public/This means all bots should avoid the /images/ folder, but they ARE allowed to access content within /images/public/.
Guiding with Sitemap
The Sitemap directive isn't about restricting access; it's about helping bots discover content.
It points to the XML Sitemap file(s) for the website. A sitemap lists all the pages and files a website owner wants search engines to crawl and index.
Example:
Sitemap: https://www.example.com/sitemap.xmlThis helps well-behaved bots find your content more efficiently.
Fetching Robots.txt with Python
You can easily fetch a website's robots.txt file using Python's requests library. This allows your script to programmatically read and interpret the rules.
Try running this example to see the robots.txt for Wikipedia:
import requests
def get_robots_txt(domain):
try:
response = requests.get(f"https://{domain}/robots.txt")
response.raise_for_status() # Raise HTTPError for bad responses
print(f"--- {domain}/robots.txt ---")
print(response.text)
print("--------------------------")
except requests.exceptions.RequestException as e:
print(f"Error fetching robots.txt for {domain}: {e}")
if __name__ == "__main__":
get_robots_txt("www.wikipedia.org")
# You can try other domains too!
# get_robots_txt("www.google.com")Interpreting Complex Rules
Let's look at a combined example to understand how rules interact:
User-agent: *
Disallow: /temp/
Disallow: /admin/
Allow: /admin/public/
User-agent: MyBot
Disallow: /- A general bot (
*) cannot access/temp/or/admin/, but it CAN access/admin/public/. - A bot named
MyBotcannot access ANYTHING on the site.
The most specific rule usually wins, especially Allow over Disallow for sub-paths.
Robots.txt is a Guideline, Not Security
It's crucial to understand that robots.txt is a voluntary agreement for well-behaved bots. It's not a security mechanism!
- Malicious bots can (and often will) ignore these rules.
- The content of
robots.txtitself is public. Don't put sensitive information there. - It's for managing server load and respecting content preferences, not hiding data.
Always scrape ethically and respect website policies.
Quick Check: Robots.txt Rules
Consider the following robots.txt content:
User-agent: *
Disallow: /private/
Allow: /private/data.html
Disallow: /temp/According to these rules, which path is a general bot (User-agent: *) explicitly allowed to access?
Recap: Respecting Robots.txt
In this lesson, we explored the robots.txt file, a fundamental component of ethical web scraping.
- You learned how to locate it and its core directives:
User-agent,Disallow,Allow, andSitemap. - We saw how to fetch and interpret these rules using Python.
- Crucially, we emphasized that
robots.txtis a guideline for respectful bots, not a security measure.
Always check and respect a website's robots.txt before scraping!
Preguntas frecuentes
¿La lección «Comprensión de Robots.txt» es gratis?
Sí — el texto completo de «Comprensión de Robots.txt» es gratis para leer aquí en la web. Para practicarla de forma interactiva (editor de código integrado y tutor de IA 24/7) y desbloquear el resto del curso de Web Scraping & Bots, actualiza a CoddyKit PRO. El curso de Web Scraping & Bots incluye 4 lecciones en total.
¿Qué aprenderé en «Comprensión de Robots.txt»?
Interprete y respete el archivo `robots.txt` para comprender las políticas y restricciones de scraping de un sitio web. Practicas Web Scraping & Bots con código real que ejecutas directamente en el navegador, y un tutor de IA 24/7 responde tus preguntas mientras trabajas en la lección.
¿Necesito experiencia previa para empezar Web Scraping & Bots?
No se requiere experiencia previa. Web Scraping & Bots en CoddyKit está estructurado para principiantes hasta estudiantes avanzados, así que puedes empezar aquí o desde el inicio y avanzar a tu ritmo. Esta es la lección 1 de 4.
¿Cuánto tiempo toma la lección «Comprensión de Robots.txt»?
La mayoría de las lecciones de CoddyKit toman alrededor de 5–10 minutos. Cada una es compacta e interactiva, así que avanzas constantemente y retomas exactamente por donde dejaste en la web y la app.
¿Puedo escribir y ejecutar código en esta lección de Web Scraping & Bots?
Sí. Cada lección de Web Scraping & Bots incluye un editor de código integrado, así que escribes y ejecutas código real directamente en tu navegador y obtienes retroalimentación instantánea de IA — sin configuración local necesaria.
Todas las lecciones de este curso
- Comprensión de Robots.txt
- Términos del servicio y derechos de autor
- Prácticas éticas de scraping
- Limitación de velocidad y rastreo respetuoso