0Pricing
Web Scraping & Bots · Lección

Extracción de datos de tablas HTML

Aprenda a analizar de forma fiable datos tabulares HTML, gestionar rowspan y colspan, y convertir tablas desordenadas en filas estructuradas y limpias.

Extracción de datos de tablas HTML es una lección gratuita de Web Scraping & Bots en CoddyKit. Esta es la lección 4 de 4. Puedes leer la lección completa abajo gratuitamente — luego la practicas en el navegador con un editor de código integrado y un tutor de IA 24/7. Forma parte de la ruta de aprendizaje de Web Scraping & Bots, y tu progreso se sincroniza en la web y la app de CoddyKit. El curso de Web Scraping & Bots incluye 4 lecciones en total.

Partes de esta lección aún no han sido traducidas y se muestran en inglés.

Why Tables Are Tricky

HTML <table> elements look simple but are one of the most error-prone targets in scraping. Rows can merge cells, headers can repeat, and layout tables masquerade as data tables.

  • Data tables hold real records you want.
  • Layout tables only control visual structure.

This lesson focuses on extracting clean rows from genuine data tables.

Anatomy of a Table

A table is built from a few key tags:

  • <thead> / <tbody> group header and body rows.
  • <tr> is a single row.
  • <th> is a header cell, <td> is a data cell.

Knowing these landmarks lets you target rows precisely instead of grabbing raw text.

<table>
  <thead><tr><th>Name</th><th>Price</th></tr></thead>
  <tbody>
    <tr><td>Widget</td><td>$9.99</td></tr>
    <tr><td>Gadget</td><td>$14.50</td></tr>
  </tbody>
</table>

Selecting Rows

With a parser like BeautifulSoup you first find the table, then iterate its rows. Always scope your selection to tbody when present so the header row does not contaminate your data.

from bs4 import BeautifulSoup

soup = BeautifulSoup(html, 'html.parser')
table = soup.select_one('table')
rows = table.select('tbody tr')
print(len(rows), 'data rows found')

Reading Cell Values

For each row, collect the cell text. Use get_text(strip=True) to drop surrounding whitespace and nested tag noise.

for row in rows:
    cells = [c.get_text(strip=True) for c in row.select('td')]
    print(cells)

Mapping Headers to Values

Raw lists of cells are fragile. Pair each value with its column header so your output is self-describing and column-order changes do not break downstream code.

headers = [h.get_text(strip=True) for h in table.select('thead th')]
records = []
for row in rows:
    cells = [c.get_text(strip=True) for c in row.select('td')]
    records.append(dict(zip(headers, cells)))
print(records[0])

Handling colspan

A cell with colspan="2" visually spans two columns. If you ignore it, every following cell shifts left and misaligns with its header. Read the attribute and pad accordingly.

span = int(cell.get('colspan', 1))
values.extend([text] * span)

Handling rowspan

rowspan is harder: a cell carries down into rows below it. Track a buffer of pending values keyed by column index and inject them into subsequent rows until the span is exhausted.

pending = {}
for r_idx, row in enumerate(rows):
    col = 0
    for cell in row.select('td'):
        while col in pending and pending[col][1] > 0:
            col += 1
        rs = int(cell.get('rowspan', 1))
        if rs > 1:
            pending[col] = [cell.get_text(strip=True), rs]
        col += 1

Cleaning Extracted Values

Cells often contain currency symbols, thousands separators, or stray unicode. Normalize before storing:

  • Strip $, , and whitespace.
  • Cast numeric strings to numbers.
  • Replace non-breaking spaces.
def clean_price(text):
    text = text.replace('$', '').replace(',', '').strip()
    return float(text) if text else None

print(clean_price('$1,299.00'))

Pandas read_html Shortcut

For well-formed tables, pandas.read_html parses every table on a page into DataFrames in one call. Use it for quick wins, then fall back to manual parsing for tables with merged cells.

import pandas as pd
tables = pd.read_html(html)
df = tables[0]
print(df.head())

Detecting Layout vs Data Tables

Before extracting, confirm the table holds real data. Heuristics:

  • Has a <thead> or repeated <th> cells.
  • Multiple rows with consistent column counts.
  • No nested tables used purely for spacing.

Skip tables that fail these checks.

Putting It Together

A robust table extractor: locate the data table, read headers, walk rows while resolving spans, clean each value, and emit a list of dictionaries. This pipeline survives most real-world markup.

def extract_table(table):
    headers = [h.get_text(strip=True) for h in table.select('thead th')]
    out = []
    for row in table.select('tbody tr'):
        vals = [c.get_text(strip=True) for c in row.select('td')]
        out.append(dict(zip(headers, vals)))
    return out

Quick Check

Test your understanding of table parsing.

Recap

You learned to extract clean records from HTML tables: scope to tbody, map headers to values, resolve colspan and rowspan, clean cell text, and use pandas.read_html for simple cases.

With these skills you can turn even messy tabular markup into reliable structured data.

Preguntas frecuentes

¿La lección «Extracción de datos de tablas HTML» es gratis?

Sí — el texto completo de «Extracción de datos de tablas HTML» es gratis para leer aquí en la web. Para practicarla de forma interactiva (editor de código integrado y tutor de IA 24/7) y desbloquear el resto del curso de Web Scraping & Bots, actualiza a CoddyKit PRO. El curso de Web Scraping & Bots incluye 4 lecciones en total.

¿Qué aprenderé en «Extracción de datos de tablas HTML»?

Aprenda a analizar de forma fiable datos tabulares HTML, gestionar rowspan y colspan, y convertir tablas desordenadas en filas estructuradas y limpias. Practicas Web Scraping & Bots con código real que ejecutas directamente en el navegador, y un tutor de IA 24/7 responde tus preguntas mientras trabajas en la lección.

¿Necesito experiencia previa para empezar Web Scraping & Bots?

No se requiere experiencia previa. Web Scraping & Bots en CoddyKit está estructurado para principiantes hasta estudiantes avanzados, así que puedes empezar aquí o desde el inicio y avanzar a tu ritmo. Esta es la lección 4 de 4.

¿Cuánto tiempo toma la lección «Extracción de datos de tablas HTML»?

La mayoría de las lecciones de CoddyKit toman alrededor de 5–10 minutos. Cada una es compacta e interactiva, así que avanzas constantemente y retomas exactamente por donde dejaste en la web y la app.

¿Puedo escribir y ejecutar código en esta lección de Web Scraping & Bots?

Sí. Cada lección de Web Scraping & Bots incluye un editor de código integrado, así que escribes y ejecutas código real directamente en tu navegador y obtienes retroalimentación instantánea de IA — sin configuración local necesaria.

Todas las lecciones de este curso

  1. Navegación por estructuras HTML complejas
  2. Selectores CSS para mayor precisión
  3. XPath para una selección sólida
  4. Extracción de datos de tablas HTML
← Volver a Web Scraping & Bots