Web Scraping & Bots · Lekcja

Wyodrębnianie danych z tabel HTML

Dowiedz się, jak niezawodnie parsować tabelaryczne dane HTML, obsługiwać rowspan i colspan oraz przekształcać nieuporządkowane tabele w czyste wiersze strukturalne.

Lekcja 4 z 413 kroki

Wyodrębnianie danych z tabel HTML to bezpłatna lekcja Web Scraping & Bots na CoddyKit. To lekcja 4 z 4. Możesz przeczytać całą lekcję poniżej za darmo — a potem ćwiczyć ją interaktywnie w przeglądarce z wbudowanym edytorem kodu i tutorem AI dostępnym 24/7. To część ścieżki edukacyjnej Web Scraping & Bots, a Twój postęp synchronizuje się między webem a aplikacją CoddyKit. Kurs Web Scraping & Bots zawiera 4 lekcji w sumie.

Części tej lekcji nie zostały jeszcze przetłumaczone i są wyświetlane po angielsku.

Why Tables Are Tricky

HTML <table> elements look simple but are one of the most error-prone targets in scraping. Rows can merge cells, headers can repeat, and layout tables masquerade as data tables.

  • Data tables hold real records you want.
  • Layout tables only control visual structure.

This lesson focuses on extracting clean rows from genuine data tables.

Anatomy of a Table

A table is built from a few key tags:

  • <thead> / <tbody> group header and body rows.
  • <tr> is a single row.
  • <th> is a header cell, <td> is a data cell.

Knowing these landmarks lets you target rows precisely instead of grabbing raw text.

<table>
  <thead><tr><th>Name</th><th>Price</th></tr></thead>
  <tbody>
    <tr><td>Widget</td><td>$9.99</td></tr>
    <tr><td>Gadget</td><td>$14.50</td></tr>
  </tbody>
</table>

Selecting Rows

With a parser like BeautifulSoup you first find the table, then iterate its rows. Always scope your selection to tbody when present so the header row does not contaminate your data.

from bs4 import BeautifulSoup

soup = BeautifulSoup(html, 'html.parser')
table = soup.select_one('table')
rows = table.select('tbody tr')
print(len(rows), 'data rows found')

Reading Cell Values

For each row, collect the cell text. Use get_text(strip=True) to drop surrounding whitespace and nested tag noise.

for row in rows:
    cells = [c.get_text(strip=True) for c in row.select('td')]
    print(cells)

Mapping Headers to Values

Raw lists of cells are fragile. Pair each value with its column header so your output is self-describing and column-order changes do not break downstream code.

headers = [h.get_text(strip=True) for h in table.select('thead th')]
records = []
for row in rows:
    cells = [c.get_text(strip=True) for c in row.select('td')]
    records.append(dict(zip(headers, cells)))
print(records[0])

Handling colspan

A cell with colspan="2" visually spans two columns. If you ignore it, every following cell shifts left and misaligns with its header. Read the attribute and pad accordingly.

span = int(cell.get('colspan', 1))
values.extend([text] * span)

Handling rowspan

rowspan is harder: a cell carries down into rows below it. Track a buffer of pending values keyed by column index and inject them into subsequent rows until the span is exhausted.

pending = {}
for r_idx, row in enumerate(rows):
    col = 0
    for cell in row.select('td'):
        while col in pending and pending[col][1] > 0:
            col += 1
        rs = int(cell.get('rowspan', 1))
        if rs > 1:
            pending[col] = [cell.get_text(strip=True), rs]
        col += 1

Cleaning Extracted Values

Cells often contain currency symbols, thousands separators, or stray unicode. Normalize before storing:

  • Strip $, , and whitespace.
  • Cast numeric strings to numbers.
  • Replace non-breaking spaces.
def clean_price(text):
    text = text.replace('$', '').replace(',', '').strip()
    return float(text) if text else None

print(clean_price('$1,299.00'))

Pandas read_html Shortcut

For well-formed tables, pandas.read_html parses every table on a page into DataFrames in one call. Use it for quick wins, then fall back to manual parsing for tables with merged cells.

import pandas as pd
tables = pd.read_html(html)
df = tables[0]
print(df.head())

Detecting Layout vs Data Tables

Before extracting, confirm the table holds real data. Heuristics:

  • Has a <thead> or repeated <th> cells.
  • Multiple rows with consistent column counts.
  • No nested tables used purely for spacing.

Skip tables that fail these checks.

Putting It Together

A robust table extractor: locate the data table, read headers, walk rows while resolving spans, clean each value, and emit a list of dictionaries. This pipeline survives most real-world markup.

def extract_table(table):
    headers = [h.get_text(strip=True) for h in table.select('thead th')]
    out = []
    for row in table.select('tbody tr'):
        vals = [c.get_text(strip=True) for c in row.select('td')]
        out.append(dict(zip(headers, vals)))
    return out

Quick Check

Test your understanding of table parsing.

Recap

You learned to extract clean records from HTML tables: scope to tbody, map headers to values, resolve colspan and rowspan, clean cell text, and use pandas.read_html for simple cases.

With these skills you can turn even messy tabular markup into reliable structured data.

Bezpłatny start

Ucz się Python dzięki korepetycjom AI — za darmo

Pisz i uruchamiaj kod w przeglądarce, otrzymuj natychmiastową pomoc od korepetytora AI dostępnego 24/7 i kontynuuj naukę w sieci lub w aplikacji.

Kursy
12
Lekcje
48

Często zadawane pytania

Czy lekcja „Wyodrębnianie danych z tabel HTML” jest bezpłatna?

Tak — pełny tekst „Wyodrębnianie danych z tabel HTML” jest dostępny za darmo tutaj w sieci. Aby ćwiczyć ją interaktywnie (wbudowany edytor kodu i tutor AI dostępny 24/7) i odblokować resztę kursu Web Scraping & Bots, przejdź na CoddyKit PRO. Kurs Web Scraping & Bots zawiera 4 lekcji w sumie.

Co nauczysz się w „Wyodrębnianie danych z tabel HTML”?

Dowiedz się, jak niezawodnie parsować tabelaryczne dane HTML, obsługiwać rowspan i colspan oraz przekształcać nieuporządkowane tabele w czyste wiersze strukturalne. Ćwiczysz Web Scraping & Bots z praktycznym kodem, który uruchamiasz bezpośrednio w przeglądarce, a tutor AI dostępny 24/7 odpowiada na Twoje pytania podczas pracy nad lekcją.

Czy potrzebuję doświadczenia, aby zacząć Web Scraping & Bots?

Nie wymagamy żadnego doświadczenia. Web Scraping & Bots w CoddyKit jest strukturyzowany dla początkujących i zaawansowanych użytkowników, więc możesz zacząć tutaj lub od początku i uczyć się w swoim tempie. To lekcja 4 z 4.

Ile czasu zajmuje lekcja „Wyodrębnianie danych z tabel HTML”?

Większość lekcji CoddyKit trwa około 5–10 minut. Każda lekcja to mały, interaktywny krok, dzięki czemu robisz systematyczne postępy i zawsze wracasz dokładnie do tego samego miejsca — na webie i w aplikacji.

Czy mogę pisać i uruchamiać kod w tej lekcji Web Scraping & Bots?

Tak. Każda lekcja Web Scraping & Bots zawiera wbudowany edytor kodu, więc piszesz i uruchamiasz prawdziwy kod bezpośrednio w przeglądarce i od razu otrzymujesz sprzężenie zwrotne od AI — bez konfiguracji na komputerze.

Wszystkie lekcje w tym kursie

  1. Poruszanie się po złożonych strukturach HTML
  2. Selektory CSS zapewniające precyzję
  3. XPath do niezawodnego wyboru elementów
  4. Wyodrębnianie danych z tabel HTML
← Powrót do Web Scraping & Bots