Извлечение данных из HTML-таблиц
Научитесь надёжно разбирать табличные данные HTML, обрабатывать rowspan и colspan и превращать неаккуратные таблицы в чистые структурированные строки.
«Извлечение данных из HTML-таблиц» — бесплатный урок Web Scraping & Bots на CoddyKit. Это урок 4 из 4. Ты можешь прочитать весь урок бесплатно ниже — а потом практиковать его прямо в браузере с встроенным редактором кода и ИИ-репетитором 24/7. Это часть пути обучения Web Scraping & Bots, и твой прогресс синхронизируется между веб-версией и приложением CoddyKit. Курс Web Scraping & Bots содержит 4 уроков всего.
Части этого урока еще не переведены и отображаются на английском.
Why Tables Are Tricky
HTML <table> elements look simple but are one of the most error-prone targets in scraping. Rows can merge cells, headers can repeat, and layout tables masquerade as data tables.
- Data tables hold real records you want.
- Layout tables only control visual structure.
This lesson focuses on extracting clean rows from genuine data tables.
Anatomy of a Table
A table is built from a few key tags:
<thead>/<tbody>group header and body rows.<tr>is a single row.<th>is a header cell,<td>is a data cell.
Knowing these landmarks lets you target rows precisely instead of grabbing raw text.
<table>
<thead><tr><th>Name</th><th>Price</th></tr></thead>
<tbody>
<tr><td>Widget</td><td>$9.99</td></tr>
<tr><td>Gadget</td><td>$14.50</td></tr>
</tbody>
</table>Selecting Rows
With a parser like BeautifulSoup you first find the table, then iterate its rows. Always scope your selection to tbody when present so the header row does not contaminate your data.
from bs4 import BeautifulSoup
soup = BeautifulSoup(html, 'html.parser')
table = soup.select_one('table')
rows = table.select('tbody tr')
print(len(rows), 'data rows found')Reading Cell Values
For each row, collect the cell text. Use get_text(strip=True) to drop surrounding whitespace and nested tag noise.
for row in rows:
cells = [c.get_text(strip=True) for c in row.select('td')]
print(cells)Mapping Headers to Values
Raw lists of cells are fragile. Pair each value with its column header so your output is self-describing and column-order changes do not break downstream code.
headers = [h.get_text(strip=True) for h in table.select('thead th')]
records = []
for row in rows:
cells = [c.get_text(strip=True) for c in row.select('td')]
records.append(dict(zip(headers, cells)))
print(records[0])Handling colspan
A cell with colspan="2" visually spans two columns. If you ignore it, every following cell shifts left and misaligns with its header. Read the attribute and pad accordingly.
span = int(cell.get('colspan', 1))
values.extend([text] * span)Handling rowspan
rowspan is harder: a cell carries down into rows below it. Track a buffer of pending values keyed by column index and inject them into subsequent rows until the span is exhausted.
pending = {}
for r_idx, row in enumerate(rows):
col = 0
for cell in row.select('td'):
while col in pending and pending[col][1] > 0:
col += 1
rs = int(cell.get('rowspan', 1))
if rs > 1:
pending[col] = [cell.get_text(strip=True), rs]
col += 1Cleaning Extracted Values
Cells often contain currency symbols, thousands separators, or stray unicode. Normalize before storing:
- Strip
$,,and whitespace. - Cast numeric strings to numbers.
- Replace non-breaking spaces.
def clean_price(text):
text = text.replace('$', '').replace(',', '').strip()
return float(text) if text else None
print(clean_price('$1,299.00'))Pandas read_html Shortcut
For well-formed tables, pandas.read_html parses every table on a page into DataFrames in one call. Use it for quick wins, then fall back to manual parsing for tables with merged cells.
import pandas as pd
tables = pd.read_html(html)
df = tables[0]
print(df.head())Detecting Layout vs Data Tables
Before extracting, confirm the table holds real data. Heuristics:
- Has a
<thead>or repeated<th>cells. - Multiple rows with consistent column counts.
- No nested tables used purely for spacing.
Skip tables that fail these checks.
Putting It Together
A robust table extractor: locate the data table, read headers, walk rows while resolving spans, clean each value, and emit a list of dictionaries. This pipeline survives most real-world markup.
def extract_table(table):
headers = [h.get_text(strip=True) for h in table.select('thead th')]
out = []
for row in table.select('tbody tr'):
vals = [c.get_text(strip=True) for c in row.select('td')]
out.append(dict(zip(headers, vals)))
return outQuick Check
Test your understanding of table parsing.
Recap
You learned to extract clean records from HTML tables: scope to tbody, map headers to values, resolve colspan and rowspan, clean cell text, and use pandas.read_html for simple cases.
With these skills you can turn even messy tabular markup into reliable structured data.
Часто задаваемые вопросы
Урок «Извлечение данных из HTML-таблиц» бесплатный?
Да — полный текст урока «Извлечение данных из HTML-таблиц» бесплатно доступен здесь в веб-версии. Чтобы практиковать его интерактивно (встроенный редактор кода и ИИ-репетитор 24/7) и разблокировать остальной курс Web Scraping & Bots, подпишись на CoddyKit PRO. Курс Web Scraping & Bots содержит 4 уроков всего.
Чему я научусь в уроке «Извлечение данных из HTML-таблиц»?
Научитесь надёжно разбирать табличные данные HTML, обрабатывать rowspan и colspan и превращать неаккуратные таблицы в чистые структурированные строки. Ты практикуешь Web Scraping & Bots с помощью реального кода, который запускаешь прямо в браузере, и ИИ-репетитор 24/7 отвечает на твои вопросы во время урока.
Нужен ли мне опыт, чтобы начать Web Scraping & Bots?
Предыдущий опыт не требуется. Web Scraping & Bots на CoddyKit структурирован для всех уровней — от новичков до продвинутых, поэтому ты можешь начать отсюда или с самого начала и учиться в своем темпе. Это урок 4 из 4.
Сколько времени занимает урок «Извлечение данных из HTML-таблиц»?
Большинство уроков CoddyKit занимают около 5–10 минут. Каждый из них компактный и интерактивный, поэтому ты постоянно делаешь прогресс и продолжаешь с того же места в веб-версии и приложении.
Можно ли писать и запускать код в этом уроке Web Scraping & Bots?
Да. Каждый урок Web Scraping & Bots включает встроенный редактор кода, поэтому ты пишешь и запускаешь реальный код прямо в браузере и получаешь моментальную обратную связь от AI — локальная установка не требуется.
Все уроки этого курса
- Навигация по сложным структурам HTML
- Селекторы CSS для точного выбора
- XPath для надежного выбора
- Извлечение данных из HTML-таблиц