0Pricing
Web Scraping & Bots · レッスン

HTML テーブルからデータを抽出する

表形式の HTML データを確実に解析し、rowspan と colspan を扱い、乱雑なテーブルを整った構造化行に変換する方法を学びます。

「HTML テーブルからデータを抽出する」はCoddyKit上の無料Web Scraping & Botsレッスンです。 これはレッスン4/4です。 下記で完全なレッスンを無料で読むことができます。その後、ブラウザ内の組み込みコードエディタと24時間対応のAIチューターでハンズオン演習できます。 これはWeb Scraping & Bots学習パスの一部であり、ウェブとCoddyKitアプリ全体で進捗が同期されます。 Web Scraping & Botsコースには全4レッスンが含まれています。

このレッスンの一部はまだ翻訳されておらず、英語で表示されています。

Why Tables Are Tricky

HTML <table> elements look simple but are one of the most error-prone targets in scraping. Rows can merge cells, headers can repeat, and layout tables masquerade as data tables.

  • Data tables hold real records you want.
  • Layout tables only control visual structure.

This lesson focuses on extracting clean rows from genuine data tables.

Anatomy of a Table

A table is built from a few key tags:

  • <thead> / <tbody> group header and body rows.
  • <tr> is a single row.
  • <th> is a header cell, <td> is a data cell.

Knowing these landmarks lets you target rows precisely instead of grabbing raw text.

<table>
  <thead><tr><th>Name</th><th>Price</th></tr></thead>
  <tbody>
    <tr><td>Widget</td><td>$9.99</td></tr>
    <tr><td>Gadget</td><td>$14.50</td></tr>
  </tbody>
</table>

Selecting Rows

With a parser like BeautifulSoup you first find the table, then iterate its rows. Always scope your selection to tbody when present so the header row does not contaminate your data.

from bs4 import BeautifulSoup

soup = BeautifulSoup(html, 'html.parser')
table = soup.select_one('table')
rows = table.select('tbody tr')
print(len(rows), 'data rows found')

Reading Cell Values

For each row, collect the cell text. Use get_text(strip=True) to drop surrounding whitespace and nested tag noise.

for row in rows:
    cells = [c.get_text(strip=True) for c in row.select('td')]
    print(cells)

Mapping Headers to Values

Raw lists of cells are fragile. Pair each value with its column header so your output is self-describing and column-order changes do not break downstream code.

headers = [h.get_text(strip=True) for h in table.select('thead th')]
records = []
for row in rows:
    cells = [c.get_text(strip=True) for c in row.select('td')]
    records.append(dict(zip(headers, cells)))
print(records[0])

Handling colspan

A cell with colspan="2" visually spans two columns. If you ignore it, every following cell shifts left and misaligns with its header. Read the attribute and pad accordingly.

span = int(cell.get('colspan', 1))
values.extend([text] * span)

Handling rowspan

rowspan is harder: a cell carries down into rows below it. Track a buffer of pending values keyed by column index and inject them into subsequent rows until the span is exhausted.

pending = {}
for r_idx, row in enumerate(rows):
    col = 0
    for cell in row.select('td'):
        while col in pending and pending[col][1] > 0:
            col += 1
        rs = int(cell.get('rowspan', 1))
        if rs > 1:
            pending[col] = [cell.get_text(strip=True), rs]
        col += 1

Cleaning Extracted Values

Cells often contain currency symbols, thousands separators, or stray unicode. Normalize before storing:

  • Strip $, , and whitespace.
  • Cast numeric strings to numbers.
  • Replace non-breaking spaces.
def clean_price(text):
    text = text.replace('$', '').replace(',', '').strip()
    return float(text) if text else None

print(clean_price('$1,299.00'))

Pandas read_html Shortcut

For well-formed tables, pandas.read_html parses every table on a page into DataFrames in one call. Use it for quick wins, then fall back to manual parsing for tables with merged cells.

import pandas as pd
tables = pd.read_html(html)
df = tables[0]
print(df.head())

Detecting Layout vs Data Tables

Before extracting, confirm the table holds real data. Heuristics:

  • Has a <thead> or repeated <th> cells.
  • Multiple rows with consistent column counts.
  • No nested tables used purely for spacing.

Skip tables that fail these checks.

Putting It Together

A robust table extractor: locate the data table, read headers, walk rows while resolving spans, clean each value, and emit a list of dictionaries. This pipeline survives most real-world markup.

def extract_table(table):
    headers = [h.get_text(strip=True) for h in table.select('thead th')]
    out = []
    for row in table.select('tbody tr'):
        vals = [c.get_text(strip=True) for c in row.select('td')]
        out.append(dict(zip(headers, vals)))
    return out

Quick Check

Test your understanding of table parsing.

Recap

You learned to extract clean records from HTML tables: scope to tbody, map headers to values, resolve colspan and rowspan, clean cell text, and use pandas.read_html for simple cases.

With these skills you can turn even messy tabular markup into reliable structured data.

よくある質問

「HTML テーブルからデータを抽出する」レッスンは無料ですか?

はい。「HTML テーブルからデータを抽出する」の完全なテキストはこのウェブで無料で読めます。インタラクティブに演習し(組み込みコードエディタと24時間対応のAIチューター)、Web Scraping & Botsコースの残りをアンロックするには、CoddyKit PROにアップグレードしてください。 Web Scraping & Botsコースには全4レッスンが含まれています。

「HTML テーブルからデータを抽出する」で何を学びますか?

表形式の HTML データを確実に解析し、rowspan と colspan を扱い、乱雑なテーブルを整った構造化行に変換する方法を学びます。 ブラウザで直接実行するハンズオンコードでWeb Scraping & Botsを演習し、24時間対応のAIチューターがレッスンを進める中での質問に答えます。

Web Scraping & Botsを始めるのに経験は必要ですか?

事前経験は必要ありません。CoddyKitのWeb Scraping & Botsは初級者から上級者向けに構成されているため、ここから始めるか最初から始めて、自分のペースで進むことができます。 これはレッスン4/4です。

「HTML テーブルからデータを抽出する」レッスンにはどのくらい時間がかかりますか?

ほとんどのCoddyKitレッスンは約5~10分かかります。各レッスンはコンパクトでインタラクティブなので、着実に進歩し、ウェブとアプリ全体で正確に前回の場所から再開できます。

このWeb Scraping & Botsレッスンでコードを書いて実行できますか?

はい。すべてのWeb Scraping & Botsレッスンに組み込みコードエディタが含まれているため、ブラウザでリアルコードを書いて実行し、即座のAIフィードバックを取得できます。ローカル設定は不要です。

このコースのすべてのレッスン

  1. 複雑なHTML構造のナビゲーション
  2. 精密な抽出のためのCSSセレクター
  3. 堅牢な選択のためのXPath
  4. HTML テーブルからデータを抽出する
← Web Scraping & Botsに戻る