Mengekstrak Data dari Tabel HTML
Pelajari cara mengurai data tabular HTML secara andal, menangani rowspan dan colspan, serta mengubah tabel yang berantakan menjadi baris terstruktur yang rapi.
Mengekstrak Data dari Tabel HTML adalah pelajaran Web Scraping & Bots gratis di CoddyKit. Ini adalah pelajaran 4 dari 4. Kamu bisa membaca pelajaran lengkapnya di bawah secara gratis — lalu praktikkan langsung di browser dengan editor kode bawaan dan tutor AI 24/7. Ini adalah bagian dari jalur belajar Web Scraping & Bots, dan progresmu tersinkronisasi di web dan aplikasi CoddyKit. Kursus Web Scraping & Bots mencakup 4 pelajaran total.
Bagian dari pelajaran ini belum diterjemahkan dan ditampilkan dalam bahasa Inggris.
Why Tables Are Tricky
HTML <table> elements look simple but are one of the most error-prone targets in scraping. Rows can merge cells, headers can repeat, and layout tables masquerade as data tables.
- Data tables hold real records you want.
- Layout tables only control visual structure.
This lesson focuses on extracting clean rows from genuine data tables.
Anatomy of a Table
A table is built from a few key tags:
<thead>/<tbody>group header and body rows.<tr>is a single row.<th>is a header cell,<td>is a data cell.
Knowing these landmarks lets you target rows precisely instead of grabbing raw text.
<table>
<thead><tr><th>Name</th><th>Price</th></tr></thead>
<tbody>
<tr><td>Widget</td><td>$9.99</td></tr>
<tr><td>Gadget</td><td>$14.50</td></tr>
</tbody>
</table>Selecting Rows
With a parser like BeautifulSoup you first find the table, then iterate its rows. Always scope your selection to tbody when present so the header row does not contaminate your data.
from bs4 import BeautifulSoup
soup = BeautifulSoup(html, 'html.parser')
table = soup.select_one('table')
rows = table.select('tbody tr')
print(len(rows), 'data rows found')Reading Cell Values
For each row, collect the cell text. Use get_text(strip=True) to drop surrounding whitespace and nested tag noise.
for row in rows:
cells = [c.get_text(strip=True) for c in row.select('td')]
print(cells)Mapping Headers to Values
Raw lists of cells are fragile. Pair each value with its column header so your output is self-describing and column-order changes do not break downstream code.
headers = [h.get_text(strip=True) for h in table.select('thead th')]
records = []
for row in rows:
cells = [c.get_text(strip=True) for c in row.select('td')]
records.append(dict(zip(headers, cells)))
print(records[0])Handling colspan
A cell with colspan="2" visually spans two columns. If you ignore it, every following cell shifts left and misaligns with its header. Read the attribute and pad accordingly.
span = int(cell.get('colspan', 1))
values.extend([text] * span)Handling rowspan
rowspan is harder: a cell carries down into rows below it. Track a buffer of pending values keyed by column index and inject them into subsequent rows until the span is exhausted.
pending = {}
for r_idx, row in enumerate(rows):
col = 0
for cell in row.select('td'):
while col in pending and pending[col][1] > 0:
col += 1
rs = int(cell.get('rowspan', 1))
if rs > 1:
pending[col] = [cell.get_text(strip=True), rs]
col += 1Cleaning Extracted Values
Cells often contain currency symbols, thousands separators, or stray unicode. Normalize before storing:
- Strip
$,,and whitespace. - Cast numeric strings to numbers.
- Replace non-breaking spaces.
def clean_price(text):
text = text.replace('$', '').replace(',', '').strip()
return float(text) if text else None
print(clean_price('$1,299.00'))Pandas read_html Shortcut
For well-formed tables, pandas.read_html parses every table on a page into DataFrames in one call. Use it for quick wins, then fall back to manual parsing for tables with merged cells.
import pandas as pd
tables = pd.read_html(html)
df = tables[0]
print(df.head())Detecting Layout vs Data Tables
Before extracting, confirm the table holds real data. Heuristics:
- Has a
<thead>or repeated<th>cells. - Multiple rows with consistent column counts.
- No nested tables used purely for spacing.
Skip tables that fail these checks.
Putting It Together
A robust table extractor: locate the data table, read headers, walk rows while resolving spans, clean each value, and emit a list of dictionaries. This pipeline survives most real-world markup.
def extract_table(table):
headers = [h.get_text(strip=True) for h in table.select('thead th')]
out = []
for row in table.select('tbody tr'):
vals = [c.get_text(strip=True) for c in row.select('td')]
out.append(dict(zip(headers, vals)))
return outQuick Check
Test your understanding of table parsing.
Recap
You learned to extract clean records from HTML tables: scope to tbody, map headers to values, resolve colspan and rowspan, clean cell text, and use pandas.read_html for simple cases.
With these skills you can turn even messy tabular markup into reliable structured data.
Pertanyaan yang Sering Diajukan
Apakah pelajaran “Mengekstrak Data dari Tabel HTML” gratis?
Ya — teks lengkap “Mengekstrak Data dari Tabel HTML” gratis dibaca di sini di web. Untuk praktiknya secara interaktif (editor kode bawaan dan tutor AI 24/7) dan buka sisa kursus Web Scraping & Bots, upgrade ke CoddyKit PRO. Kursus Web Scraping & Bots mencakup 4 pelajaran total.
Apa yang akan aku pelajari di “Mengekstrak Data dari Tabel HTML”?
Pelajari cara mengurai data tabular HTML secara andal, menangani rowspan dan colspan, serta mengubah tabel yang berantakan menjadi baris terstruktur yang rapi. Kamu berlatih Web Scraping & Bots dengan kode praktik yang langsung kamu jalankan di browser, dan tutor AI 24/7 menjawab pertanyaanmu saat kamu mengerjakan pelajaran ini.
Apakah aku perlu pengalaman untuk memulai Web Scraping & Bots?
Tidak diperlukan pengalaman sebelumnya. Web Scraping & Bots di CoddyKit dirancang untuk pemula hingga pelajar tingkat lanjut, jadi kamu bisa memulai di sini atau dari awal dan belajar sesuai kecepatan kamu sendiri. Ini adalah pelajaran 4 dari 4.
Berapa lama pelajaran “Mengekstrak Data dari Tabel HTML” memakan waktu?
Sebagian besar pelajaran CoddyKit memakan waktu sekitar 5–10 menit. Setiap pelajaran ringkas dan interaktif, jadi kamu membuat kemajuan stabil dan melanjutkan dari tempat kamu tinggalkan di web dan aplikasi.
Bisakah aku menulis dan menjalankan kode dalam pelajaran Web Scraping & Bots ini?
Ya. Setiap pelajaran Web Scraping & Bots menyertakan editor kode bawaan, jadi kamu menulis dan menjalankan kode nyata langsung di browser dan mendapatkan umpan balik AI instan — tidak diperlukan penyiapan lokal.
Semua pelajaran dalam kursus ini
- Menavigasi Struktur HTML yang Kompleks
- Selector CSS untuk Presisi
- XPath untuk Seleksi Andal
- Mengekstrak Data dari Tabel HTML