การดึงข้อมูลจากตาราง HTML
เรียนรู้การแยกวิเคราะห์ข้อมูลแบบตารางใน HTML อย่างน่าเชื่อถือ จัดการ rowspan และ colspan และแปลงตารางยุ่งเหยิงให้เป็นแถวที่มีโครงสร้างสะอาด
การดึงข้อมูลจากตาราง HTML เป็นบทเรียน Web Scraping & Bots ฟรีบน CoddyKit นี่คือบทเรียนที่ 4 จากทั้งหมด 4 บทเรียน คุณสามารถอ่านบทเรียนทั้งหมดด้านล่างฟรี — จากนั้นลองปฏิบัติด้วยตัวคุณเองในเบราว์เซอร์พร้อมตัวแก้ไขโค้ดในตัวและติวเตอร์ AI ตลอด 24/7 บทเรียนนี้เป็นส่วนหนึ่งของเส้นทางการเรียน Web Scraping & Bots และความก้าวหน้าของคุณจะซิงค์ข้ามเว็บและแอป CoddyKit คอร์ส Web Scraping & Bots มีบทเรียนทั้งหมด 4 บทเรียน
บางส่วนของบทเรียนนี้ยังไม่ได้รับการแปล และแสดงเป็นภาษาอังกฤษ
Why Tables Are Tricky
HTML <table> elements look simple but are one of the most error-prone targets in scraping. Rows can merge cells, headers can repeat, and layout tables masquerade as data tables.
- Data tables hold real records you want.
- Layout tables only control visual structure.
This lesson focuses on extracting clean rows from genuine data tables.
Anatomy of a Table
A table is built from a few key tags:
<thead>/<tbody>group header and body rows.<tr>is a single row.<th>is a header cell,<td>is a data cell.
Knowing these landmarks lets you target rows precisely instead of grabbing raw text.
<table>
<thead><tr><th>Name</th><th>Price</th></tr></thead>
<tbody>
<tr><td>Widget</td><td>$9.99</td></tr>
<tr><td>Gadget</td><td>$14.50</td></tr>
</tbody>
</table>Selecting Rows
With a parser like BeautifulSoup you first find the table, then iterate its rows. Always scope your selection to tbody when present so the header row does not contaminate your data.
from bs4 import BeautifulSoup
soup = BeautifulSoup(html, 'html.parser')
table = soup.select_one('table')
rows = table.select('tbody tr')
print(len(rows), 'data rows found')Reading Cell Values
For each row, collect the cell text. Use get_text(strip=True) to drop surrounding whitespace and nested tag noise.
for row in rows:
cells = [c.get_text(strip=True) for c in row.select('td')]
print(cells)Mapping Headers to Values
Raw lists of cells are fragile. Pair each value with its column header so your output is self-describing and column-order changes do not break downstream code.
headers = [h.get_text(strip=True) for h in table.select('thead th')]
records = []
for row in rows:
cells = [c.get_text(strip=True) for c in row.select('td')]
records.append(dict(zip(headers, cells)))
print(records[0])Handling colspan
A cell with colspan="2" visually spans two columns. If you ignore it, every following cell shifts left and misaligns with its header. Read the attribute and pad accordingly.
span = int(cell.get('colspan', 1))
values.extend([text] * span)Handling rowspan
rowspan is harder: a cell carries down into rows below it. Track a buffer of pending values keyed by column index and inject them into subsequent rows until the span is exhausted.
pending = {}
for r_idx, row in enumerate(rows):
col = 0
for cell in row.select('td'):
while col in pending and pending[col][1] > 0:
col += 1
rs = int(cell.get('rowspan', 1))
if rs > 1:
pending[col] = [cell.get_text(strip=True), rs]
col += 1Cleaning Extracted Values
Cells often contain currency symbols, thousands separators, or stray unicode. Normalize before storing:
- Strip
$,,and whitespace. - Cast numeric strings to numbers.
- Replace non-breaking spaces.
def clean_price(text):
text = text.replace('$', '').replace(',', '').strip()
return float(text) if text else None
print(clean_price('$1,299.00'))Pandas read_html Shortcut
For well-formed tables, pandas.read_html parses every table on a page into DataFrames in one call. Use it for quick wins, then fall back to manual parsing for tables with merged cells.
import pandas as pd
tables = pd.read_html(html)
df = tables[0]
print(df.head())Detecting Layout vs Data Tables
Before extracting, confirm the table holds real data. Heuristics:
- Has a
<thead>or repeated<th>cells. - Multiple rows with consistent column counts.
- No nested tables used purely for spacing.
Skip tables that fail these checks.
Putting It Together
A robust table extractor: locate the data table, read headers, walk rows while resolving spans, clean each value, and emit a list of dictionaries. This pipeline survives most real-world markup.
def extract_table(table):
headers = [h.get_text(strip=True) for h in table.select('thead th')]
out = []
for row in table.select('tbody tr'):
vals = [c.get_text(strip=True) for c in row.select('td')]
out.append(dict(zip(headers, vals)))
return outQuick Check
Test your understanding of table parsing.
Recap
You learned to extract clean records from HTML tables: scope to tbody, map headers to values, resolve colspan and rowspan, clean cell text, and use pandas.read_html for simple cases.
With these skills you can turn even messy tabular markup into reliable structured data.
เรียนรู้ Python ด้วย AI tutor — ฟรี
เขียนและเรียกใช้โค้ดจริงในเบราว์เซอร์ของคุณ รับความช่วยเหลือทันทีจาก AI tutor 24/7 และเรียนรู้ต่อจากที่คุณหยุดบนเว็บหรือในแอป
- คอร์ส
- 12
- บทเรียน
- 48
คำถามที่พบบ่อย
บทเรียน “การดึงข้อมูลจากตาราง HTML” ฟรีหรือไม่
ใช่ — ข้อความเต็มของ “การดึงข้อมูลจากตาราง HTML” ฟรีให้อ่านที่นี่บนเว็บ เพื่อปฏิบัติแบบโต้ตอบ (ตัวแก้ไขโค้ดในตัวและติวเตอร์ AI ตลอด 24/7) และปลดล็อคส่วนที่เหลือของคอร์ส Web Scraping & Bots ให้อัปเกรดเป็น CoddyKit PRO คอร์ส Web Scraping & Bots มีบทเรียนทั้งหมด 4 บทเรียน
คุณจะเรียนรู้อะไรในบทเรียน “การดึงข้อมูลจากตาราง HTML”
เรียนรู้การแยกวิเคราะห์ข้อมูลแบบตารางใน HTML อย่างน่าเชื่อถือ จัดการ rowspan และ colspan และแปลงตารางยุ่งเหยิงให้เป็นแถวที่มีโครงสร้างสะอาด คุณปฏิบัติ Web Scraping & Bots ด้วยโค้ดที่ใช้งานได้จริงที่คุณเรียกใช้โดยตรงในเบราว์เซอร์ และติวเตอร์ AI ตลอด 24/7 ตอบคำถามของคุณขณะที่คุณไปผ่านบทเรียน
คุณต้องมีประสบการณ์ก่อนที่จะเริ่มเรียน Web Scraping & Bots หรือไม่
ไม่จำเป็นต้องมีประสบการณ์มาก่อน Web Scraping & Bots บน CoddyKit ออกแบบมาสำหรับผู้เริ่มต้นไปจนถึงผู้เรียนขั้นสูง คุณสามารถเริ่มต้นที่นี่หรือเริ่มจากตัวแรกและเรียนด้วยความเร็วของคุณเอง นี่คือบทเรียนที่ 4 จากทั้งหมด 4 บทเรียน
บทเรียน “การดึงข้อมูลจากตาราง HTML” ใช้เวลานานแค่ไหน
บทเรียน CoddyKit ส่วนใหญ่ใช้เวลาประมาณ 5–10 นาที แต่ละบทเรียนจึงสั้นและเป็นแบบโต้ตอบ คุณสามารถก้าวหน้าอย่างต่อเนื่องและกลับมาเรียนต่อจากตรงที่เพิ่งหยุดบนเว็บและแอปได้เลย
ฉันเขียนและรันโค้ดในบทเรียน Web Scraping & Bots นี้ได้ไหม
ได้ บทเรียน Web Scraping & Bots ทุกบทมีตัวแก้ไขโค้ดในตัว คุณจึงเขียนและรันโค้ดจริงได้เลยในเบราว์เซอร์ และได้รับข้อเสนอแนะจาก AI ในทันที — ไม่ต้องติดตั้งในเครื่องของคุณ
บทเรียนทั้งหมดในหลักสูตรนี้
- การนำทางโครงสร้าง HTML ที่ซับซ้อน
- ตัวเลือก CSS เพื่อความแม่นยำ
- XPath สำหรับการเลือกที่มีประสิทธิภาพ
- การดึงข้อมูลจากตาราง HTML