정밀한 CSS 선택자
CSS 선택자를 적용해 스타일과 속성을 기준으로 요소를 정확히 찾아 데이터를 추출합니다.
정밀한 CSS 선택자은(는) CoddyKit의 무료 Web Scraping & Bots 강의입니다. 이것은 4개 중 2번째 강의입니다. 아래에서 전체 강의를 무료로 읽을 수 있으며, 내장 코드 에디터와 24/7 AI 튜터와 함께 브라우저에서 직접 실습할 수 있습니다. 이 강의는 Web Scraping & Bots 학습 경로의 일부이며, 진행 상황이 웹과 CoddyKit 앱에 동기화됩니다. Web Scraping & Bots 강의에는 총 4개의 강의가 포함되어 있습니다.
이 강의의 일부는 아직 번역되지 않았으며 영어로 표시됩니다.
What are CSS Selectors?
CSS selectors are patterns used to select elements on a web page. Think of them as precise instructions for finding specific pieces of information.
Web browsers use them to apply styles (CSS), and we'll use them to extract data efficiently.
Selectors for Data Extraction
In web scraping, CSS selectors provide a powerful way to pinpoint exactly the data you need from complex HTML.
- They are often more concise than XPath.
- Many developers are already familiar with CSS.
- BeautifulSoup has excellent support for them.
Selecting by Tag Name
The simplest selector is the tag name. This selects all elements of that type.
For example, p selects all paragraph tags, and a selects all anchor (link) tags.
Example: To find all list items, you'd use li.
from bs4 import BeautifulSoup
html_doc = """
<html><body>
<h1>My Title</h1>
<p>First paragraph.</p>
<ul>
<li>Item 1</li>
<li>Item 2</li>
</ul>
</body></html>
"""
soup = BeautifulSoup(html_doc, 'html.parser')
# Select all 'p' tags
paragraphs = soup.select('p')
for p in paragraphs:
print(p.get_text())
Class and ID Selectors
You can select elements based on their class or ID attributes. These are very common for styling and unique identification.
- Class: Use a dot
.before the class name (e.g.,.product-title). - ID: Use a hash
#before the ID name (e.g.,#main-content). IDs should be unique!
from bs4 import BeautifulSoup
html_doc = """
<div id="header">Welcome</div>
<p class="intro">Hello there!</p>
<p class="intro">Another intro.</p>
"""
soup = BeautifulSoup(html_doc, 'html.parser')
header = soup.select_one('#header')
print("Header:", header.get_text())
intros = soup.select('.intro')
for p in intros:
print("Intro:", p.get_text())
Selecting Nested Elements
To select elements that are inside other elements, you use a space between selectors. This is called a descendant selector.
It means "find an element (B) that is anywhere inside another element (A)".
Example: div p selects all <p> tags that are inside any <div> tag.
from bs4 import BeautifulSoup
html_doc = """
<div>
<p>Inside div paragraph 1</p>
<span>
<p>Inside span inside div</p>
</span>
</div>
<p>Outside div paragraph</p>
"""
soup = BeautifulSoup(html_doc, 'html.parser')
div_paragraphs = soup.select('div p')
for p in div_paragraphs:
print(p.get_text())
Direct Children Only
Sometimes you only want elements that are direct children of another element, not just any descendant.
Use the greater than symbol > for this.
Example: ul > li selects all <li> tags that are direct children of a <ul> tag.
from bs4 import BeautifulSoup
html_doc = """
<div class="container">
<p>Direct child P</p>
<div>
<p>Nested P (not direct)</p>
</div>
</div>
"""
soup = BeautifulSoup(html_doc, 'html.parser')
direct_p = soup.select('.container > p')
for p in direct_p:
print(p.get_text())
Selecting by Attributes
You can select elements based on their attributes and even their attribute values!
[attr]: Has the attribute (e.g.,[href]).[attr="value"]: Has attribute with exact value (e.g.,[target="_blank"]).[attr^="value"]: Attribute value starts with (e.g.,[src^="data:"]).[attr$="value"]: Attribute value ends with (e.g.,[alt$="logo"]).[attr*="value"]: Attribute value contains (e.g.,[id*="item"]).
from bs4 import BeautifulSoup
html_doc = """
<a href="/about">About Us</a>
<a href="https://example.com/contact" target="_blank">Contact</a>
<img src="image.jpg" alt="product image">
"""
soup = BeautifulSoup(html_doc, 'html.parser')
# Select links with target="_blank"
external_links = soup.select('a[target="_blank"]')
for link in external_links:
print("External:", link.get('href'))
# Select images with alt containing "image"
product_images = soup.select('img[alt*="image"]')
for img in product_images:
print("Image:", img.get('src'))
Combining with Commas
To select elements that match any of several different selectors, you can separate them with a comma ,.
This is useful when you want to gather data from different types of elements or locations.
Example: h1, h2, h3 selects all heading tags of level 1, 2, or 3.
from bs4 import BeautifulSoup
html_doc = """
<h1>Main Heading</h1>
<p>Some text.</p>
<h2>Sub Heading</h2>
<div>Another div.</div>
"""
soup = BeautifulSoup(html_doc, 'html.parser')
headings = soup.select('h1, h2')
for h in headings:
print(h.get_text())
Pseudo-classes for Position
CSS pseudo-classes allow selection based on state or position, not just attributes. For scraping, position-based ones are very useful.
:first-child: Selects the first child element.:last-child: Selects the last child element.:nth-of-type(n): Selects the Nth element of a specific type (e.g.,li:nth-of-type(2)for the second list item).
from bs4 import BeautifulSoup
html_doc = """
<ul>
<li>First item</li>
<li>Second item</li>
<li>Third item</li>
</ul>
"""
soup = BeautifulSoup(html_doc, 'html.parser')
first_item = soup.select_one('li:first-child')
print("First:", first_item.get_text())
second_item = soup.select_one('li:nth-of-type(2)')
print("Second:", second_item.get_text())
Practical CSS Selector Use
Let's combine what we've learned to extract specific data from a sample product listing.
We want the title and price of the first product.
from bs4 import BeautifulSoup
html_doc = """
<div class="product-list">
<div class="product-card">
<h3 class="product-title">Laptop X1</h3>
<p class="product-price">$999.99</p>
<button class="add-to-cart">Add</button>
</div>
<div class="product-card">
<h3 class="product-title">Mouse Z2</h3>
<p class="product-price">$29.99</p>
<button class="add-to-cart">Add</button>
</div>
</div>
"""
soup = BeautifulSoup(html_doc, 'html.parser')
# Select the first product card
first_product = soup.select_one('.product-card:first-of-type')
if first_product:
title = first_product.select_one('.product-title')
price = first_product.select_one('.product-price')
print("Title:", title.get_text())
print("Price:", price.get_text())
else:
print("No product found.")
Quick Check on Selectors
Given the HTML below, what CSS selector would correctly select the text "Product Name 2"?
<div class="items">
<div id="item-1">
<span class="name">Product Name 1</span>
</div>
<div id="item-2">
<span class="name">Product Name 2</span>
</div>
<p class="name">Other Name</p>
</div>Recap: CSS Selectors
Great job! You've mastered the basics of CSS selectors for web scraping.
- We learned to select by tag, class, and ID.
- We explored descendant (space) and direct child (
>) selectors. - You can filter by attributes (
[attr="value"]) and use pseudo-classes like:first-child. - BeautifulSoup's
.select()and.select_one()methods make using them easy in Python.
Next, we'll look into XPath for even more powerful selections!
자주 묻는 질문
“정밀한 CSS 선택자” 강의는 무료인가요?
네 — “정밀한 CSS 선택자” 전체 내용을 이 웹사이트에서 무료로 읽을 수 있습니다. 인터랙티브하게 실습하려면(내장 코드 에디터와 24/7 AI 튜터), CoddyKit PRO로 업그레이드하면 Web Scraping & Bots 강의 전체를 잠금 해제할 수 있습니다. Web Scraping & Bots 강의에는 총 4개의 강의가 포함되어 있습니다.
“정밀한 CSS 선택자”에서 뭘 배우나요?
CSS 선택자를 적용해 스타일과 속성을 기준으로 요소를 정확히 찾아 데이터를 추출합니다. 브라우저에서 직접 실행하는 실습 코드로 Web Scraping & Bots을(를) 배우며, 24/7 AI 튜터가 강의를 진행하면서 질문에 답변해줍니다.
Web Scraping & Bots을(를) 시작하는 데 경험이 필요한가요?
사전 경험은 필요하지 않습니다. CoddyKit의 Web Scraping & Bots은(는) 초급자부터 고급 학습자까지를 위해 구성되어 있으므로, 여기서 시작하거나 처음부터 시작할 수 있으며 자신의 속도대로 진행할 수 있습니다. 이것은 4개 중 2번째 강의입니다.
“정밀한 CSS 선택자” 강의는 얼마나 걸리나요?
대부분의 CoddyKit 강의는 약 5~10분이 소요됩니다. 각 강의는 간결하고 인터랙티브하여 꾸준한 진행이 가능하며, 웹과 앱에서 중단한 부분부터 바로 시작할 수 있습니다.
이 Web Scraping & Bots 강의에서 코드를 작성하고 실행할 수 있나요?
네. 모든 Web Scraping & Bots 강의에는 내장 코드 에디터가 포함되어 있으므로, 브라우저에서 바로 실제 코드를 작성하고 실행한 후 즉시 AI 피드백을 받을 수 있습니다 — 로컬 설정이 필요 없습니다.