웹 스크래핑과 데이터 보강
에이전트가 웹사이트에서 정보를 추출하고 지식 기반을 동적으로 보강하도록 하는 기법을 구현합니다.
웹 스크래핑과 데이터 보강은(는) CoddyKit의 무료 AI Agents with LangChain & Autonomous Workflows 강의입니다. 이것은 4개 중 3번째 강의입니다. 아래에서 전체 강의를 무료로 읽을 수 있으며, 내장 코드 에디터와 24/7 AI 튜터와 함께 브라우저에서 직접 실습할 수 있습니다. 이 강의는 AI Agents with LangChain & Autonomous Workflows 학습 경로의 일부이며, 진행 상황이 웹과 CoddyKit 앱에 동기화됩니다. AI Agents with LangChain & Autonomous Workflows 강의에는 총 4개의 강의가 포함되어 있습니다.
이 강의의 일부는 아직 번역되지 않았으며 영어로 표시됩니다.
Web Scraping for Agents
AI agents often need up-to-date, specific information that isn't part of their pre-trained knowledge or available via simple API calls.
This is where web scraping comes in! It allows agents to "read" and extract data directly from web pages, just like a human would.
What is Web Scraping?
Web scraping is the automated process of extracting data from websites. Instead of manually copying information, a program does it for you.
For AI agents, this means they can dynamically gather facts, prices, news, or any other publicly available information from the internet to inform their decisions or responses.
Essential Python Libraries
Python has powerful libraries that make web scraping straightforward:
requests: Used to make HTTP requests (like visiting a website) and get the raw HTML content.BeautifulSoup(often imported asbs4): A library for parsing HTML and XML documents, making it easy to extract data.
Together, they form a common duo for scraping tasks.
Getting Web Page Content
First, we use the requests library to fetch the content of a web page. This sends a request to the server and gets the HTML response.
Try running this example to see how we fetch the content from example.com:
import requests
def main():
try:
# Send a GET request to the URL
response = requests.get("https://www.example.com")
# Check if the request was successful (status code 200)
if response.status_code == 200:
print(f"Successfully fetched example.com!")
print(f"Content Length: {len(response.text)} characters")
else:
print(f"Failed to fetch. Status Code: {response.status_code}")
except requests.exceptions.RequestException as e:
print(f"Error fetching URL: {e}")
if __name__ == "__main__":
main()HTML Basics for Scraping
Web pages are built with HTML (HyperText Markup Language). HTML uses tags like <h1>, <p>, <a> to structure content.
BeautifulSoup helps us navigate this structure. To extract data, we need to know what tags, classes, or IDs the information is wrapped in.
Parsing with BeautifulSoup
Once you have the raw HTML, BeautifulSoup turns it into a parse tree, allowing you to search for elements easily.
Here's how to create a BeautifulSoup object and extract a simple title from a string:
from bs4 import BeautifulSoup
def main():
html_doc = """
<html><head><title>CoddyKit Lesson</title></head>
<body>
<p class="intro"><b>Hello Learners!</b></p>
</body></html>
"""
# Create a BeautifulSoup object
soup = BeautifulSoup(html_doc, 'html.parser')
# Access the title tag and its string content
print(f"Page Title: {soup.title.string}")
if __name__ == "__main__":
main()Finding Specific Data
BeautifulSoup offers powerful methods like find() (for the first match) and find_all() (for all matches) to locate elements by tag name, attributes, or CSS selectors.
Let's find the text inside an <h1> tag:
from bs4 import BeautifulSoup
def main():
html_doc = """
<html>
<body>
<h1>Welcome to Our Course</h1>
<p>This is an example paragraph.</p>
<div id="footer">Contact Us</div>
</body>
</html>
"""
soup = BeautifulSoup(html_doc, 'html.parser')
# Find the first h1 tag
first_h1 = soup.find('h1')
if first_h1:
print(f"Extracted H1 Text: {first_h1.get_text()}")
else:
print("H1 tag not found.")
if __name__ == "__main__":
main()Data Augmentation for Agents
Data augmentation, in this context, refers to the process of enhancing an agent's internal knowledge base or current context with newly acquired information.
When an agent scrapes a website, it's not just "reading" for itself; it's gathering data that can be used to improve its understanding, answer questions, or make more informed decisions.
Enhancing Agent Capabilities
Imagine an agent that needs to provide real-time stock prices or the latest news headlines. Its pre-trained model won't have this info.
- Real-time data: Scrape current stock prices or news.
- Specific facts: Extract product details from an e-commerce site.
- Contextual understanding: Get background info on a topic not covered in its training.
This scraped data can then be passed to the LLM as part of the prompt, augmenting its knowledge.
Ethical & Legal Considerations
When scraping, always be mindful of:
robots.txt: A file on websites that tells bots which parts of the site they shouldn't access. Respect it!- Terms of Service: Many sites prohibit scraping in their terms.
- Rate Limiting: Don't bombard a server with too many requests too quickly; it can be seen as a DoS attack. Introduce delays.
Scrape responsibly and ethically.
Quick Check
You've learned about fetching web content and parsing HTML. Which Python library is primarily used for navigating and extracting data from HTML documents?
Recap & Next Steps
In this lesson, you learned how to use Python's requests and BeautifulSoup libraries for web scraping. We covered fetching raw HTML and then parsing it to extract specific information.
You also understood how this scraped data contributes to data augmentation, enabling AI agents to access current and specific information, thereby enriching their knowledge and capabilities dynamically. Always remember to scrape responsibly!
자주 묻는 질문
“웹 스크래핑과 데이터 보강” 강의는 무료인가요?
네 — “웹 스크래핑과 데이터 보강” 전체 내용을 이 웹사이트에서 무료로 읽을 수 있습니다. 인터랙티브하게 실습하려면(내장 코드 에디터와 24/7 AI 튜터), CoddyKit PRO로 업그레이드하면 AI Agents with LangChain & Autonomous Workflows 강의 전체를 잠금 해제할 수 있습니다. AI Agents with LangChain & Autonomous Workflows 강의에는 총 4개의 강의가 포함되어 있습니다.
“웹 스크래핑과 데이터 보강”에서 뭘 배우나요?
에이전트가 웹사이트에서 정보를 추출하고 지식 기반을 동적으로 보강하도록 하는 기법을 구현합니다. 브라우저에서 직접 실행하는 실습 코드로 AI Agents with LangChain & Autonomous Workflows을(를) 배우며, 24/7 AI 튜터가 강의를 진행하면서 질문에 답변해줍니다.
AI Agents with LangChain & Autonomous Workflows을(를) 시작하는 데 경험이 필요한가요?
사전 경험은 필요하지 않습니다. CoddyKit의 AI Agents with LangChain & Autonomous Workflows은(는) 초급자부터 고급 학습자까지를 위해 구성되어 있으므로, 여기서 시작하거나 처음부터 시작할 수 있으며 자신의 속도대로 진행할 수 있습니다. 이것은 4개 중 3번째 강의입니다.
“웹 스크래핑과 데이터 보강” 강의는 얼마나 걸리나요?
대부분의 CoddyKit 강의는 약 5~10분이 소요됩니다. 각 강의는 간결하고 인터랙티브하여 꾸준한 진행이 가능하며, 웹과 앱에서 중단한 부분부터 바로 시작할 수 있습니다.
이 AI Agents with LangChain & Autonomous Workflows 강의에서 코드를 작성하고 실행할 수 있나요?
네. 모든 AI Agents with LangChain & Autonomous Workflows 강의에는 내장 코드 에디터가 포함되어 있으므로, 브라우저에서 바로 실제 코드를 작성하고 실행한 후 즉시 AI 피드백을 받을 수 있습니다 — 로컬 설정이 필요 없습니다.
이 강의의 모든 강의
- 사용자 지정 LangChain 도구 만들기
- 외부 API 통합
- 웹 스크래핑과 데이터 보강
- 도구 모음 및 구조화된 도구 입력