0Pricing
AI Agents with LangChain & Autonomous Workflows · 课时

网页抓取与数据增强

实施相关技术,使智能体能够从网站提取信息,并动态丰富其知识库

网页抓取与数据增强 是 CoddyKit 上的免费 AI Agents with LangChain & Autonomous Workflows 课时。 这是第 3 节课,共 4 节。 你可以在下方免费阅读本课时的完整内容 — 然后在浏览器中使用内置代码编辑器和全天候 AI 导师进行实践。 这是 AI Agents with LangChain & Autonomous Workflows 学习路径的一部分,你的进度在网页和 CoddyKit 应用中同步。 AI Agents with LangChain & Autonomous Workflows 课程共包含 4 节课。

本课时的部分内容尚未翻译,以英文显示。

Web Scraping for Agents

AI agents often need up-to-date, specific information that isn't part of their pre-trained knowledge or available via simple API calls.

This is where web scraping comes in! It allows agents to "read" and extract data directly from web pages, just like a human would.

What is Web Scraping?

Web scraping is the automated process of extracting data from websites. Instead of manually copying information, a program does it for you.

For AI agents, this means they can dynamically gather facts, prices, news, or any other publicly available information from the internet to inform their decisions or responses.

Essential Python Libraries

Python has powerful libraries that make web scraping straightforward:

  • requests: Used to make HTTP requests (like visiting a website) and get the raw HTML content.
  • BeautifulSoup (often imported as bs4): A library for parsing HTML and XML documents, making it easy to extract data.

Together, they form a common duo for scraping tasks.

Getting Web Page Content

First, we use the requests library to fetch the content of a web page. This sends a request to the server and gets the HTML response.

Try running this example to see how we fetch the content from example.com:

import requests

def main():
    try:
        # Send a GET request to the URL
        response = requests.get("https://www.example.com")
        
        # Check if the request was successful (status code 200)
        if response.status_code == 200:
            print(f"Successfully fetched example.com!")
            print(f"Content Length: {len(response.text)} characters")
        else:
            print(f"Failed to fetch. Status Code: {response.status_code}")
    except requests.exceptions.RequestException as e:
        print(f"Error fetching URL: {e}")

if __name__ == "__main__":
    main()

HTML Basics for Scraping

Web pages are built with HTML (HyperText Markup Language). HTML uses tags like <h1>, <p>, <a> to structure content.

BeautifulSoup helps us navigate this structure. To extract data, we need to know what tags, classes, or IDs the information is wrapped in.

Parsing with BeautifulSoup

Once you have the raw HTML, BeautifulSoup turns it into a parse tree, allowing you to search for elements easily.

Here's how to create a BeautifulSoup object and extract a simple title from a string:

from bs4 import BeautifulSoup

def main():
    html_doc = """
    <html><head><title>CoddyKit Lesson</title></head>
    <body>
    <p class="intro"><b>Hello Learners!</b></p>
    </body></html>
    """
    
    # Create a BeautifulSoup object
    soup = BeautifulSoup(html_doc, 'html.parser')
    
    # Access the title tag and its string content
    print(f"Page Title: {soup.title.string}")

if __name__ == "__main__":
    main()

Finding Specific Data

BeautifulSoup offers powerful methods like find() (for the first match) and find_all() (for all matches) to locate elements by tag name, attributes, or CSS selectors.

Let's find the text inside an <h1> tag:

from bs4 import BeautifulSoup

def main():
    html_doc = """
    <html>
    <body>
        <h1>Welcome to Our Course</h1>
        <p>This is an example paragraph.</p>
        <div id="footer">Contact Us</div>
    </body>
    </html>
    """
    soup = BeautifulSoup(html_doc, 'html.parser')
    
    # Find the first h1 tag
    first_h1 = soup.find('h1')
    if first_h1:
        print(f"Extracted H1 Text: {first_h1.get_text()}")
    else:
        print("H1 tag not found.")

if __name__ == "__main__":
    main()

Data Augmentation for Agents

Data augmentation, in this context, refers to the process of enhancing an agent's internal knowledge base or current context with newly acquired information.

When an agent scrapes a website, it's not just "reading" for itself; it's gathering data that can be used to improve its understanding, answer questions, or make more informed decisions.

Enhancing Agent Capabilities

Imagine an agent that needs to provide real-time stock prices or the latest news headlines. Its pre-trained model won't have this info.

  • Real-time data: Scrape current stock prices or news.
  • Specific facts: Extract product details from an e-commerce site.
  • Contextual understanding: Get background info on a topic not covered in its training.

This scraped data can then be passed to the LLM as part of the prompt, augmenting its knowledge.

Ethical & Legal Considerations

When scraping, always be mindful of:

  • robots.txt: A file on websites that tells bots which parts of the site they shouldn't access. Respect it!
  • Terms of Service: Many sites prohibit scraping in their terms.
  • Rate Limiting: Don't bombard a server with too many requests too quickly; it can be seen as a DoS attack. Introduce delays.

Scrape responsibly and ethically.

Quick Check

You've learned about fetching web content and parsing HTML. Which Python library is primarily used for navigating and extracting data from HTML documents?

Recap & Next Steps

In this lesson, you learned how to use Python's requests and BeautifulSoup libraries for web scraping. We covered fetching raw HTML and then parsing it to extract specific information.

You also understood how this scraped data contributes to data augmentation, enabling AI agents to access current and specific information, thereby enriching their knowledge and capabilities dynamically. Always remember to scrape responsibly!

常见问题解答

「网页抓取与数据增强」课时是免费的吗?

是的 — 「网页抓取与数据增强」的完整文本可在网页上免费阅读。要进行交互式练习(内置代码编辑器和全天候 AI 导师)并解锁 AI Agents with LangChain & Autonomous Workflows 课程的其余内容,请升级到 CoddyKit PRO。 AI Agents with LangChain & Autonomous Workflows 课程共包含 4 节课。

「网页抓取与数据增强」这节课中我会学到什么?

实施相关技术,使智能体能够从网站提取信息,并动态丰富其知识库 你通过在浏览器中直接运行的动手代码来练习 AI Agents with LangChain & Autonomous Workflows,全天候 AI 导师会在你学习这节课的过程中回答你的问题。

学习 AI Agents with LangChain & Autonomous Workflows 需要有经验吗?

无需任何先前经验。CoddyKit 上的 AI Agents with LangChain & Autonomous Workflows 课程适合初学者到高级学习者,你可以从这里开始或从头开始,按照自己的节奏学习。 这是第 3 节课,共 4 节。

「网页抓取与数据增强」课时需要多长时间?

大多数 CoddyKit 课程大约需要 5–10 分钟。每节课都很精短且互动,所以你能稳步进步,并在网页和应用中从离开的地方继续。

我能在这节 AI Agents with LangChain & Autonomous Workflows 课中编写并运行代码吗?

能。每节 AI Agents with LangChain & Autonomous Workflows 课都包含内置代码编辑器,你可以在浏览器中直接编写并运行真实代码,并获得即时 AI 反馈 — 无需本地设置。

此课程中的所有课时

  1. 创建自定义 LangChain 工具
  2. 与外部 API 集成
  3. 网页抓取与数据增强
  4. 工具包与结构化工具输入
← 返回 AI Agents with LangChain & Autonomous Workflows