文档加载器详解
了解如何将 PDF、网页和数据库等多种来源的数据加载为 LangChain 可用的格式
文档加载器详解 是 CoddyKit 上的免费 AI Agents with LangChain & Autonomous Workflows 课时。 这是第 1 节课,共 4 节。 你可以在下方免费阅读本课时的完整内容 — 然后在浏览器中使用内置代码编辑器和全天候 AI 导师进行实践。 这是 AI Agents with LangChain & Autonomous Workflows 学习路径的一部分,你的进度在网页和 CoddyKit 应用中同步。 AI Agents with LangChain & Autonomous Workflows 课程共包含 4 节课。
本课时的部分内容尚未翻译,以英文显示。
Data for Smarter LLMs
Large Language Models (LLMs) are powerful, but their knowledge is limited to their training data. To make AI agents truly useful, they often need access to fresh, external information.
- Imagine an agent that needs to answer questions about today's news.
- Or one that summarizes a specific document from your company's internal drive.
This is where data loading comes in.
What are Document Loaders?
Document Loaders are LangChain's way of bringing external data into your agent's workflow. They act as bridges, converting raw data from various sources into a standardized format that LLMs can understand.
Think of them as specialized data connectors. They handle the messy details of reading files, fetching web pages, or querying databases, and then present the data cleanly to LangChain.
The LangChain Document Object
When a loader fetches data, it wraps it into a Document object. This is LangChain's universal container for data.
Each Document typically has two main parts:
page_content: The actual text content extracted from the source (e.g., the text from a PDF page, a web article).metadata: A dictionary containing extra information about the document (e.g., source URL, page number, file path, creation date).
Loading from Local Text Files
One of the simplest ways to load data is from a plain text file on your computer. LangChain provides the TextLoader for this purpose.
It takes a file path, reads the content, and creates a Document object. Let's see how it works by creating a temporary file and loading it.
TextLoader in Action
This example creates a temporary text file, writes some content, then uses TextLoader to read it and prints the result.
from langchain_community.document_loaders import TextLoader
import os
import tempfile
if __name__ == "__main__":
# Create a temporary file
with tempfile.NamedTemporaryFile(mode='w', delete=False, encoding='utf-8') as temp_file:
temp_file.write("Hello CoddyKit!\nThis is a test document.")
temp_file_path = temp_file.name
print(f"Created temp file: {temp_file_path}")
try:
# Load the document using TextLoader
loader = TextLoader(temp_file_path)
documents = loader.load()
# Print the content of the loaded document
for doc in documents:
print("--- Document Content ---")
print(doc.page_content)
print("--- Metadata ---")
print(doc.metadata)
finally:
# Clean up the temporary file
os.remove(temp_file_path)
print(f"Cleaned up temp file: {temp_file_path}")Web Content with WebBaseLoader
What if your data is on the internet? The WebBaseLoader is perfect for fetching content directly from web pages.
It's smart enough to extract the main text content, ignoring navigation, ads, and other irrelevant elements. You just provide the URL!
WebBaseLoader Example
Here's how to fetch content from a simple example web page. Note: This loader typically requires beautifulsoup4 and lxml to be installed.
from langchain_community.document_loaders import WebBaseLoader
if __name__ == "__main__":
# Define the URL to load
url = "https://www.google.com/search?q=hello"
print(f"Loading content from: {url}")
# Initialize the WebBaseLoader
loader = WebBaseLoader(url)
# Load the documents
documents = loader.load()
# Print the content of the first loaded document
if documents:
print("--- Extracted Content (first 200 chars) ---")
print(documents[0].page_content[:200] + "...")
print("--- Metadata ---")
print(documents[0].metadata)
else:
print("No documents loaded.")PDFs and Structured Documents
Loading data from PDFs is a common requirement. LangChain offers loaders like PyPDFLoader (which uses the pypdf library) to handle these.
These loaders are designed to extract text from complex formats, often preserving some structure. For PDFs, each page might become a separate Document object, with metadata indicating its page number.
There are also loaders for other formats like CSV, JSON, Markdown, and even specific data structures like Notion databases or Confluence pages.
Beyond Files: Database Loaders
Your data might live in databases. LangChain supports various database loaders to connect directly to your data sources:
- SQL Database Loader: Connects to relational databases (e.g., PostgreSQL, MySQL) to fetch data based on queries.
- MongoDB Loader: For NoSQL document databases.
- Elasticsearch Loader: To pull data from Elasticsearch indices.
These loaders allow your agents to dynamically query and retrieve data from your existing data infrastructure.
Quick Check
Which of the following is NOT a primary purpose of a LangChain Document Loader?
Recap: Document Loaders
In this lesson, we explored LangChain's Document Loaders, essential tools for bringing external data into your AI agents.
- Loaders convert diverse data sources (text files, web pages, PDFs, databases) into LangChain's universal
Documentobject. - Each
Documentcontainspage_contentandmetadata. - We saw practical examples with
TextLoaderfor local files andWebBaseLoaderfor web content.
Next, we'll learn how to handle large documents by splitting them into manageable chunks.
常见问题解答
「文档加载器详解」课时是免费的吗?
是的 — 「文档加载器详解」的完整文本可在网页上免费阅读。要进行交互式练习(内置代码编辑器和全天候 AI 导师)并解锁 AI Agents with LangChain & Autonomous Workflows 课程的其余内容,请升级到 CoddyKit PRO。 AI Agents with LangChain & Autonomous Workflows 课程共包含 4 节课。
「文档加载器详解」这节课中我会学到什么?
了解如何将 PDF、网页和数据库等多种来源的数据加载为 LangChain 可用的格式 你通过在浏览器中直接运行的动手代码来练习 AI Agents with LangChain & Autonomous Workflows,全天候 AI 导师会在你学习这节课的过程中回答你的问题。
学习 AI Agents with LangChain & Autonomous Workflows 需要有经验吗?
无需任何先前经验。CoddyKit 上的 AI Agents with LangChain & Autonomous Workflows 课程适合初学者到高级学习者,你可以从这里开始或从头开始,按照自己的节奏学习。 这是第 1 节课,共 4 节。
「文档加载器详解」课时需要多长时间?
大多数 CoddyKit 课程大约需要 5–10 分钟。每节课都很精短且互动,所以你能稳步进步,并在网页和应用中从离开的地方继续。
我能在这节 AI Agents with LangChain & Autonomous Workflows 课中编写并运行代码吗?
能。每节 AI Agents with LangChain & Autonomous Workflows 课都包含内置代码编辑器,你可以在浏览器中直接编写并运行真实代码,并获得即时 AI 反馈 — 无需本地设置。