0Pricing
LLM Apps in Production (RAG + Vector DB + Caching) · 课时

加载多种文档格式

探索从 PDF、网页、数据库和自定义文件类型等各种来源摄取数据的方法。

加载多种文档格式 是 CoddyKit 上的免费 LLM Apps in Production (RAG + Vector DB + Caching) 课时。 这是第 1 节课,共 4 节。 你可以在下方免费阅读本课时的完整内容 — 然后在浏览器中使用内置代码编辑器和全天候 AI 导师进行实践。 这是 LLM Apps in Production (RAG + Vector DB + Caching) 学习路径的一部分,你的进度在网页和 CoddyKit 应用中同步。 LLM Apps in Production (RAG + Vector DB + Caching) 课程共包含 4 节课。

本课时的部分内容尚未翻译,以英文显示。

Ingesting Diverse Document Types

Welcome! In RAG, your LLM needs information from various sources. This lesson explores how to load data from different document formats into your application.

The goal is to get raw text from places like web pages, PDFs, and databases, preparing it for the next steps in your RAG pipeline.

Loading Web Pages (HTML)

Web pages are a common source of information. To ingest them, you typically:

  • Fetch the HTML: Use an HTTP client to download the page content from a URL.
  • Parse the HTML: Extract the main text and discard navigation, ads, and other irrelevant elements.

Libraries like requests for fetching and BeautifulSoup for parsing are very popular in Python.

Web Page Loading Example

This Python snippet demonstrates fetching a simple web page and extracting its title. Imagine doing this for many pages to build your knowledge base!

import requests
from bs4 import BeautifulSoup

def main():
    url = "http://quotes.toscrape.com/"
    try:
        response = requests.get(url)
        response.raise_for_status() # Raise HTTPError for bad responses
        soup = BeautifulSoup(response.text, 'html.parser')
        title = soup.find('title').get_text()
        print(f"Page Title: {title}")
    except requests.exceptions.RequestException as e:
        print(f"Error fetching URL: {e}")

if __name__ == "__main__":
    main()

Extracting Text from PDFs

PDFs (Portable Document Format) are widely used for reports and documents. Extracting text from PDFs can be tricky due to their complex structure, which combines text, images, and formatting.

Fortunately, programming libraries exist to help. They can read the PDF structure and pull out the textual content, often page by page.

PDF Text Extraction Example

This conceptual Python code shows how you might extract text from the first page of a PDF using a library like pypdf. For a real run, you'd need a sample.pdf file.

from pypdf import PdfReader

def main():
    # In a real scenario, 'sample.pdf' would exist
    # For this example, we'll simulate the output
    pdf_file_path = "sample.pdf"
    print(f"Attempting to read from: {pdf_file_path}")
    print("\n--- Simulated PDF Content ---")
    print("This is some text from the first page of a sample PDF document.")
    print("It contains important information for our RAG system.")
    print("-----------------------------")
    # Actual code might look like this:
    # reader = PdfReader(pdf_file_path)
    # page = reader.pages[0]
    # text = page.extract_text()
    # print(text)

if __name__ == "__main__":
    main()

Loading Data from Databases

Databases, both SQL (like PostgreSQL, MySQL) and NoSQL (like MongoDB, Cassandra), store structured data that can be valuable for RAG.

To ingest from databases:

  • Connect: Establish a connection using database drivers.
  • Query: Write queries (e.g., SQL statements) to retrieve relevant data.
  • Process: Extract text fields from the query results.

Database Loading Example

Here's a Python example using SQLite, an embedded SQL database. It creates a simple table, inserts data, and then retrieves it. This is a common pattern for database ingestion.

import sqlite3

def main():
    # Connect to an in-memory SQLite database
    conn = sqlite3.connect(':memory:')
    cursor = conn.cursor()

    # Create a simple table
    cursor.execute('''
        CREATE TABLE IF NOT EXISTS documents (
            id INTEGER PRIMARY KEY,
            title TEXT,
            content TEXT
        )
    ''')

    # Insert some data
    cursor.execute("INSERT INTO documents (title, content) VALUES (?, ?)", 
                   ("RAG Overview", "RAG enhances LLMs by retrieving relevant docs."))
    cursor.execute("INSERT INTO documents (title, content) VALUES (?, ?)", 
                   ("Vector DBs", "Store embeddings for fast similarity search."))
    conn.commit()

    # Retrieve data
    cursor.execute("SELECT title, content FROM documents")
    rows = cursor.fetchall()

    print("--- Retrieved Documents ---")
    for row in rows:
        print(f"Title: {row[0]}, Content: {row[1]}")
    print("---------------------------")

    conn.close()

if __name__ == "__main__":
    main()

Handling Plain Text & Custom Files

Beyond specific formats, you'll often deal with plain text files (.txt, .md) or custom formats (e.g., CSV, JSON). For these:

  • Plain Text: Read directly, paying attention to encoding (UTF-8 is common).
  • Custom Formats: Use libraries specific to the format (e.g., csv, json modules in Python) to parse and extract text fields.

The key is transforming the data into a usable text string.

Unified Data Loading with Libraries

For complex RAG systems, you don't always need to write custom loaders for every format. Libraries like LlamaIndex and LangChain offer 'Document Loaders' that abstract away much of this complexity.

  • They provide connectors for many data sources (web, PDF, databases, cloud storage).
  • They often handle basic parsing and text extraction automatically.

These tools simplify the initial ingestion step, letting you focus on retrieval and generation.

Check Your Knowledge

Which of the following is typically a primary challenge when extracting text content from PDF documents for a RAG system?

Recap: Loading Diverse Data

Great job! You've learned the fundamentals of loading diverse document formats for your RAG system.

  • We covered fetching and parsing web pages.
  • Discussed extracting text from complex PDFs.
  • Explored querying databases for structured content.
  • Touched upon handling plain text and other custom files.
  • Recognized the value of unified data loading libraries.

The next step is to prepare this raw text for effective retrieval!

常见问题解答

「加载多种文档格式」课时是免费的吗?

是的 — 「加载多种文档格式」的完整文本可在网页上免费阅读。要进行交互式练习(内置代码编辑器和全天候 AI 导师)并解锁 LLM Apps in Production (RAG + Vector DB + Caching) 课程的其余内容,请升级到 CoddyKit PRO。 LLM Apps in Production (RAG + Vector DB + Caching) 课程共包含 4 节课。

「加载多种文档格式」这节课中我会学到什么?

探索从 PDF、网页、数据库和自定义文件类型等各种来源摄取数据的方法。 你通过在浏览器中直接运行的动手代码来练习 LLM Apps in Production (RAG + Vector DB + Caching),全天候 AI 导师会在你学习这节课的过程中回答你的问题。

学习 LLM Apps in Production (RAG + Vector DB + Caching) 需要有经验吗?

无需任何先前经验。CoddyKit 上的 LLM Apps in Production (RAG + Vector DB + Caching) 课程适合初学者到高级学习者,你可以从这里开始或从头开始,按照自己的节奏学习。 这是第 1 节课,共 4 节。

「加载多种文档格式」课时需要多长时间?

大多数 CoddyKit 课程大约需要 5–10 分钟。每节课都很精短且互动,所以你能稳步进步,并在网页和应用中从离开的地方继续。

我能在这节 LLM Apps in Production (RAG + Vector DB + Caching) 课中编写并运行代码吗?

能。每节 LLM Apps in Production (RAG + Vector DB + Caching) 课都包含内置代码编辑器,你可以在浏览器中直接编写并运行真实代码,并获得即时 AI 反馈 — 无需本地设置。

此课程中的所有课时

  1. 加载多种文档格式
  2. 上下文感知的分块策略
  3. 元数据管理与筛选
  4. 清洗源数据并去除重复内容
← 返回 LLM Apps in Production (RAG + Vector DB + Caching)