0Pricing
AI Agents with LangChain & Autonomous Workflows · Pelajaran

Penjelasan Pemuat Dokumen

Temukan cara memuat data dari berbagai sumber seperti PDF, halaman web, dan basis data ke dalam format yang dapat digunakan LangChain.

Penjelasan Pemuat Dokumen adalah pelajaran AI Agents with LangChain & Autonomous Workflows gratis di CoddyKit. Ini adalah pelajaran 1 dari 4. Kamu bisa membaca pelajaran lengkapnya di bawah secara gratis — lalu praktikkan langsung di browser dengan editor kode bawaan dan tutor AI 24/7. Ini adalah bagian dari jalur belajar AI Agents with LangChain & Autonomous Workflows, dan progresmu tersinkronisasi di web dan aplikasi CoddyKit. Kursus AI Agents with LangChain & Autonomous Workflows mencakup 4 pelajaran total.

Bagian dari pelajaran ini belum diterjemahkan dan ditampilkan dalam bahasa Inggris.

Data for Smarter LLMs

Large Language Models (LLMs) are powerful, but their knowledge is limited to their training data. To make AI agents truly useful, they often need access to fresh, external information.

  • Imagine an agent that needs to answer questions about today's news.
  • Or one that summarizes a specific document from your company's internal drive.

This is where data loading comes in.

What are Document Loaders?

Document Loaders are LangChain's way of bringing external data into your agent's workflow. They act as bridges, converting raw data from various sources into a standardized format that LLMs can understand.

Think of them as specialized data connectors. They handle the messy details of reading files, fetching web pages, or querying databases, and then present the data cleanly to LangChain.

The LangChain Document Object

When a loader fetches data, it wraps it into a Document object. This is LangChain's universal container for data.

Each Document typically has two main parts:

  • page_content: The actual text content extracted from the source (e.g., the text from a PDF page, a web article).
  • metadata: A dictionary containing extra information about the document (e.g., source URL, page number, file path, creation date).

Loading from Local Text Files

One of the simplest ways to load data is from a plain text file on your computer. LangChain provides the TextLoader for this purpose.

It takes a file path, reads the content, and creates a Document object. Let's see how it works by creating a temporary file and loading it.

TextLoader in Action

This example creates a temporary text file, writes some content, then uses TextLoader to read it and prints the result.

from langchain_community.document_loaders import TextLoader
import os
import tempfile

if __name__ == "__main__":
    # Create a temporary file
    with tempfile.NamedTemporaryFile(mode='w', delete=False, encoding='utf-8') as temp_file:
        temp_file.write("Hello CoddyKit!\nThis is a test document.")
        temp_file_path = temp_file.name

    print(f"Created temp file: {temp_file_path}")

    try:
        # Load the document using TextLoader
        loader = TextLoader(temp_file_path)
        documents = loader.load()

        # Print the content of the loaded document
        for doc in documents:
            print("--- Document Content ---")
            print(doc.page_content)
            print("--- Metadata ---")
            print(doc.metadata)
    finally:
        # Clean up the temporary file
        os.remove(temp_file_path)
        print(f"Cleaned up temp file: {temp_file_path}")

Web Content with WebBaseLoader

What if your data is on the internet? The WebBaseLoader is perfect for fetching content directly from web pages.

It's smart enough to extract the main text content, ignoring navigation, ads, and other irrelevant elements. You just provide the URL!

WebBaseLoader Example

Here's how to fetch content from a simple example web page. Note: This loader typically requires beautifulsoup4 and lxml to be installed.

from langchain_community.document_loaders import WebBaseLoader

if __name__ == "__main__":
    # Define the URL to load
    url = "https://www.google.com/search?q=hello"
    print(f"Loading content from: {url}")

    # Initialize the WebBaseLoader
    loader = WebBaseLoader(url)

    # Load the documents
    documents = loader.load()

    # Print the content of the first loaded document
    if documents:
        print("--- Extracted Content (first 200 chars) ---")
        print(documents[0].page_content[:200] + "...")
        print("--- Metadata ---")
        print(documents[0].metadata)
    else:
        print("No documents loaded.")

PDFs and Structured Documents

Loading data from PDFs is a common requirement. LangChain offers loaders like PyPDFLoader (which uses the pypdf library) to handle these.

These loaders are designed to extract text from complex formats, often preserving some structure. For PDFs, each page might become a separate Document object, with metadata indicating its page number.

There are also loaders for other formats like CSV, JSON, Markdown, and even specific data structures like Notion databases or Confluence pages.

Beyond Files: Database Loaders

Your data might live in databases. LangChain supports various database loaders to connect directly to your data sources:

  • SQL Database Loader: Connects to relational databases (e.g., PostgreSQL, MySQL) to fetch data based on queries.
  • MongoDB Loader: For NoSQL document databases.
  • Elasticsearch Loader: To pull data from Elasticsearch indices.

These loaders allow your agents to dynamically query and retrieve data from your existing data infrastructure.

Quick Check

Which of the following is NOT a primary purpose of a LangChain Document Loader?

Recap: Document Loaders

In this lesson, we explored LangChain's Document Loaders, essential tools for bringing external data into your AI agents.

  • Loaders convert diverse data sources (text files, web pages, PDFs, databases) into LangChain's universal Document object.
  • Each Document contains page_content and metadata.
  • We saw practical examples with TextLoader for local files and WebBaseLoader for web content.

Next, we'll learn how to handle large documents by splitting them into manageable chunks.

Pertanyaan yang Sering Diajukan

Apakah pelajaran “Penjelasan Pemuat Dokumen” gratis?

Ya — teks lengkap “Penjelasan Pemuat Dokumen” gratis dibaca di sini di web. Untuk praktiknya secara interaktif (editor kode bawaan dan tutor AI 24/7) dan buka sisa kursus AI Agents with LangChain & Autonomous Workflows, upgrade ke CoddyKit PRO. Kursus AI Agents with LangChain & Autonomous Workflows mencakup 4 pelajaran total.

Apa yang akan aku pelajari di “Penjelasan Pemuat Dokumen”?

Temukan cara memuat data dari berbagai sumber seperti PDF, halaman web, dan basis data ke dalam format yang dapat digunakan LangChain. Kamu berlatih AI Agents with LangChain & Autonomous Workflows dengan kode praktik yang langsung kamu jalankan di browser, dan tutor AI 24/7 menjawab pertanyaanmu saat kamu mengerjakan pelajaran ini.

Apakah aku perlu pengalaman untuk memulai AI Agents with LangChain & Autonomous Workflows?

Tidak diperlukan pengalaman sebelumnya. AI Agents with LangChain & Autonomous Workflows di CoddyKit dirancang untuk pemula hingga pelajar tingkat lanjut, jadi kamu bisa memulai di sini atau dari awal dan belajar sesuai kecepatan kamu sendiri. Ini adalah pelajaran 1 dari 4.

Berapa lama pelajaran “Penjelasan Pemuat Dokumen” memakan waktu?

Sebagian besar pelajaran CoddyKit memakan waktu sekitar 5–10 menit. Setiap pelajaran ringkas dan interaktif, jadi kamu membuat kemajuan stabil dan melanjutkan dari tempat kamu tinggalkan di web dan aplikasi.

Bisakah aku menulis dan menjalankan kode dalam pelajaran AI Agents with LangChain & Autonomous Workflows ini?

Ya. Setiap pelajaran AI Agents with LangChain & Autonomous Workflows menyertakan editor kode bawaan, jadi kamu menulis dan menjalankan kode nyata langsung di browser dan mendapatkan umpan balik AI instan — tidak diperlukan penyiapan lokal.

Semua pelajaran dalam kursus ini

  1. Penjelasan Pemuat Dokumen
  2. Pemisah Teks dan Embedding
  3. Penyimpanan Vektor untuk Pengambilan
  4. Pengambil dan Kompresi Kontekstual
← Kembali ke AI Agents with LangChain & Autonomous Workflows