0Pricing
AI Agents with LangChain & Autonomous Workflows · 강의

문서 로더 해설

PDF, 웹 페이지, 데이터베이스 등 다양한 원천의 데이터를 LangChain에서 사용할 수 있는 형식으로 불러오는 방법을 알아봅니다.

문서 로더 해설은(는) CoddyKit의 무료 AI Agents with LangChain & Autonomous Workflows 강의입니다. 이것은 4개 중 1번째 강의입니다. 아래에서 전체 강의를 무료로 읽을 수 있으며, 내장 코드 에디터와 24/7 AI 튜터와 함께 브라우저에서 직접 실습할 수 있습니다. 이 강의는 AI Agents with LangChain & Autonomous Workflows 학습 경로의 일부이며, 진행 상황이 웹과 CoddyKit 앱에 동기화됩니다. AI Agents with LangChain & Autonomous Workflows 강의에는 총 4개의 강의가 포함되어 있습니다.

이 강의의 일부는 아직 번역되지 않았으며 영어로 표시됩니다.

Data for Smarter LLMs

Large Language Models (LLMs) are powerful, but their knowledge is limited to their training data. To make AI agents truly useful, they often need access to fresh, external information.

  • Imagine an agent that needs to answer questions about today's news.
  • Or one that summarizes a specific document from your company's internal drive.

This is where data loading comes in.

What are Document Loaders?

Document Loaders are LangChain's way of bringing external data into your agent's workflow. They act as bridges, converting raw data from various sources into a standardized format that LLMs can understand.

Think of them as specialized data connectors. They handle the messy details of reading files, fetching web pages, or querying databases, and then present the data cleanly to LangChain.

The LangChain Document Object

When a loader fetches data, it wraps it into a Document object. This is LangChain's universal container for data.

Each Document typically has two main parts:

  • page_content: The actual text content extracted from the source (e.g., the text from a PDF page, a web article).
  • metadata: A dictionary containing extra information about the document (e.g., source URL, page number, file path, creation date).

Loading from Local Text Files

One of the simplest ways to load data is from a plain text file on your computer. LangChain provides the TextLoader for this purpose.

It takes a file path, reads the content, and creates a Document object. Let's see how it works by creating a temporary file and loading it.

TextLoader in Action

This example creates a temporary text file, writes some content, then uses TextLoader to read it and prints the result.

from langchain_community.document_loaders import TextLoader
import os
import tempfile

if __name__ == "__main__":
    # Create a temporary file
    with tempfile.NamedTemporaryFile(mode='w', delete=False, encoding='utf-8') as temp_file:
        temp_file.write("Hello CoddyKit!\nThis is a test document.")
        temp_file_path = temp_file.name

    print(f"Created temp file: {temp_file_path}")

    try:
        # Load the document using TextLoader
        loader = TextLoader(temp_file_path)
        documents = loader.load()

        # Print the content of the loaded document
        for doc in documents:
            print("--- Document Content ---")
            print(doc.page_content)
            print("--- Metadata ---")
            print(doc.metadata)
    finally:
        # Clean up the temporary file
        os.remove(temp_file_path)
        print(f"Cleaned up temp file: {temp_file_path}")

Web Content with WebBaseLoader

What if your data is on the internet? The WebBaseLoader is perfect for fetching content directly from web pages.

It's smart enough to extract the main text content, ignoring navigation, ads, and other irrelevant elements. You just provide the URL!

WebBaseLoader Example

Here's how to fetch content from a simple example web page. Note: This loader typically requires beautifulsoup4 and lxml to be installed.

from langchain_community.document_loaders import WebBaseLoader

if __name__ == "__main__":
    # Define the URL to load
    url = "https://www.google.com/search?q=hello"
    print(f"Loading content from: {url}")

    # Initialize the WebBaseLoader
    loader = WebBaseLoader(url)

    # Load the documents
    documents = loader.load()

    # Print the content of the first loaded document
    if documents:
        print("--- Extracted Content (first 200 chars) ---")
        print(documents[0].page_content[:200] + "...")
        print("--- Metadata ---")
        print(documents[0].metadata)
    else:
        print("No documents loaded.")

PDFs and Structured Documents

Loading data from PDFs is a common requirement. LangChain offers loaders like PyPDFLoader (which uses the pypdf library) to handle these.

These loaders are designed to extract text from complex formats, often preserving some structure. For PDFs, each page might become a separate Document object, with metadata indicating its page number.

There are also loaders for other formats like CSV, JSON, Markdown, and even specific data structures like Notion databases or Confluence pages.

Beyond Files: Database Loaders

Your data might live in databases. LangChain supports various database loaders to connect directly to your data sources:

  • SQL Database Loader: Connects to relational databases (e.g., PostgreSQL, MySQL) to fetch data based on queries.
  • MongoDB Loader: For NoSQL document databases.
  • Elasticsearch Loader: To pull data from Elasticsearch indices.

These loaders allow your agents to dynamically query and retrieve data from your existing data infrastructure.

Quick Check

Which of the following is NOT a primary purpose of a LangChain Document Loader?

Recap: Document Loaders

In this lesson, we explored LangChain's Document Loaders, essential tools for bringing external data into your AI agents.

  • Loaders convert diverse data sources (text files, web pages, PDFs, databases) into LangChain's universal Document object.
  • Each Document contains page_content and metadata.
  • We saw practical examples with TextLoader for local files and WebBaseLoader for web content.

Next, we'll learn how to handle large documents by splitting them into manageable chunks.

자주 묻는 질문

“문서 로더 해설” 강의는 무료인가요?

네 — “문서 로더 해설” 전체 내용을 이 웹사이트에서 무료로 읽을 수 있습니다. 인터랙티브하게 실습하려면(내장 코드 에디터와 24/7 AI 튜터), CoddyKit PRO로 업그레이드하면 AI Agents with LangChain & Autonomous Workflows 강의 전체를 잠금 해제할 수 있습니다. AI Agents with LangChain & Autonomous Workflows 강의에는 총 4개의 강의가 포함되어 있습니다.

“문서 로더 해설”에서 뭘 배우나요?

PDF, 웹 페이지, 데이터베이스 등 다양한 원천의 데이터를 LangChain에서 사용할 수 있는 형식으로 불러오는 방법을 알아봅니다. 브라우저에서 직접 실행하는 실습 코드로 AI Agents with LangChain & Autonomous Workflows을(를) 배우며, 24/7 AI 튜터가 강의를 진행하면서 질문에 답변해줍니다.

AI Agents with LangChain & Autonomous Workflows을(를) 시작하는 데 경험이 필요한가요?

사전 경험은 필요하지 않습니다. CoddyKit의 AI Agents with LangChain & Autonomous Workflows은(는) 초급자부터 고급 학습자까지를 위해 구성되어 있으므로, 여기서 시작하거나 처음부터 시작할 수 있으며 자신의 속도대로 진행할 수 있습니다. 이것은 4개 중 1번째 강의입니다.

“문서 로더 해설” 강의는 얼마나 걸리나요?

대부분의 CoddyKit 강의는 약 5~10분이 소요됩니다. 각 강의는 간결하고 인터랙티브하여 꾸준한 진행이 가능하며, 웹과 앱에서 중단한 부분부터 바로 시작할 수 있습니다.

이 AI Agents with LangChain & Autonomous Workflows 강의에서 코드를 작성하고 실행할 수 있나요?

네. 모든 AI Agents with LangChain & Autonomous Workflows 강의에는 내장 코드 에디터가 포함되어 있으므로, 브라우저에서 바로 실제 코드를 작성하고 실행한 후 즉시 AI 피드백을 받을 수 있습니다 — 로컬 설정이 필요 없습니다.

이 강의의 모든 강의

  1. 문서 로더 해설
  2. 텍스트 분할기와 임베딩
  3. 검색을 위한 벡터 저장소
  4. 검색기 및 맥락 압축
← AI Agents with LangChain & Autonomous Workflows(으)로 돌아가기