LangChain / RAG / Vector DBs · Lezione

Sviluppo di loader personalizzati per documenti

Crei loader personalizzati per acquisire dati da fonti uniche o proprietarie non supportate direttamente da LangChain.

Lezione 1 di 411 passaggi

Sviluppo di loader personalizzati per documenti è una lezione LangChain / RAG / Vector DBs gratuita su CoddyKit. Questa è la lezione 1 di 4. Puoi leggere la lezione completa qui gratuitamente — poi esercitati direttamente nel browser con un editor di codice integrato e un tutor IA disponibile 24/7. Fa parte del percorso di apprendimento LangChain / RAG / Vector DBs, e i tuoi progressi si sincronizzano tra il web e l'app CoddyKit. Il corso LangChain / RAG / Vector DBs include 4 lezioni in totale.

Parti di questa lezione non sono ancora state tradotte e vengono mostrate in inglese.

Why Custom Document Loaders?

LangChain offers many built-in document loaders for common formats like PDFs, web pages, and databases. But what if your data is unique?

Sometimes, you'll encounter:

  • Proprietary file formats
  • Internal APIs or data sources
  • Complex data structures needing custom parsing

This is where custom document loaders shine!

Meet LangChain's BaseLoader

To create your own loader, you'll inherit from LangChain's BaseLoader class. This is an abstract class, meaning it provides a template for what your loader needs to do.

The most important method you'll implement is load(). This method is responsible for fetching your data and transforming it into a list of Document objects.

The LangChain Document Object

All data processed by LangChain, especially for RAG, is standardized into Document objects. Each Document has two main parts:

  • page_content: The actual text content.
  • metadata: A dictionary of key-value pairs describing the document (e.g., source file, page number, author).

Your custom loader's job is to create these Document objects from your unique data.

Basic Custom Loader Structure

Let's start with a very simple custom loader that just returns a fixed text string as a document. This shows the basic structure of inheriting from BaseLoader and implementing load().

from langchain_core.documents import Document
from langchain_core.document_loaders import BaseLoader

class MySimpleTextLoader(BaseLoader):
    def load(self):
        content = "This is a custom text document from MySimpleTextLoader."
        doc = Document(page_content=content)
        return [doc]

# Example usage:
if __name__ == "__main__":
    loader = MySimpleTextLoader()
    documents = loader.load()
    for doc in documents:
        print(f"Content: {doc.page_content}")
        print(f"Metadata: {doc.metadata}")

Making Your Loader Dynamic

A fixed string loader isn't very useful! Real-world loaders need to take parameters, like a file path, a URL, or API credentials.

You can achieve this by adding an __init__ method to your custom loader class. This allows you to pass arguments when you create an instance of your loader.

Loading a 'Custom' Log File

Imagine you have a simple application log file (app.log) where each line is an event. Let's create a custom loader to read this file, treating each line as a separate document.

We'll create a dummy app.log file content directly in the code for simplicity.

from langchain_core.documents import Document
from langchain_core.document_loaders import BaseLoader

# Simulate a log file content
log_file_content = (
    "[INFO] User logged in: user123\n"
    "[ERROR] Database connection failed\n"
    "[DEBUG] Processing request for /api/data\n"
    "[INFO] Data retrieved successfully"
)

class CustomLogLoader(BaseLoader):
    def __init__(self, log_data: str):
        self.log_data = log_data.split('\n')

    def load(self):
        documents = []
        for line in self.log_data:
            if line.strip(): # Avoid empty lines
                doc = Document(page_content=line)
                documents.append(doc)
        return documents

# Example usage:
if __name__ == "__main__":
    loader = CustomLogLoader(log_file_content)
    documents = loader.load()
    for i, doc in enumerate(documents):
        print(f"Doc {i+1}: {doc.page_content[:40]}...")

Adding Rich Metadata

Metadata is incredibly useful! It helps the LLM understand the context of the text and can be used for filtering or improving retrieval. For our log file example, knowing the original log line number or the source file could be very helpful.

You can add any relevant information as key-value pairs to the metadata dictionary of a Document.

Log Loader with Metadata

Let's enhance our CustomLogLoader to include metadata like the original source and the line number for each log entry. This makes the retrieved information much richer!

from langchain_core.documents import Document
from langchain_core.document_loaders import BaseLoader

# Simulate a log file content
log_file_content = (
    "[INFO] User logged in: user123\n"
    "[ERROR] Database connection failed\n"
    "[DEBUG] Processing request for /api/data\n"
    "[INFO] Data retrieved successfully"
)

class CustomLogLoaderWithMeta(BaseLoader):
    def __init__(self, log_data: str, source_name: str = "app.log"):
        self.log_data = log_data.split('\n')
        self.source_name = source_name

    def load(self):
        documents = []
        for i, line in enumerate(self.log_data):
            if line.strip():
                metadata = {
                    "source": self.source_name,
                    "line_number": i + 1
                }
                doc = Document(page_content=line, metadata=metadata)
                documents.append(doc)
        return documents

# Example usage:
if __name__ == "__main__":
    loader = CustomLogLoaderWithMeta(log_file_content, "my_custom_app_logs")
    documents = loader.load()
    for i, doc in enumerate(documents):
        print(f"Doc {i+1}:")
        print(f"  Content: {doc.page_content[:40]}...")
        print(f"  Metadata: {doc.metadata}")

Integrating Custom Documents

Once your custom loader produces Document objects, they behave just like documents loaded by any other LangChain loader.

You can then pass them into subsequent steps of your RAG pipeline:

  • Text Splitting: Break large documents into smaller chunks.
  • Embeddings: Convert text chunks into numerical vectors.
  • Vector Stores: Store these embeddings for efficient similarity search.

Your custom data is now ready for advanced LLM applications!

Quick Check: Custom Loaders

You're building a custom document loader for LangChain. Which of the following statements is TRUE about the Document object you must return?

Recap: Custom Document Loaders

Great job! You've learned how to develop custom document loaders in LangChain.

  • You inherit from BaseLoader and implement the load() method.
  • Your loader converts unique data into a list of Document objects.
  • Document objects contain page_content and a flexible metadata dictionary.
  • Custom loaders are essential for integrating proprietary data sources into your RAG applications.

Next, we'll explore how to integrate custom embedding models!

Gratis per iniziare

Impara LangChain / RAG / Vector DBs con un tutor IA — gratis

Scrivi ed esegui vero codice nel tuo browser, ricevi aiuto istantaneo da un tutor IA disponibile 24/7, e riprendi da dove hai lasciato sul web o nell'app.

Corsi
12
Lezioni
48

Domande Frequenti

La lezione «Sviluppo di loader personalizzati per documenti» è gratuita?

Sì — il testo completo di «Sviluppo di loader personalizzati per documenti» è gratuito qui sul web. Per esercitarvi in modo interattivo (un editor di codice integrato e un tutor IA 24/7) e sbloccare il resto del corso LangChain / RAG / Vector DBs, passa a CoddyKit PRO. Il corso LangChain / RAG / Vector DBs include 4 lezioni in totale.

Cosa imparerò in «Sviluppo di loader personalizzati per documenti»?

Crei loader personalizzati per acquisire dati da fonti uniche o proprietarie non supportate direttamente da LangChain. Eserciti LangChain / RAG / Vector DBs con codice pratico che esegui direttamente nel browser, e un tutor IA 24/7 risponde alle tue domande mentre lavori sulla lezione.

Ho bisogno di esperienza per iniziare LangChain / RAG / Vector DBs?

Non è richiesta alcuna esperienza precedente. LangChain / RAG / Vector DBs su CoddyKit è strutturato per principianti e studenti avanzati, quindi puoi iniziare da qui o dall'inizio e procedere al tuo ritmo. Questa è la lezione 1 di 4.

Quanto tempo richiede la lezione «Sviluppo di loader personalizzati per documenti»?

La maggior parte delle lezioni CoddyKit richiede circa 5–10 minuti. Ogni lezione è breve e interattiva, quindi fai progressi costanti e riprendi esattamente da dove hai lasciato su web e app.

Posso scrivere ed eseguire codice in questa lezione LangChain / RAG / Vector DBs?

Sì. Ogni lezione LangChain / RAG / Vector DBs include un editor di codice integrato, quindi scrivi ed esegui codice reale direttamente nel tuo browser e ricevi feedback istantaneo dall'IA — nessuna configurazione locale necessaria.

Tutte le lezioni di questo corso

  1. Sviluppo di loader personalizzati per documenti
  2. Integrazione di modelli di embedding personalizzati
  3. Estensione delle catene di retrieval con logica personalizzata
  4. Creare parser di output personalizzati
← Torna a LangChain / RAG / Vector DBs