カスタムドキュメントローダーの開発
LangChainが直接サポートしていない独自形式や proprietary なソースからデータを取り込むための、専用ドキュメントローダーを作成します。
「カスタムドキュメントローダーの開発」はCoddyKit上の無料LangChain / RAG / Vector DBsレッスンです。 これはレッスン1/4です。 下記で完全なレッスンを無料で読むことができます。その後、ブラウザ内の組み込みコードエディタと24時間対応のAIチューターでハンズオン演習できます。 これはLangChain / RAG / Vector DBs学習パスの一部であり、ウェブとCoddyKitアプリ全体で進捗が同期されます。 LangChain / RAG / Vector DBsコースには全4レッスンが含まれています。
このレッスンの一部はまだ翻訳されておらず、英語で表示されています。
Why Custom Document Loaders?
LangChain offers many built-in document loaders for common formats like PDFs, web pages, and databases. But what if your data is unique?
Sometimes, you'll encounter:
- Proprietary file formats
- Internal APIs or data sources
- Complex data structures needing custom parsing
This is where custom document loaders shine!
Meet LangChain's BaseLoader
To create your own loader, you'll inherit from LangChain's BaseLoader class. This is an abstract class, meaning it provides a template for what your loader needs to do.
The most important method you'll implement is load(). This method is responsible for fetching your data and transforming it into a list of Document objects.
The LangChain Document Object
All data processed by LangChain, especially for RAG, is standardized into Document objects. Each Document has two main parts:
page_content: The actual text content.metadata: A dictionary of key-value pairs describing the document (e.g., source file, page number, author).
Your custom loader's job is to create these Document objects from your unique data.
Basic Custom Loader Structure
Let's start with a very simple custom loader that just returns a fixed text string as a document. This shows the basic structure of inheriting from BaseLoader and implementing load().
from langchain_core.documents import Document
from langchain_core.document_loaders import BaseLoader
class MySimpleTextLoader(BaseLoader):
def load(self):
content = "This is a custom text document from MySimpleTextLoader."
doc = Document(page_content=content)
return [doc]
# Example usage:
if __name__ == "__main__":
loader = MySimpleTextLoader()
documents = loader.load()
for doc in documents:
print(f"Content: {doc.page_content}")
print(f"Metadata: {doc.metadata}")Making Your Loader Dynamic
A fixed string loader isn't very useful! Real-world loaders need to take parameters, like a file path, a URL, or API credentials.
You can achieve this by adding an __init__ method to your custom loader class. This allows you to pass arguments when you create an instance of your loader.
Loading a 'Custom' Log File
Imagine you have a simple application log file (app.log) where each line is an event. Let's create a custom loader to read this file, treating each line as a separate document.
We'll create a dummy app.log file content directly in the code for simplicity.
from langchain_core.documents import Document
from langchain_core.document_loaders import BaseLoader
# Simulate a log file content
log_file_content = (
"[INFO] User logged in: user123\n"
"[ERROR] Database connection failed\n"
"[DEBUG] Processing request for /api/data\n"
"[INFO] Data retrieved successfully"
)
class CustomLogLoader(BaseLoader):
def __init__(self, log_data: str):
self.log_data = log_data.split('\n')
def load(self):
documents = []
for line in self.log_data:
if line.strip(): # Avoid empty lines
doc = Document(page_content=line)
documents.append(doc)
return documents
# Example usage:
if __name__ == "__main__":
loader = CustomLogLoader(log_file_content)
documents = loader.load()
for i, doc in enumerate(documents):
print(f"Doc {i+1}: {doc.page_content[:40]}...")Adding Rich Metadata
Metadata is incredibly useful! It helps the LLM understand the context of the text and can be used for filtering or improving retrieval. For our log file example, knowing the original log line number or the source file could be very helpful.
You can add any relevant information as key-value pairs to the metadata dictionary of a Document.
Log Loader with Metadata
Let's enhance our CustomLogLoader to include metadata like the original source and the line number for each log entry. This makes the retrieved information much richer!
from langchain_core.documents import Document
from langchain_core.document_loaders import BaseLoader
# Simulate a log file content
log_file_content = (
"[INFO] User logged in: user123\n"
"[ERROR] Database connection failed\n"
"[DEBUG] Processing request for /api/data\n"
"[INFO] Data retrieved successfully"
)
class CustomLogLoaderWithMeta(BaseLoader):
def __init__(self, log_data: str, source_name: str = "app.log"):
self.log_data = log_data.split('\n')
self.source_name = source_name
def load(self):
documents = []
for i, line in enumerate(self.log_data):
if line.strip():
metadata = {
"source": self.source_name,
"line_number": i + 1
}
doc = Document(page_content=line, metadata=metadata)
documents.append(doc)
return documents
# Example usage:
if __name__ == "__main__":
loader = CustomLogLoaderWithMeta(log_file_content, "my_custom_app_logs")
documents = loader.load()
for i, doc in enumerate(documents):
print(f"Doc {i+1}:")
print(f" Content: {doc.page_content[:40]}...")
print(f" Metadata: {doc.metadata}")Integrating Custom Documents
Once your custom loader produces Document objects, they behave just like documents loaded by any other LangChain loader.
You can then pass them into subsequent steps of your RAG pipeline:
- Text Splitting: Break large documents into smaller chunks.
- Embeddings: Convert text chunks into numerical vectors.
- Vector Stores: Store these embeddings for efficient similarity search.
Your custom data is now ready for advanced LLM applications!
Quick Check: Custom Loaders
You're building a custom document loader for LangChain. Which of the following statements is TRUE about the Document object you must return?
Recap: Custom Document Loaders
Great job! You've learned how to develop custom document loaders in LangChain.
- You inherit from
BaseLoaderand implement theload()method. - Your loader converts unique data into a list of
Documentobjects. Documentobjects containpage_contentand a flexiblemetadatadictionary.- Custom loaders are essential for integrating proprietary data sources into your RAG applications.
Next, we'll explore how to integrate custom embedding models!
よくある質問
「カスタムドキュメントローダーの開発」レッスンは無料ですか?
はい。「カスタムドキュメントローダーの開発」の完全なテキストはこのウェブで無料で読めます。インタラクティブに演習し(組み込みコードエディタと24時間対応のAIチューター)、LangChain / RAG / Vector DBsコースの残りをアンロックするには、CoddyKit PROにアップグレードしてください。 LangChain / RAG / Vector DBsコースには全4レッスンが含まれています。
「カスタムドキュメントローダーの開発」で何を学びますか?
LangChainが直接サポートしていない独自形式や proprietary なソースからデータを取り込むための、専用ドキュメントローダーを作成します。 ブラウザで直接実行するハンズオンコードでLangChain / RAG / Vector DBsを演習し、24時間対応のAIチューターがレッスンを進める中での質問に答えます。
LangChain / RAG / Vector DBsを始めるのに経験は必要ですか?
事前経験は必要ありません。CoddyKitのLangChain / RAG / Vector DBsは初級者から上級者向けに構成されているため、ここから始めるか最初から始めて、自分のペースで進むことができます。 これはレッスン1/4です。
「カスタムドキュメントローダーの開発」レッスンにはどのくらい時間がかかりますか?
ほとんどのCoddyKitレッスンは約5~10分かかります。各レッスンはコンパクトでインタラクティブなので、着実に進歩し、ウェブとアプリ全体で正確に前回の場所から再開できます。
このLangChain / RAG / Vector DBsレッスンでコードを書いて実行できますか?
はい。すべてのLangChain / RAG / Vector DBsレッスンに組み込みコードエディタが含まれているため、ブラウザでリアルコードを書いて実行し、即座のAIフィードバックを取得できます。ローカル設定は不要です。
このコースのすべてのレッスン
- カスタムドキュメントローダーの開発
- カスタム埋め込みモデルの統合
- カスタムロジックによる検索チェーンの拡張
- カスタム出力パーサーを作成する