Belge Yükleyicileri Açıklaması
PDF'ler, web sayfaları ve veritabanları gibi çeşitli kaynaklardaki verileri LangChain için kullanılabilir bir biçime nasıl yükleyeceğinizi keşfedin.
Belge Yükleyicileri Açıklaması, CoddyKit'te ücretsiz bir AI Agents with LangChain & Autonomous Workflows dersidir. Bu, 4 dersinin 1. dersidir. Aşağıdan dersin tamamını ücretsiz okuyabilir, sonra tarayıcıda yerleşik kod editörü ve 7/24 yapay zeka koçu ile uygulamalı olarak pratik yapabilirsin. Bu, AI Agents with LangChain & Autonomous Workflows öğrenme yolunun bir parçasıdır ve ilerlemeniz web ve CoddyKit uygulaması arasında senkronize olur. AI Agents with LangChain & Autonomous Workflows kursu toplamda 4 dersten oluşur.
Bu dersin bazı bölümleri henüz çevrilmemiş olup İngilizce olarak gösterilmektedir.
Data for Smarter LLMs
Large Language Models (LLMs) are powerful, but their knowledge is limited to their training data. To make AI agents truly useful, they often need access to fresh, external information.
- Imagine an agent that needs to answer questions about today's news.
- Or one that summarizes a specific document from your company's internal drive.
This is where data loading comes in.
What are Document Loaders?
Document Loaders are LangChain's way of bringing external data into your agent's workflow. They act as bridges, converting raw data from various sources into a standardized format that LLMs can understand.
Think of them as specialized data connectors. They handle the messy details of reading files, fetching web pages, or querying databases, and then present the data cleanly to LangChain.
The LangChain Document Object
When a loader fetches data, it wraps it into a Document object. This is LangChain's universal container for data.
Each Document typically has two main parts:
page_content: The actual text content extracted from the source (e.g., the text from a PDF page, a web article).metadata: A dictionary containing extra information about the document (e.g., source URL, page number, file path, creation date).
Loading from Local Text Files
One of the simplest ways to load data is from a plain text file on your computer. LangChain provides the TextLoader for this purpose.
It takes a file path, reads the content, and creates a Document object. Let's see how it works by creating a temporary file and loading it.
TextLoader in Action
This example creates a temporary text file, writes some content, then uses TextLoader to read it and prints the result.
from langchain_community.document_loaders import TextLoader
import os
import tempfile
if __name__ == "__main__":
# Create a temporary file
with tempfile.NamedTemporaryFile(mode='w', delete=False, encoding='utf-8') as temp_file:
temp_file.write("Hello CoddyKit!\nThis is a test document.")
temp_file_path = temp_file.name
print(f"Created temp file: {temp_file_path}")
try:
# Load the document using TextLoader
loader = TextLoader(temp_file_path)
documents = loader.load()
# Print the content of the loaded document
for doc in documents:
print("--- Document Content ---")
print(doc.page_content)
print("--- Metadata ---")
print(doc.metadata)
finally:
# Clean up the temporary file
os.remove(temp_file_path)
print(f"Cleaned up temp file: {temp_file_path}")Web Content with WebBaseLoader
What if your data is on the internet? The WebBaseLoader is perfect for fetching content directly from web pages.
It's smart enough to extract the main text content, ignoring navigation, ads, and other irrelevant elements. You just provide the URL!
WebBaseLoader Example
Here's how to fetch content from a simple example web page. Note: This loader typically requires beautifulsoup4 and lxml to be installed.
from langchain_community.document_loaders import WebBaseLoader
if __name__ == "__main__":
# Define the URL to load
url = "https://www.google.com/search?q=hello"
print(f"Loading content from: {url}")
# Initialize the WebBaseLoader
loader = WebBaseLoader(url)
# Load the documents
documents = loader.load()
# Print the content of the first loaded document
if documents:
print("--- Extracted Content (first 200 chars) ---")
print(documents[0].page_content[:200] + "...")
print("--- Metadata ---")
print(documents[0].metadata)
else:
print("No documents loaded.")PDFs and Structured Documents
Loading data from PDFs is a common requirement. LangChain offers loaders like PyPDFLoader (which uses the pypdf library) to handle these.
These loaders are designed to extract text from complex formats, often preserving some structure. For PDFs, each page might become a separate Document object, with metadata indicating its page number.
There are also loaders for other formats like CSV, JSON, Markdown, and even specific data structures like Notion databases or Confluence pages.
Beyond Files: Database Loaders
Your data might live in databases. LangChain supports various database loaders to connect directly to your data sources:
- SQL Database Loader: Connects to relational databases (e.g., PostgreSQL, MySQL) to fetch data based on queries.
- MongoDB Loader: For NoSQL document databases.
- Elasticsearch Loader: To pull data from Elasticsearch indices.
These loaders allow your agents to dynamically query and retrieve data from your existing data infrastructure.
Quick Check
Which of the following is NOT a primary purpose of a LangChain Document Loader?
Recap: Document Loaders
In this lesson, we explored LangChain's Document Loaders, essential tools for bringing external data into your AI agents.
- Loaders convert diverse data sources (text files, web pages, PDFs, databases) into LangChain's universal
Documentobject. - Each
Documentcontainspage_contentandmetadata. - We saw practical examples with
TextLoaderfor local files andWebBaseLoaderfor web content.
Next, we'll learn how to handle large documents by splitting them into manageable chunks.
Sıkça Sorulan Sorular
“Belge Yükleyicileri Açıklaması” dersi ücretsiz mi?
Evet — “Belge Yükleyicileri Açıklaması” dersin tüm metni burada web'de ücretsiz olarak okunabilir. Etkileşimli olarak pratik yapmak (yerleşik kod editörü ve 7/24 yapay zeka koçu) ve AI Agents with LangChain & Autonomous Workflows kursunun geri kalanını açmak için CoddyKit PRO'ya yükselt. AI Agents with LangChain & Autonomous Workflows kursu toplamda 4 dersten oluşur.
“Belge Yükleyicileri Açıklaması” dersinde ne öğreneceğim?
PDF'ler, web sayfaları ve veritabanları gibi çeşitli kaynaklardaki verileri LangChain için kullanılabilir bir biçime nasıl yükleyeceğinizi keşfedin. AI Agents with LangChain & Autonomous Workflows ile uygulamalı kodu tarayıcıda doğrudan çalıştırarak pratik yaparsın ve 7/24 yapay zeka koçu dersi çalışırken sorularını yanıtlar.
AI Agents with LangChain & Autonomous Workflows öğrenmeye başlamak için deneyim gerekli mi?
Önceden deneyim gerekmez. CoddyKit'te AI Agents with LangChain & Autonomous Workflows, başlangıçtan ileri seviyeye kadar yapılandırıldığı için buradan başlayabilir veya başından başlayıp kendi hızında ilerleme yapabilirsin. Bu, 4 dersinin 1. dersidir.
“Belge Yükleyicileri Açıklaması” dersi ne kadar sürer?
Çoğu CoddyKit dersi yaklaşık 5–10 dakika sürer. Her biri kısa ve etkileşimli olduğu için sabit ilerleme yaparsın ve web ile uygulama arasında tam olarak bıraktığın yerden devam edebilirsin.
Bu AI Agents with LangChain & Autonomous Workflows dersinde kod yazıp çalıştırabilir miyim?
Evet. Her AI Agents with LangChain & Autonomous Workflows dersi yerleşik bir kod editörü içerir, bu sayede tarayıcıda gerçek kod yazıp çalıştırabilir ve anlık yapay zeka geri bildirimi alırsın — yerel kurulum gerekli değildir.
Bu kursun tüm dersleri
- Belge Yükleyicileri Açıklaması
- Metin Bölücüler ve Gömüler
- Getirme için Vektör Depoları
- Getiriciler ve Bağlamsal Sıkıştırma