文本分割器与嵌入
掌握将大型文档拆分为易于处理的文本块,并生成用于语义搜索的数值嵌入的技术
文本分割器与嵌入 是 CoddyKit 上的免费 AI Agents with LangChain & Autonomous Workflows 课时。 这是第 2 节课,共 4 节。 你可以在下方免费阅读本课时的完整内容 — 然后在浏览器中使用内置代码编辑器和全天候 AI 导师进行实践。 这是 AI Agents with LangChain & Autonomous Workflows 学习路径的一部分,你的进度在网页和 CoddyKit 应用中同步。 AI Agents with LangChain & Autonomous Workflows 课程共包含 4 节课。
本课时的部分内容尚未翻译,以英文显示。
Why Split & Embed Text?
When working with large documents, directly feeding them to a Large Language Model (LLM) often causes problems. LLMs have strict input limits, known as context windows.
This lesson teaches you how to prepare large texts for LLMs using two key techniques: text splitting and embeddings. These are essential for building advanced AI agents.
The Problem: Long Documents
Imagine you have a 100-page PDF document. If you try to ask an LLM a question about it, you can't just send the whole document.
- Context Window Limits: LLMs can only process a certain amount of text at once (e.g., 4,000 to 128,000 tokens).
- Cost: Longer inputs mean higher API costs.
- Relevance: Filling the context window with irrelevant information can make the LLM 'forget' the important parts.
Text splitting solves this by breaking documents into smaller, manageable chunks.
Introducing Text Splitters
LangChain provides various text splitters to divide documents efficiently. Their goal is to keep semantically related pieces of text together while respecting size limits.
Instead of just cutting at arbitrary character counts, smart splitters try to break text at logical points, like paragraphs or sentences.
A common and versatile splitter is the RecursiveCharacterTextSplitter.
Recursive Character Text Splitter
The RecursiveCharacterTextSplitter is a powerful tool. It attempts to split text using a list of characters, trying them in order until the chunks are small enough.
- It starts by trying to split by
\n\n(double newline for paragraphs). - If chunks are still too big, it tries
\n(single newline for lines). - Then spaces, and finally individual characters.
This recursive approach helps maintain semantic coherence.
Code: Basic Splitting Demo
Let's see how RecursiveCharacterTextSplitter works. We'll split a short story into chunks.
from langchain_text_splitters import RecursiveCharacterTextSplitter
story = (
"Alice was beginning to get very tired of sitting by her sister on the bank, "
"and of having nothing to do: once or twice she had peeped into the book her "
"sister was reading, but it had no pictures or conversations in it, 'and what "
"is the use of a book,' thought Alice 'without pictures or conversation?'"
"So she was considering in her own mind (as well as she could, for the hot "
"day made her feel very sleepy and stupid), whether the pleasure of making "
"a daisy-chain would be worth the trouble of getting up and picking the "
"daisies, when suddenly a White Rabbit with pink eyes ran close by her."
)
text_splitter = RecursiveCharacterTextSplitter(
chunk_size=100,
chunk_overlap=0
)
chunks = text_splitter.split_text(story)
for i, chunk in enumerate(chunks):
print(f"Chunk {i+1}: {chunk}\n")Chunk Size & Overlap Explained
Two crucial parameters for text splitting are chunk_size and chunk_overlap:
- Chunk Size: This is the maximum number of characters (or tokens, depending on the splitter) in each chunk. Choose a size that fits within your LLM's context window.
- Chunk Overlap: This specifies how many characters (or tokens) should overlap between consecutive chunks. Overlap helps preserve context across splits, ensuring that important information isn't lost at chunk boundaries.
Finding the right balance for these parameters is key to effective retrieval.
Code: Splitting with Overlap
Let's modify our previous example to use a chunk_overlap. Notice how parts of the text are repeated in adjacent chunks, providing continuity.
from langchain_text_splitters import RecursiveCharacterTextSplitter
story = (
"Alice was beginning to get very tired of sitting by her sister on the bank, "
"and of having nothing to do: once or twice she had peeped into the book her "
"sister was reading, but it had no pictures or conversations in it, 'and what "
"is the use of a book,' thought Alice 'without pictures or conversation?'"
"So she was considering in her own mind (as well as she could, for the hot "
"day made her feel very sleepy and stupid), whether the pleasure of making "
"a daisy-chain would be worth the trouble of getting up and picking the "
"daisies, when suddenly a White Rabbit with pink eyes ran close by her."
)
text_splitter = RecursiveCharacterTextSplitter(
chunk_size=100,
chunk_overlap=20 # Added overlap
)
chunks = text_splitter.split_text(story)
for i, chunk in enumerate(chunks):
print(f"Chunk {i+1}: {chunk}\n")What are Text Embeddings?
Once you've split your documents, how do you find the most relevant chunks for a user's query? This is where text embeddings come in.
- An embedding is a numerical representation (a vector) of text.
- Texts with similar meanings have embeddings that are 'closer' to each other in a high-dimensional space.
- This allows us to perform semantic search: finding chunks that are conceptually similar to a query, not just keyword matches.
Embeddings are the backbone of Retrieval Augmented Generation (RAG).
Generating Embeddings with LangChain
LangChain makes it easy to generate embeddings using various models. You interact with an Embeddings object, which abstracts away the underlying model details.
Popular embedding models include those from OpenAI, Hugging Face, Cohere, and many open-source options like `all-MiniLM-L6-v2`.
You typically initialize an embedding model and then call its embed_query() for a single text or embed_documents() for a list of chunks.
Code: Creating Embeddings
Here's how to generate an embedding for a simple text using OpenAI's embedding model. Remember, you'll need an OpenAI API key for this to run successfully.
import os
# Set your OpenAI API key as an environment variable
# os.environ["OPENAI_API_KEY"] = "YOUR_API_KEY"
from langchain_openai import OpenAIEmbeddings
# Initialize the embedding model
# Requires OPENAI_API_KEY env var or direct pass
embeddings_model = OpenAIEmbeddings()
text_to_embed = "The quick brown fox jumps over the lazy dog."
# Generate the embedding vector
embedding_vector = embeddings_model.embed_query(text_to_embed)
print(f"Original Text: '{text_to_embed}'")
print(f"Embedding Vector (first 5 values): {embedding_vector[:5]}...")
print(f"Vector Dimension: {len(embedding_vector)}")Quick Check: Splitting & Embeddings
Consider the following statements about text splitting and embeddings:
Recap: Splitting & Embedding for RAG
You've learned two fundamental techniques for handling large documents in AI agents:
- Text Splitting: Breaking large texts into smaller, manageable chunks using tools like
RecursiveCharacterTextSplitter, controlled bychunk_sizeandchunk_overlap. - Embeddings: Converting text chunks into numerical vectors using embedding models (e.g.,
OpenAIEmbeddings) to enable semantic similarity search.
These techniques are crucial for building effective Retrieval Augmented Generation (RAG) systems, allowing your agents to intelligently find and use relevant information from vast knowledge bases.
常见问题解答
「文本分割器与嵌入」课时是免费的吗?
是的 — 「文本分割器与嵌入」的完整文本可在网页上免费阅读。要进行交互式练习(内置代码编辑器和全天候 AI 导师)并解锁 AI Agents with LangChain & Autonomous Workflows 课程的其余内容,请升级到 CoddyKit PRO。 AI Agents with LangChain & Autonomous Workflows 课程共包含 4 节课。
「文本分割器与嵌入」这节课中我会学到什么?
掌握将大型文档拆分为易于处理的文本块,并生成用于语义搜索的数值嵌入的技术 你通过在浏览器中直接运行的动手代码来练习 AI Agents with LangChain & Autonomous Workflows,全天候 AI 导师会在你学习这节课的过程中回答你的问题。
学习 AI Agents with LangChain & Autonomous Workflows 需要有经验吗?
无需任何先前经验。CoddyKit 上的 AI Agents with LangChain & Autonomous Workflows 课程适合初学者到高级学习者,你可以从这里开始或从头开始,按照自己的节奏学习。 这是第 2 节课,共 4 节。
「文本分割器与嵌入」课时需要多长时间?
大多数 CoddyKit 课程大约需要 5–10 分钟。每节课都很精短且互动,所以你能稳步进步,并在网页和应用中从离开的地方继续。
我能在这节 AI Agents with LangChain & Autonomous Workflows 课中编写并运行代码吗?
能。每节 AI Agents with LangChain & Autonomous Workflows 课都包含内置代码编辑器,你可以在浏览器中直接编写并运行真实代码,并获得即时 AI 反馈 — 无需本地设置。