埋め込み類似度による意味的チャンク分割
連続する文の間で意味的距離が最大になる位置でテキストを分割し、テーマ上まとまりのある内容を一緒に保つ意味的チャンク分割を実装します。
「埋め込み類似度による意味的チャンク分割」はCoddyKit上の無料AI Engineering Academyレッスンです。 これはレッスン2/4です。 下記で完全なレッスンを無料で読むことができます。その後、ブラウザ内の組み込みコードエディタと24時間対応のAIチューターでハンズオン演習できます。 これはAI Engineering Academy学習パスの一部であり、ウェブとCoddyKitアプリ全体で進捗が同期されます。 AI Engineering Academyコースには全4レッスンが含まれています。
このレッスンの一部はまだ翻訳されておらず、英語で表示されています。
What Is Semantic Chunking?
Semantic chunking is a technique that splits text at points where the topic changes significantly, rather than at fixed character counts. Instead of asking 'have we hit 500 tokens?', it asks 'does the next sentence belong to the same topic as the current chunk?' — using embedding similarity to answer that question.
The Core Idea: Embedding Distance
The algorithm works by embedding each sentence (or small group of sentences) and computing the cosine similarity between consecutive sentence embeddings. When the similarity drops sharply — meaning the topic has shifted — the algorithm inserts a chunk boundary. Sentences that discuss the same concept stay together in the same chunk.
Step 1: Sentence-Level Embeddings
The first step is to split the document into individual sentences using a sentence tokenizer, then embed each sentence with a fast embedding model. You need sentence-level embeddings — not document-level — so you can detect local topic changes as you move through the text.
from openai import OpenAI
from nltk.tokenize import sent_tokenize
import numpy as np
client = OpenAI()
def embed_sentences(text):
sentences = sent_tokenize(text)
response = client.embeddings.create(
model='text-embedding-3-small',
input=sentences
)
vectors = [item.embedding for item in response.data]
return sentences, np.array(vectors)Step 2: Computing Adjacent Similarity
Once you have sentence embeddings, compute the cosine similarity between each consecutive pair: sentence i and sentence i+1. The result is a list of similarity scores, one per sentence boundary. Low scores indicate that the adjacent sentences cover different topics — these are your candidate split points.
def cosine_similarity_adjacent(vectors):
similarities = []
for i in range(len(vectors) - 1):
a = vectors[i]
b = vectors[i + 1]
sim = np.dot(a, b) / (np.linalg.norm(a) * np.linalg.norm(b))
similarities.append(sim)
return similaritiesStep 3: Detecting Breakpoints
A breakpoint is a sentence boundary where the similarity drops below a threshold. You can use a fixed threshold (e.g., 0.6) or a percentile-based threshold that adapts to the document — for example, split whenever similarity falls below the 25th percentile of all similarity scores in that document.
def find_breakpoints(similarities, percentile=25):
threshold = np.percentile(similarities, percentile)
breakpoints = []
for i, sim in enumerate(similarities):
if sim < threshold:
breakpoints.append(i + 1) # split AFTER sentence i
return breakpointsStep 4: Assembling Chunks
With breakpoints identified, you can now assemble chunks by joining consecutive sentences between each breakpoint. Each resulting chunk contains a coherent sequence of sentences about the same topic. The chunk boundaries align with natural topic transitions in the original document.
def assemble_chunks(sentences, breakpoints):
chunks = []
start = 0
for bp in breakpoints:
chunk = ' '.join(sentences[start:bp])
chunks.append(chunk)
start = bp
chunks.append(' '.join(sentences[start:])) # last chunk
return chunksFull Semantic Chunker Example
Putting it all together into a single function: embed sentences, compute adjacent similarities, find breakpoints, assemble chunks. The output is a list of semantically coherent text segments ready to be embedded as whole chunks and stored in your vector database.
def semantic_chunk(text, percentile=25):
sentences, vectors = embed_sentences(text)
similarities = cosine_similarity_adjacent(vectors)
breakpoints = find_breakpoints(similarities, percentile)
chunks = assemble_chunks(sentences, breakpoints)
return chunks
chunks = semantic_chunk(my_document)
print(f'Produced {len(chunks)} semantic chunks')
for i, c in enumerate(chunks):
print(f'Chunk {i+1}: {len(c)} chars')Choosing the Percentile Threshold
The percentile threshold controls chunk granularity. A low percentile (e.g., 10th) means you only split at major topic shifts — resulting in fewer, longer chunks. A high percentile (e.g., 40th) splits more aggressively — resulting in many small, highly focused chunks. Tune this against your retrieval hit rate evaluation set.
LangChain SemanticChunker
LangChain provides a built-in SemanticChunker that implements this algorithm. It accepts an embedding model and a breakpoint threshold type. This saves you from implementing the algorithm from scratch and integrates directly with LangChain document loaders and vector stores.
from langchain_experimental.text_splitter import SemanticChunker
from langchain_openai import OpenAIEmbeddings
embeddings = OpenAIEmbeddings(model='text-embedding-3-small')
chunker = SemanticChunker(
embeddings,
breakpoint_threshold_type='percentile',
breakpoint_threshold_amount=25
)
chunks = chunker.create_documents([long_document_text])
print(f'{len(chunks)} semantic chunks created')Trade-offs vs Fixed-Size Chunking
Semantic chunking produces higher-quality chunks with better topic coherence, but it is more expensive: every sentence must be embedded just to determine chunk boundaries — before the chunks are even indexed. For a 100-page document, this means thousands of embedding calls just for chunking. Use semantic chunking when retrieval quality matters more than indexing speed.
When to Use Semantic Chunking
Semantic chunking excels on long-form narrative content — blog posts, research papers, legal documents, and books — where topics shift organically. It is less necessary for highly structured documents like product catalogs, FAQ lists, or code files, which are better handled with document-aware splitters that respect the structure directly.
Quick Check
Test your understanding of semantic chunking from this lesson.
Lesson Recap
In this lesson you learned: semantic chunking uses embedding similarity to detect topic shifts, the percentile threshold controls granularity, and LangChain's SemanticChunker implements this out of the box. Next up we explore parent-child chunking, which combines small precise chunks with large context-rich parent passages.
よくある質問
「埋め込み類似度による意味的チャンク分割」レッスンは無料ですか?
はい。「埋め込み類似度による意味的チャンク分割」の完全なテキストはこのウェブで無料で読めます。インタラクティブに演習し(組み込みコードエディタと24時間対応のAIチューター)、AI Engineering Academyコースの残りをアンロックするには、CoddyKit PROにアップグレードしてください。 AI Engineering Academyコースには全4レッスンが含まれています。
「埋め込み類似度による意味的チャンク分割」で何を学びますか?
連続する文の間で意味的距離が最大になる位置でテキストを分割し、テーマ上まとまりのある内容を一緒に保つ意味的チャンク分割を実装します。 ブラウザで直接実行するハンズオンコードでAI Engineering Academyを演習し、24時間対応のAIチューターがレッスンを進める中での質問に答えます。
AI Engineering Academyを始めるのに経験は必要ですか?
事前経験は必要ありません。CoddyKitのAI Engineering Academyは初級者から上級者向けに構成されているため、ここから始めるか最初から始めて、自分のペースで進むことができます。 これはレッスン2/4です。
「埋め込み類似度による意味的チャンク分割」レッスンにはどのくらい時間がかかりますか?
ほとんどのCoddyKitレッスンは約5~10分かかります。各レッスンはコンパクトでインタラクティブなので、着実に進歩し、ウェブとアプリ全体で正確に前回の場所から再開できます。
このAI Engineering Academyレッスンでコードを書いて実行できますか?
はい。すべてのAI Engineering Academyレッスンに組み込みコードエディタが含まれているため、ブラウザでリアルコードを書いて実行し、即座のAIフィードバックを取得できます。ローカル設定は不要です。
このコースのすべてのレッスン
- 素朴なチャンク分割が検索に悪影響を与える理由
- 埋め込み類似度による意味的チャンク分割
- 親子チャンクと小から大への検索
- コードとHTMLに特化したドキュメント戦略