Découpage sémantique avec la similarité des plongements
Implémentez un découpage sémantique qui sépare le texte aux points de distance sémantique maximale entre des phrases consécutives, tout en conservant ensemble les contenus cohérents sur le plan thématique.
Découpage sémantique avec la similarité des plongements est une leçon AI Engineering Academy gratuite sur CoddyKit. Ceci est la leçon 2 sur 4. Tu peux lire la leçon complète ci-dessous gratuitement — puis la pratiquer en direct dans le navigateur avec un éditeur de code intégré et un tuteur IA 24/7. Elle fait partie du parcours d'apprentissage AI Engineering Academy, et ta progression se synchronise sur le web et l'application CoddyKit. Le cours AI Engineering Academy comprend 4 leçons au total.
Certaines parties de cette leçon n'ont pas encore été traduites et s'affichent en anglais.
What Is Semantic Chunking?
Semantic chunking is a technique that splits text at points where the topic changes significantly, rather than at fixed character counts. Instead of asking 'have we hit 500 tokens?', it asks 'does the next sentence belong to the same topic as the current chunk?' — using embedding similarity to answer that question.
The Core Idea: Embedding Distance
The algorithm works by embedding each sentence (or small group of sentences) and computing the cosine similarity between consecutive sentence embeddings. When the similarity drops sharply — meaning the topic has shifted — the algorithm inserts a chunk boundary. Sentences that discuss the same concept stay together in the same chunk.
Step 1: Sentence-Level Embeddings
The first step is to split the document into individual sentences using a sentence tokenizer, then embed each sentence with a fast embedding model. You need sentence-level embeddings — not document-level — so you can detect local topic changes as you move through the text.
from openai import OpenAI
from nltk.tokenize import sent_tokenize
import numpy as np
client = OpenAI()
def embed_sentences(text):
sentences = sent_tokenize(text)
response = client.embeddings.create(
model='text-embedding-3-small',
input=sentences
)
vectors = [item.embedding for item in response.data]
return sentences, np.array(vectors)Step 2: Computing Adjacent Similarity
Once you have sentence embeddings, compute the cosine similarity between each consecutive pair: sentence i and sentence i+1. The result is a list of similarity scores, one per sentence boundary. Low scores indicate that the adjacent sentences cover different topics — these are your candidate split points.
def cosine_similarity_adjacent(vectors):
similarities = []
for i in range(len(vectors) - 1):
a = vectors[i]
b = vectors[i + 1]
sim = np.dot(a, b) / (np.linalg.norm(a) * np.linalg.norm(b))
similarities.append(sim)
return similaritiesStep 3: Detecting Breakpoints
A breakpoint is a sentence boundary where the similarity drops below a threshold. You can use a fixed threshold (e.g., 0.6) or a percentile-based threshold that adapts to the document — for example, split whenever similarity falls below the 25th percentile of all similarity scores in that document.
def find_breakpoints(similarities, percentile=25):
threshold = np.percentile(similarities, percentile)
breakpoints = []
for i, sim in enumerate(similarities):
if sim < threshold:
breakpoints.append(i + 1) # split AFTER sentence i
return breakpointsStep 4: Assembling Chunks
With breakpoints identified, you can now assemble chunks by joining consecutive sentences between each breakpoint. Each resulting chunk contains a coherent sequence of sentences about the same topic. The chunk boundaries align with natural topic transitions in the original document.
def assemble_chunks(sentences, breakpoints):
chunks = []
start = 0
for bp in breakpoints:
chunk = ' '.join(sentences[start:bp])
chunks.append(chunk)
start = bp
chunks.append(' '.join(sentences[start:])) # last chunk
return chunksFull Semantic Chunker Example
Putting it all together into a single function: embed sentences, compute adjacent similarities, find breakpoints, assemble chunks. The output is a list of semantically coherent text segments ready to be embedded as whole chunks and stored in your vector database.
def semantic_chunk(text, percentile=25):
sentences, vectors = embed_sentences(text)
similarities = cosine_similarity_adjacent(vectors)
breakpoints = find_breakpoints(similarities, percentile)
chunks = assemble_chunks(sentences, breakpoints)
return chunks
chunks = semantic_chunk(my_document)
print(f'Produced {len(chunks)} semantic chunks')
for i, c in enumerate(chunks):
print(f'Chunk {i+1}: {len(c)} chars')Choosing the Percentile Threshold
The percentile threshold controls chunk granularity. A low percentile (e.g., 10th) means you only split at major topic shifts — resulting in fewer, longer chunks. A high percentile (e.g., 40th) splits more aggressively — resulting in many small, highly focused chunks. Tune this against your retrieval hit rate evaluation set.
LangChain SemanticChunker
LangChain provides a built-in SemanticChunker that implements this algorithm. It accepts an embedding model and a breakpoint threshold type. This saves you from implementing the algorithm from scratch and integrates directly with LangChain document loaders and vector stores.
from langchain_experimental.text_splitter import SemanticChunker
from langchain_openai import OpenAIEmbeddings
embeddings = OpenAIEmbeddings(model='text-embedding-3-small')
chunker = SemanticChunker(
embeddings,
breakpoint_threshold_type='percentile',
breakpoint_threshold_amount=25
)
chunks = chunker.create_documents([long_document_text])
print(f'{len(chunks)} semantic chunks created')Trade-offs vs Fixed-Size Chunking
Semantic chunking produces higher-quality chunks with better topic coherence, but it is more expensive: every sentence must be embedded just to determine chunk boundaries — before the chunks are even indexed. For a 100-page document, this means thousands of embedding calls just for chunking. Use semantic chunking when retrieval quality matters more than indexing speed.
When to Use Semantic Chunking
Semantic chunking excels on long-form narrative content — blog posts, research papers, legal documents, and books — where topics shift organically. It is less necessary for highly structured documents like product catalogs, FAQ lists, or code files, which are better handled with document-aware splitters that respect the structure directly.
Quick Check
Test your understanding of semantic chunking from this lesson.
Lesson Recap
In this lesson you learned: semantic chunking uses embedding similarity to detect topic shifts, the percentile threshold controls granularity, and LangChain's SemanticChunker implements this out of the box. Next up we explore parent-child chunking, which combines small precise chunks with large context-rich parent passages.
Questions Fréquemment Posées
La leçon « Découpage sémantique avec la similarité des plongements » est-elle gratuite ?
Oui — le texte complet de « Découpage sémantique avec la similarité des plongements » est gratuit à lire ici sur le web. Pour la pratiquer de manière interactive (un éditeur de code intégré et un tuteur IA 24/7) et déverrouiller le reste du cours AI Engineering Academy, passe à CoddyKit PRO. Le cours AI Engineering Academy comprend 4 leçons au total.
Qu'est-ce que j'apprendrai dans « Découpage sémantique avec la similarité des plongements » ?
Implémentez un découpage sémantique qui sépare le texte aux points de distance sémantique maximale entre des phrases consécutives, tout en conservant ensemble les contenus cohérents sur le plan théma… Tu pratiques AI Engineering Academy avec du code pratique que tu exécutes directement dans le navigateur, et un tuteur IA 24/7 répond à tes questions au fur et à mesure que tu avances dans la leçon.
Dois-je avoir de l'expérience pour commencer AI Engineering Academy ?
Aucune expérience préalable n'est requise. AI Engineering Academy sur CoddyKit est structuré pour les débutants jusqu'aux apprenants avancés, donc tu peux commencer ici ou depuis le début et avancer à ton rythme. Ceci est la leçon 2 sur 4.
Combien de temps prend la leçon « Découpage sémantique avec la similarité des plongements » ?
La plupart des leçons CoddyKit prennent environ 5–10 minutes. Chacune est courte et interactive, tu progresses régulièrement et tu repiques exactement où tu t'es arrêté sur le web et l'app.
Peux-tu écrire et exécuter du code dans cette leçon AI Engineering Academy ?
Oui. Chaque leçon AI Engineering Academy inclut un éditeur de code intégré, tu écris et exécutes du vrai code directement dans ton navigateur et tu reçois des retours IA instantanés — aucune configuration locale requise.
Toutes les leçons de ce cours
- Pourquoi le découpage naïf nuit à la recherche
- Découpage sémantique avec la similarité des plongements
- Recherche parent-enfant et du petit vers le grand
- Stratégies propres aux documents pour le code et HTML