0Pricing
AI Engineering Academy · Aula

Divisão semântica com similaridade de embeddings

Implemente uma divisão semântica que separe o texto nos pontos de maior distância semântica entre frases consecutivas, mantendo juntos os conteúdos tematicamente coerentes.

Divisão semântica com similaridade de embeddings é uma aula grátis de AI Engineering Academy no CoddyKit. Esta é a aula 2 de 4. Você pode ler a aula completa abaixo gratuitamente — depois pratica ao vivo no navegador com um editor de código integrado e um tutor de IA 24/7. Faz parte do caminho de aprendizado de AI Engineering Academy, e seu progresso é sincronizado entre a web e o app CoddyKit. O curso de AI Engineering Academy inclui 4 aulas no total.

Partes desta aula ainda não foram traduzidas e aparecem em inglês.

What Is Semantic Chunking?

Semantic chunking is a technique that splits text at points where the topic changes significantly, rather than at fixed character counts. Instead of asking 'have we hit 500 tokens?', it asks 'does the next sentence belong to the same topic as the current chunk?' — using embedding similarity to answer that question.

The Core Idea: Embedding Distance

The algorithm works by embedding each sentence (or small group of sentences) and computing the cosine similarity between consecutive sentence embeddings. When the similarity drops sharply — meaning the topic has shifted — the algorithm inserts a chunk boundary. Sentences that discuss the same concept stay together in the same chunk.

Step 1: Sentence-Level Embeddings

The first step is to split the document into individual sentences using a sentence tokenizer, then embed each sentence with a fast embedding model. You need sentence-level embeddings — not document-level — so you can detect local topic changes as you move through the text.

from openai import OpenAI
from nltk.tokenize import sent_tokenize
import numpy as np

client = OpenAI()

def embed_sentences(text):
    sentences = sent_tokenize(text)
    response = client.embeddings.create(
        model='text-embedding-3-small',
        input=sentences
    )
    vectors = [item.embedding for item in response.data]
    return sentences, np.array(vectors)

Step 2: Computing Adjacent Similarity

Once you have sentence embeddings, compute the cosine similarity between each consecutive pair: sentence i and sentence i+1. The result is a list of similarity scores, one per sentence boundary. Low scores indicate that the adjacent sentences cover different topics — these are your candidate split points.

def cosine_similarity_adjacent(vectors):
    similarities = []
    for i in range(len(vectors) - 1):
        a = vectors[i]
        b = vectors[i + 1]
        sim = np.dot(a, b) / (np.linalg.norm(a) * np.linalg.norm(b))
        similarities.append(sim)
    return similarities

Step 3: Detecting Breakpoints

A breakpoint is a sentence boundary where the similarity drops below a threshold. You can use a fixed threshold (e.g., 0.6) or a percentile-based threshold that adapts to the document — for example, split whenever similarity falls below the 25th percentile of all similarity scores in that document.

def find_breakpoints(similarities, percentile=25):
    threshold = np.percentile(similarities, percentile)
    breakpoints = []
    for i, sim in enumerate(similarities):
        if sim < threshold:
            breakpoints.append(i + 1)  # split AFTER sentence i
    return breakpoints

Step 4: Assembling Chunks

With breakpoints identified, you can now assemble chunks by joining consecutive sentences between each breakpoint. Each resulting chunk contains a coherent sequence of sentences about the same topic. The chunk boundaries align with natural topic transitions in the original document.

def assemble_chunks(sentences, breakpoints):
    chunks = []
    start = 0
    for bp in breakpoints:
        chunk = ' '.join(sentences[start:bp])
        chunks.append(chunk)
        start = bp
    chunks.append(' '.join(sentences[start:]))  # last chunk
    return chunks

Full Semantic Chunker Example

Putting it all together into a single function: embed sentences, compute adjacent similarities, find breakpoints, assemble chunks. The output is a list of semantically coherent text segments ready to be embedded as whole chunks and stored in your vector database.

def semantic_chunk(text, percentile=25):
    sentences, vectors = embed_sentences(text)
    similarities = cosine_similarity_adjacent(vectors)
    breakpoints = find_breakpoints(similarities, percentile)
    chunks = assemble_chunks(sentences, breakpoints)
    return chunks

chunks = semantic_chunk(my_document)
print(f'Produced {len(chunks)} semantic chunks')
for i, c in enumerate(chunks):
    print(f'Chunk {i+1}: {len(c)} chars')

Choosing the Percentile Threshold

The percentile threshold controls chunk granularity. A low percentile (e.g., 10th) means you only split at major topic shifts — resulting in fewer, longer chunks. A high percentile (e.g., 40th) splits more aggressively — resulting in many small, highly focused chunks. Tune this against your retrieval hit rate evaluation set.

LangChain SemanticChunker

LangChain provides a built-in SemanticChunker that implements this algorithm. It accepts an embedding model and a breakpoint threshold type. This saves you from implementing the algorithm from scratch and integrates directly with LangChain document loaders and vector stores.

from langchain_experimental.text_splitter import SemanticChunker
from langchain_openai import OpenAIEmbeddings

embeddings = OpenAIEmbeddings(model='text-embedding-3-small')

chunker = SemanticChunker(
    embeddings,
    breakpoint_threshold_type='percentile',
    breakpoint_threshold_amount=25
)

chunks = chunker.create_documents([long_document_text])
print(f'{len(chunks)} semantic chunks created')

Trade-offs vs Fixed-Size Chunking

Semantic chunking produces higher-quality chunks with better topic coherence, but it is more expensive: every sentence must be embedded just to determine chunk boundaries — before the chunks are even indexed. For a 100-page document, this means thousands of embedding calls just for chunking. Use semantic chunking when retrieval quality matters more than indexing speed.

When to Use Semantic Chunking

Semantic chunking excels on long-form narrative content — blog posts, research papers, legal documents, and books — where topics shift organically. It is less necessary for highly structured documents like product catalogs, FAQ lists, or code files, which are better handled with document-aware splitters that respect the structure directly.

Quick Check

Test your understanding of semantic chunking from this lesson.

Lesson Recap

In this lesson you learned: semantic chunking uses embedding similarity to detect topic shifts, the percentile threshold controls granularity, and LangChain's SemanticChunker implements this out of the box. Next up we explore parent-child chunking, which combines small precise chunks with large context-rich parent passages.

Perguntas Frequentes

A aula “Divisão semântica com similaridade de embeddings” é grátis?

Sim — o texto completo de “Divisão semântica com similaridade de embeddings” é grátis para ler aqui na web. Para praticá-la interativamente (um editor de código integrado e um tutor de IA 24/7) e desbloquear o restante do curso de AI Engineering Academy, atualize para CoddyKit PRO. O curso de AI Engineering Academy inclui 4 aulas no total.

O que vou aprender em “Divisão semântica com similaridade de embeddings”?

Implemente uma divisão semântica que separe o texto nos pontos de maior distância semântica entre frases consecutivas, mantendo juntos os conteúdos tematicamente coerentes. Você pratica AI Engineering Academy com código prático que executa diretamente no navegador, e um tutor de IA 24/7 responde suas dúvidas enquanto trabalha na aula.

Preciso ter experiência prévia para começar AI Engineering Academy?

Nenhuma experiência prévia é necessária. AI Engineering Academy no CoddyKit é estruturado para alunos iniciantes até avançados, então você pode começar aqui ou desde o início e aprender no seu ritmo. Esta é a aula 2 de 4.

Quanto tempo leva a aula “Divisão semântica com similaridade de embeddings”?

A maioria das aulas CoddyKit leva cerca de 5–10 minutos. Cada uma é compacta e interativa, então você faz progresso constante e retoma exatamente de onde parou entre web e app.

Posso escrever e executar código nesta aula de AI Engineering Academy?

Sim. Cada aula de AI Engineering Academy inclui um editor de código integrado, então você escreve e executa código real direto no navegador e recebe feedback de IA instantaneamente — nenhuma configuração local necessária.

Todas as aulas deste curso

  1. Por que a divisão ingênua prejudica a recuperação
  2. Divisão semântica com similaridade de embeddings
  3. Recuperação pai-filho e do pequeno para o grande
  4. Estratégias específicas para código e HTML
← Voltar para AI Engineering Academy