0Pricing
AI Engineering Academy · Ders

Gömme Benzerliğiyle Anlamsal Parçalama

Ardışık cümleler arasındaki anlamsal uzaklığın en yüksek olduğu noktalarda metni bölen anlamsal parçalamayı uygulayın ve konu açısından tutarlı içeriği birlikte tutun.

Gömme Benzerliğiyle Anlamsal Parçalama, CoddyKit'te ücretsiz bir AI Engineering Academy dersidir. Bu, 4 dersinin 2. dersidir. Aşağıdan dersin tamamını ücretsiz okuyabilir, sonra tarayıcıda yerleşik kod editörü ve 7/24 yapay zeka koçu ile uygulamalı olarak pratik yapabilirsin. Bu, AI Engineering Academy öğrenme yolunun bir parçasıdır ve ilerlemeniz web ve CoddyKit uygulaması arasında senkronize olur. AI Engineering Academy kursu toplamda 4 dersten oluşur.

Bu dersin bazı bölümleri henüz çevrilmemiş olup İngilizce olarak gösterilmektedir.

What Is Semantic Chunking?

Semantic chunking is a technique that splits text at points where the topic changes significantly, rather than at fixed character counts. Instead of asking 'have we hit 500 tokens?', it asks 'does the next sentence belong to the same topic as the current chunk?' — using embedding similarity to answer that question.

The Core Idea: Embedding Distance

The algorithm works by embedding each sentence (or small group of sentences) and computing the cosine similarity between consecutive sentence embeddings. When the similarity drops sharply — meaning the topic has shifted — the algorithm inserts a chunk boundary. Sentences that discuss the same concept stay together in the same chunk.

Step 1: Sentence-Level Embeddings

The first step is to split the document into individual sentences using a sentence tokenizer, then embed each sentence with a fast embedding model. You need sentence-level embeddings — not document-level — so you can detect local topic changes as you move through the text.

from openai import OpenAI
from nltk.tokenize import sent_tokenize
import numpy as np

client = OpenAI()

def embed_sentences(text):
    sentences = sent_tokenize(text)
    response = client.embeddings.create(
        model='text-embedding-3-small',
        input=sentences
    )
    vectors = [item.embedding for item in response.data]
    return sentences, np.array(vectors)

Step 2: Computing Adjacent Similarity

Once you have sentence embeddings, compute the cosine similarity between each consecutive pair: sentence i and sentence i+1. The result is a list of similarity scores, one per sentence boundary. Low scores indicate that the adjacent sentences cover different topics — these are your candidate split points.

def cosine_similarity_adjacent(vectors):
    similarities = []
    for i in range(len(vectors) - 1):
        a = vectors[i]
        b = vectors[i + 1]
        sim = np.dot(a, b) / (np.linalg.norm(a) * np.linalg.norm(b))
        similarities.append(sim)
    return similarities

Step 3: Detecting Breakpoints

A breakpoint is a sentence boundary where the similarity drops below a threshold. You can use a fixed threshold (e.g., 0.6) or a percentile-based threshold that adapts to the document — for example, split whenever similarity falls below the 25th percentile of all similarity scores in that document.

def find_breakpoints(similarities, percentile=25):
    threshold = np.percentile(similarities, percentile)
    breakpoints = []
    for i, sim in enumerate(similarities):
        if sim < threshold:
            breakpoints.append(i + 1)  # split AFTER sentence i
    return breakpoints

Step 4: Assembling Chunks

With breakpoints identified, you can now assemble chunks by joining consecutive sentences between each breakpoint. Each resulting chunk contains a coherent sequence of sentences about the same topic. The chunk boundaries align with natural topic transitions in the original document.

def assemble_chunks(sentences, breakpoints):
    chunks = []
    start = 0
    for bp in breakpoints:
        chunk = ' '.join(sentences[start:bp])
        chunks.append(chunk)
        start = bp
    chunks.append(' '.join(sentences[start:]))  # last chunk
    return chunks

Full Semantic Chunker Example

Putting it all together into a single function: embed sentences, compute adjacent similarities, find breakpoints, assemble chunks. The output is a list of semantically coherent text segments ready to be embedded as whole chunks and stored in your vector database.

def semantic_chunk(text, percentile=25):
    sentences, vectors = embed_sentences(text)
    similarities = cosine_similarity_adjacent(vectors)
    breakpoints = find_breakpoints(similarities, percentile)
    chunks = assemble_chunks(sentences, breakpoints)
    return chunks

chunks = semantic_chunk(my_document)
print(f'Produced {len(chunks)} semantic chunks')
for i, c in enumerate(chunks):
    print(f'Chunk {i+1}: {len(c)} chars')

Choosing the Percentile Threshold

The percentile threshold controls chunk granularity. A low percentile (e.g., 10th) means you only split at major topic shifts — resulting in fewer, longer chunks. A high percentile (e.g., 40th) splits more aggressively — resulting in many small, highly focused chunks. Tune this against your retrieval hit rate evaluation set.

LangChain SemanticChunker

LangChain provides a built-in SemanticChunker that implements this algorithm. It accepts an embedding model and a breakpoint threshold type. This saves you from implementing the algorithm from scratch and integrates directly with LangChain document loaders and vector stores.

from langchain_experimental.text_splitter import SemanticChunker
from langchain_openai import OpenAIEmbeddings

embeddings = OpenAIEmbeddings(model='text-embedding-3-small')

chunker = SemanticChunker(
    embeddings,
    breakpoint_threshold_type='percentile',
    breakpoint_threshold_amount=25
)

chunks = chunker.create_documents([long_document_text])
print(f'{len(chunks)} semantic chunks created')

Trade-offs vs Fixed-Size Chunking

Semantic chunking produces higher-quality chunks with better topic coherence, but it is more expensive: every sentence must be embedded just to determine chunk boundaries — before the chunks are even indexed. For a 100-page document, this means thousands of embedding calls just for chunking. Use semantic chunking when retrieval quality matters more than indexing speed.

When to Use Semantic Chunking

Semantic chunking excels on long-form narrative content — blog posts, research papers, legal documents, and books — where topics shift organically. It is less necessary for highly structured documents like product catalogs, FAQ lists, or code files, which are better handled with document-aware splitters that respect the structure directly.

Quick Check

Test your understanding of semantic chunking from this lesson.

Lesson Recap

In this lesson you learned: semantic chunking uses embedding similarity to detect topic shifts, the percentile threshold controls granularity, and LangChain's SemanticChunker implements this out of the box. Next up we explore parent-child chunking, which combines small precise chunks with large context-rich parent passages.

Sıkça Sorulan Sorular

“Gömme Benzerliğiyle Anlamsal Parçalama” dersi ücretsiz mi?

Evet — “Gömme Benzerliğiyle Anlamsal Parçalama” dersin tüm metni burada web'de ücretsiz olarak okunabilir. Etkileşimli olarak pratik yapmak (yerleşik kod editörü ve 7/24 yapay zeka koçu) ve AI Engineering Academy kursunun geri kalanını açmak için CoddyKit PRO'ya yükselt. AI Engineering Academy kursu toplamda 4 dersten oluşur.

“Gömme Benzerliğiyle Anlamsal Parçalama” dersinde ne öğreneceğim?

Ardışık cümleler arasındaki anlamsal uzaklığın en yüksek olduğu noktalarda metni bölen anlamsal parçalamayı uygulayın ve konu açısından tutarlı içeriği birlikte tutun. AI Engineering Academy ile uygulamalı kodu tarayıcıda doğrudan çalıştırarak pratik yaparsın ve 7/24 yapay zeka koçu dersi çalışırken sorularını yanıtlar.

AI Engineering Academy öğrenmeye başlamak için deneyim gerekli mi?

Önceden deneyim gerekmez. CoddyKit'te AI Engineering Academy, başlangıçtan ileri seviyeye kadar yapılandırıldığı için buradan başlayabilir veya başından başlayıp kendi hızında ilerleme yapabilirsin. Bu, 4 dersinin 2. dersidir.

“Gömme Benzerliğiyle Anlamsal Parçalama” dersi ne kadar sürer?

Çoğu CoddyKit dersi yaklaşık 5–10 dakika sürer. Her biri kısa ve etkileşimli olduğu için sabit ilerleme yaparsın ve web ile uygulama arasında tam olarak bıraktığın yerden devam edebilirsin.

Bu AI Engineering Academy dersinde kod yazıp çalıştırabilir miyim?

Evet. Her AI Engineering Academy dersi yerleşik bir kod editörü içerir, bu sayede tarayıcıda gerçek kod yazıp çalıştırabilir ve anlık yapay zeka geri bildirimi alırsın — yerel kurulum gerekli değildir.

Bu kursun tüm dersleri

  1. Saf Parçalamanın Getirme İşlemini Zayıflatmasının Nedeni
  2. Gömme Benzerliğiyle Anlamsal Parçalama
  3. Üst-Alt ve Küçükten Büyüğe Getirme
  4. Kod ve HTML İçin Belgeye Özel Stratejiler
← AI Engineering Academy Sayfasına Dön