Dzielenie tekstu na fragmenty dla lepszych embeddingów
Dowiedz się, jak dzielić dokumenty na fragmenty, które dobrze poddają się embeddingowi, dlaczego rozmiar fragmentu i overlap mają znaczenie oraz jakie strategie maksymalizują jakość retrievalu.
Dzielenie tekstu na fragmenty dla lepszych embeddingów to bezpłatna lekcja Vector Databases: Pinecone, Weaviate & pgvector na CoddyKit. To lekcja 4 z 4. Możesz przeczytać całą lekcję poniżej za darmo — a potem ćwiczyć ją interaktywnie w przeglądarce z wbudowanym edytorem kodu i tutorem AI dostępnym 24/7. To część ścieżki edukacyjnej Vector Databases: Pinecone, Weaviate & pgvector, a Twój postęp synchronizuje się między webem a aplikacją CoddyKit. Kurs Vector Databases: Pinecone, Weaviate & pgvector zawiera 4 lekcji w sumie.
Części tej lekcji nie zostały jeszcze przetłumaczone i są wyświetlane po angielsku.
Why Chunking Matters
Embedding models have a token limit and produce one vector per input. Feeding a whole document yields a vague, averaged vector. Chunking splits text into focused pieces so each vector captures a specific idea.
The Goldilocks Problem
Chunk size is a balance:
- Too large — diluted meaning, mixed topics in one vector
- Too small — fragments lose context, more vectors to store
Aim for chunks that hold one coherent thought.
Fixed-Size Chunking
The simplest method: split every N characters or tokens.
def chunk(text, size):
return [text[i:i+size] for i in range(0, len(text), size)]
print(chunk('abcdefghij', 4))The Overlap Trick
Fixed splits can cut a sentence in half, losing context at the boundary. Adding overlap repeats the last few tokens of one chunk at the start of the next so ideas spanning a boundary survive.
def chunk_overlap(text, size, overlap):
out = []
i = 0
while i < len(text):
out.append(text[i:i+size])
i += size - overlap
return out
print(chunk_overlap('abcdefghij', 4, 1))Sentence-Aware Chunking
Better than blind character splits: break on sentence boundaries, then group sentences up to a target size. Chunks end cleanly and read coherently.
Recursive Chunking
Recursive splitting tries large separators first (paragraphs), then smaller ones (sentences, then words) until chunks fit the size limit. It respects document structure while guaranteeing size.
Structure-Aware Chunking
For Markdown, code, or HTML, split along structural elements:
- Markdown headers and sections
- Code functions or classes
- HTML sections and lists
This keeps related content together.
Token Counting
Models limit by tokens, not characters. Estimate tokens before embedding so chunks fit the model window.
def approx_tokens(text):
return max(1, len(text) // 4)
print(approx_tokens('The quick brown fox jumps'))Matching Chunks to Queries
Think about how users query. If questions target short facts, smaller chunks improve precision. If questions need broad context, larger chunks help. Sometimes you store both granularities.
Adding Context to Chunks
A chunk taken out of context can be ambiguous. Prepend a short header (document title, section name) to each chunk before embedding so the vector knows where it came from.
Iterate and Measure
There is no universal best chunk size. Pick a starting point (often a few hundred tokens with modest overlap), then measure retrieval quality and adjust. Chunking is an experiment, not a fixed rule.
Quick Check
Test your understanding of chunk overlap.
Recap
You learned that chunking turns documents into focused, embeddable pieces. Balance chunk size, add overlap to preserve boundary context, prefer sentence-aware or recursive splitting, count tokens, enrich chunks with context headers, and iterate while measuring retrieval quality.
Często zadawane pytania
Czy lekcja „Dzielenie tekstu na fragmenty dla lepszych embeddingów” jest bezpłatna?
Tak — pełny tekst „Dzielenie tekstu na fragmenty dla lepszych embeddingów” jest dostępny za darmo tutaj w sieci. Aby ćwiczyć ją interaktywnie (wbudowany edytor kodu i tutor AI dostępny 24/7) i odblokować resztę kursu Vector Databases: Pinecone, Weaviate & pgvector, przejdź na CoddyKit PRO. Kurs Vector Databases: Pinecone, Weaviate & pgvector zawiera 4 lekcji w sumie.
Co nauczysz się w „Dzielenie tekstu na fragmenty dla lepszych embeddingów”?
Dowiedz się, jak dzielić dokumenty na fragmenty, które dobrze poddają się embeddingowi, dlaczego rozmiar fragmentu i overlap mają znaczenie oraz jakie strategie maksymalizują jakość retrievalu. Ćwiczysz Vector Databases: Pinecone, Weaviate & pgvector z praktycznym kodem, który uruchamiasz bezpośrednio w przeglądarce, a tutor AI dostępny 24/7 odpowiada na Twoje pytania podczas pracy nad lekcją.
Czy potrzebuję doświadczenia, aby zacząć Vector Databases: Pinecone, Weaviate & pgvector?
Nie wymagamy żadnego doświadczenia. Vector Databases: Pinecone, Weaviate & pgvector w CoddyKit jest strukturyzowany dla początkujących i zaawansowanych użytkowników, więc możesz zacząć tutaj lub od początku i uczyć się w swoim tempie. To lekcja 4 z 4.
Ile czasu zajmuje lekcja „Dzielenie tekstu na fragmenty dla lepszych embeddingów”?
Większość lekcji CoddyKit trwa około 5–10 minut. Każda lekcja to mały, interaktywny krok, dzięki czemu robisz systematyczne postępy i zawsze wracasz dokładnie do tego samego miejsca — na webie i w aplikacji.
Czy mogę pisać i uruchamiać kod w tej lekcji Vector Databases: Pinecone, Weaviate & pgvector?
Tak. Każda lekcja Vector Databases: Pinecone, Weaviate & pgvector zawiera wbudowany edytor kodu, więc piszesz i uruchamiasz prawdziwy kod bezpośrednio w przeglądarce i od razu otrzymujesz sprzężenie zwrotne od AI — bez konfiguracji na komputerze.
Wszystkie lekcje w tym kursie
- Modele osadzania tekstu
- Korzystanie z interfejsów API osadzania
- Przechowywanie i aktualizowanie osadzeń
- Dzielenie tekstu na fragmenty dla lepszych embeddingów