Analizadores, tokenizadores y el índice invertido
Aprenda cómo Elasticsearch convierte el texto en tokens consultables mediante analizadores y los almacena en un índice invertido.
Analizadores, tokenizadores y el índice invertido es una lección gratuita de Elasticsearch & Full Text Search Systems en CoddyKit. Esta es la lección 4 de 4. Puedes leer la lección completa abajo gratuitamente — luego la practicas en el navegador con un editor de código integrado y un tutor de IA 24/7. Forma parte de la ruta de aprendizaje de Elasticsearch & Full Text Search Systems, y tu progreso se sincroniza en la web y la app de CoddyKit. El curso de Elasticsearch & Full Text Search Systems incluye 4 lecciones en total.
Partes de esta lección aún no han sido traducidas y se muestran en inglés.
From Text to Tokens
Before text is searchable, an analyzer breaks it into tokens and stores them in an inverted index. That's the first step of every search.
The Inverted Index
An inverted index maps each token to the documents that contain it — like a book's index — making full-text lookups extremely fast.
Anatomy of an Analyzer
An analyzer runs three stages in order: character filters clean raw text, a tokenizer splits it into tokens, and token filters reshape them.
The Standard Analyzer
The default standard analyzer splits on word boundaries and lowercases, so Quick Brown Fox! becomes quick, brown, fox.
Testing with _analyze
The _analyze API shows exactly which tokens an analyzer produces from your text — run it to see the output.
POST /_analyze
{
"analyzer": "standard",
"text": "The Quick Brown Foxes"
}Stemming
A stemming filter reduces words to their root, so running, runs, and ran all map to run — and searching one finds the others.
Stop Words
A stop filter removes low-value words like the, a, and is, shrinking the index and improving relevance.
Custom Analyzer
Define a custom analyzer in index settings by combining a tokenizer with filters — the code wires up lowercase, stop words, and stemming.
PUT /articles
{
"settings": { "analysis": { "analyzer": {
"my_english": {
"tokenizer": "standard",
"filter": ["lowercase", "english_stop", "english_stemmer"]
}}}}
}text vs keyword
A text field is analyzed for full-text search; a keyword field stays one exact token for filtering, sorting, and aggregations.
Index vs Search Time
Analysis runs at index time when storing a document and again at search time on the query — usually the same analyzer, so tokens match.
Multi-field Mapping
A handy trick: map a string as text for search and add a .keyword sub-field for exact matches and aggregations at once.
"title": { "type": "text",
"fields": { "raw": { "type": "keyword" } } }Quick Check
Why are text and keyword fields treated differently?
Recap
Recap: analyzers (filters, tokenizer, filters) build the inverted index, stemming and stop words refine it, and text vs keyword shapes how fields behave.
Preguntas frecuentes
¿La lección «Analizadores, tokenizadores y el índice invertido» es gratis?
Sí — el texto completo de «Analizadores, tokenizadores y el índice invertido» es gratis para leer aquí en la web. Para practicarla de forma interactiva (editor de código integrado y tutor de IA 24/7) y desbloquear el resto del curso de Elasticsearch & Full Text Search Systems, actualiza a CoddyKit PRO. El curso de Elasticsearch & Full Text Search Systems incluye 4 lecciones en total.
¿Qué aprenderé en «Analizadores, tokenizadores y el índice invertido»?
Aprenda cómo Elasticsearch convierte el texto en tokens consultables mediante analizadores y los almacena en un índice invertido. Practicas Elasticsearch & Full Text Search Systems con código real que ejecutas directamente en el navegador, y un tutor de IA 24/7 responde tus preguntas mientras trabajas en la lección.
¿Necesito experiencia previa para empezar Elasticsearch & Full Text Search Systems?
No se requiere experiencia previa. Elasticsearch & Full Text Search Systems en CoddyKit está estructurado para principiantes hasta estudiantes avanzados, así que puedes empezar aquí o desde el inicio y avanzar a tu ritmo. Esta es la lección 4 de 4.
¿Cuánto tiempo toma la lección «Analizadores, tokenizadores y el índice invertido»?
La mayoría de las lecciones de CoddyKit toman alrededor de 5–10 minutos. Cada una es compacta e interactiva, así que avanzas constantemente y retomas exactamente por donde dejaste en la web y la app.
¿Puedo escribir y ejecutar código en esta lección de Elasticsearch & Full Text Search Systems?
Sí. Cada lección de Elasticsearch & Full Text Search Systems incluye un editor de código integrado, así que escribes y ejecutas código real directamente en tu navegador y obtienes retroalimentación instantánea de IA — sin configuración local necesaria.
Todas las lecciones de este curso
- ¿Qué es la búsqueda de texto completo?
- Conceptos básicos de Elasticsearch
- Configuración de su primer clúster
- Analizadores, tokenizadores y el índice invertido