Индексация документов в Elasticsearch
Разберитесь, как индексировать один или несколько документов в индексе Elasticsearch, включая автоматическую генерацию идентификаторов и пользовательские идентификаторы.
«Индексация документов в Elasticsearch» — бесплатный урок Elasticsearch & Full Text Search Systems на CoddyKit. Это урок 1 из 4. Ты можешь прочитать весь урок бесплатно ниже — а потом практиковать его прямо в браузере с встроенным редактором кода и ИИ-репетитором 24/7. Это часть пути обучения Elasticsearch & Full Text Search Systems, и твой прогресс синхронизируется между веб-версией и приложением CoddyKit. Курс Elasticsearch & Full Text Search Systems содержит 4 уроков всего.
Части этого урока еще не переведены и отображаются на английском.
What is Indexing?
Welcome to indexing! In Elasticsearch, indexing is the process of storing data into an index to make it searchable.
Think of it like adding a new book to a library's catalog. You provide the book's details, and the library stores them in a way that makes the book easy to find later.
Documents & Indices Refresher
Before we dive in, let's quickly recap two core concepts:
- Document: A basic unit of information in Elasticsearch, similar to a row in a traditional database. It's usually a JSON object.
- Index: A collection of documents that have similar characteristics. It's like a database in a relational world.
When you index, you add a document to an index.
The Index API
You interact with Elasticsearch using its REST API. To index a document, you'll typically use HTTP POST or PUT requests.
POST /<index>/_doc: Used to index a document, often letting Elasticsearch generate an ID.PUT /<index>/_doc/<id>: Used to index a document with a specific, user-provided ID.
Let's see them in action!
Auto-Generated IDs
The simplest way to index is to let Elasticsearch generate a unique ID for your document. You use the POST method to the _doc endpoint without specifying an ID.
Here's an example using curl to index a document into an index named products:
curl -X POST "localhost:9200/products/_doc?pretty" \
-H 'Content-Type: application/json' \
-d'{"name": "Laptop", "price": 1200}'Understanding Auto IDs
After the previous POST request, Elasticsearch would return a response including a unique _id for your document, like "_id": "AbCdEfGhIjKlMnOpQrSt".
When should you use auto-generated IDs?
- When you don't have a natural unique identifier for your data.
- For logs or temporary data where a unique ID isn't critical for external reference.
- When you want to guarantee a new document is always created.
Indexing with Custom IDs
Often, your data already has a unique identifier from another system (e.g., a database primary key). In such cases, you can provide your own ID using the PUT method.
The ID is specified directly in the URL path: /<index>/_doc/<your_id>.
curl -X PUT "localhost:9200/products/_doc/prod_101?pretty" \
-H 'Content-Type: application/json' \
-d'{"name": "Smartphone", "price": 800}'Why Use Custom IDs?
Using custom IDs offers several advantages:
- Integration: Easily map Elasticsearch documents to records in an external database.
- Predictability: You know the document's ID beforehand.
- Updates: It makes updating specific documents straightforward, as you always refer to them by their known ID.
Idempotency with PUT
A key concept when using PUT with a custom ID is idempotency. This means that performing the same operation multiple times will produce the same result as performing it once.
- If a document with the specified ID already exists,
PUTwill update it. - If it doesn't exist,
PUTwill create it.
This is different from POST, which always creates a *new* document with a new ID.
Indexing Many Documents
While indexing documents one-by-one is fine for small numbers, it can be inefficient for large datasets due to network overhead.
Elasticsearch provides a powerful _bulk API that allows you to perform multiple index, update, or delete operations in a single request. This dramatically improves indexing performance.
We'll explore the _bulk API in more detail in a future lesson!
Indexing Method Check
Imagine you have a new set of sensor readings. Each reading is unique, and you don't have a predefined ID for them, but you want to store them in Elasticsearch to be searchable.
Recap: Indexing Essentials
Great job! In this lesson, you learned the fundamentals of indexing documents into Elasticsearch:
- What indexing means and its role in making data searchable.
- The difference between documents and indices.
- How to use
POST /<index>/_docto index documents with auto-generated IDs. - How to use
PUT /<index>/_doc/<id>to index documents with custom IDs. - The concept of idempotency when using
PUT. - A brief introduction to the efficiency of bulk indexing.
Next, we'll explore more operations on these documents!
Часто задаваемые вопросы
Урок «Индексация документов в Elasticsearch» бесплатный?
Да — полный текст урока «Индексация документов в Elasticsearch» бесплатно доступен здесь в веб-версии. Чтобы практиковать его интерактивно (встроенный редактор кода и ИИ-репетитор 24/7) и разблокировать остальной курс Elasticsearch & Full Text Search Systems, подпишись на CoddyKit PRO. Курс Elasticsearch & Full Text Search Systems содержит 4 уроков всего.
Чему я научусь в уроке «Индексация документов в Elasticsearch»?
Разберитесь, как индексировать один или несколько документов в индексе Elasticsearch, включая автоматическую генерацию идентификаторов и пользовательские идентификаторы. Ты практикуешь Elasticsearch & Full Text Search Systems с помощью реального кода, который запускаешь прямо в браузере, и ИИ-репетитор 24/7 отвечает на твои вопросы во время урока.
Нужен ли мне опыт, чтобы начать Elasticsearch & Full Text Search Systems?
Предыдущий опыт не требуется. Elasticsearch & Full Text Search Systems на CoddyKit структурирован для всех уровней — от новичков до продвинутых, поэтому ты можешь начать отсюда или с самого начала и учиться в своем темпе. Это урок 1 из 4.
Сколько времени занимает урок «Индексация документов в Elasticsearch»?
Большинство уроков CoddyKit занимают около 5–10 минут. Каждый из них компактный и интерактивный, поэтому ты постоянно делаешь прогресс и продолжаешь с того же места в веб-версии и приложении.
Можно ли писать и запускать код в этом уроке Elasticsearch & Full Text Search Systems?
Да. Каждый урок Elasticsearch & Full Text Search Systems включает встроенный редактор кода, поэтому ты пишешь и запускаешь реальный код прямо в браузере и получаешь моментальную обратную связь от AI — локальная установка не требуется.
Все уроки этого курса
- Индексация документов в Elasticsearch
- Операции CRUD с документами
- Основы отображения и типы данных
- Массовая индексация и Bulk API