0Pricing
Elasticsearch & Full Text Search Systems · Lección

Gestión de datos de series temporales

Optimice Elasticsearch para datos de series temporales, incluidos los flujos de datos, ILM (Index Lifecycle Management) y las arquitecturas hot-warm-cold.

Gestión de datos de series temporales es una lección gratuita de Elasticsearch & Full Text Search Systems en CoddyKit. Esta es la lección 2 de 4. Puedes leer la lección completa abajo gratuitamente — luego la practicas en el navegador con un editor de código integrado y un tutor de IA 24/7. Forma parte de la ruta de aprendizaje de Elasticsearch & Full Text Search Systems, y tu progreso se sincroniza en la web y la app de CoddyKit. El curso de Elasticsearch & Full Text Search Systems incluye 4 lecciones en total.

Partes de esta lección aún no han sido traducidas y se muestran en inglés.

What is Time-Series Data?

Time-series data is information collected over a period of time, often at regular intervals. Think of it as a sequence of data points indexed by time.

Examples include:

  • System logs: Events happening on a server.
  • IoT sensor readings: Temperature, humidity from devices.
  • Financial data: Stock prices over days or hours.

This data is typically append-only and usually immutable once recorded.

Why Elasticsearch for Time-Series?

Elasticsearch is an excellent choice for managing time-series data due to its:

  • Scalability: Handles massive volumes of data.
  • Speed: Fast indexing and search capabilities.
  • Analytics: Powerful aggregations for insights.
  • Flexibility: Can store structured and unstructured data.

It's particularly strong for use cases like log analytics, metric monitoring, and security event management.

Introducing Data Streams

Data streams are a core feature in Elasticsearch designed specifically for time-series data. They simplify management by automatically rolling over to new indices as needed.

When you index data into a data stream, it always writes to the current write index. Queries, however, search across all backing indices transparently. This makes managing large, continuously growing datasets much easier.

Creating Your First Data Stream

To create a data stream, you first need an index template that defines its structure and points to a data stream.

Then, simply indexing a document into the stream's name will create it automatically based on the template:

PUT _index_template/my-data-stream-template
{
  "index_patterns": ["my-data-stream-*"],
  "data_stream": {},
  "priority": 200,
  "template": {
    "settings": {
      "number_of_shards": 1
    },
    "mappings": {
      "properties": {
        "@timestamp": {
          "type": "date",
          "format": "strict_date_optional_time||epoch_millis"
        },
        "message": { "type": "text" }
      }
    }
  }
}

// Indexing a document creates the stream
POST my-data-stream/_doc
{
  "@timestamp": "2023-10-27T10:00:00Z",
  "message": "Sensor reading: 25.5C"
}

Automating with ILM

Index Lifecycle Management (ILM) is a powerful feature that automates the management of your indices through various phases of their life.

For time-series data, ILM helps you:

  • Automatically roll over indices when they reach a certain size or age.
  • Move older, less frequently accessed data to cheaper storage.
  • Delete data that is no longer needed.

The Four ILM Phases

ILM policies define actions for indices across four main phases:

  • Hot: Indices are actively being written to and frequently queried. Requires fast storage.
  • Warm: Indices are no longer being written to, but are still queried. Can be read-only and potentially on slower storage.
  • Cold: Indices are rarely queried and often stored on very cheap, slow storage.
  • Delete: Indices are no longer needed and are safely removed from the cluster.

Crafting an ILM Policy

Here's an example of an ILM policy that defines rollover, forcemerge (for warm phase), and deletion stages for time-series data:

PUT _ilm/policy/my-ts-policy
{
  "policy": {
    "phases": {
      "hot": {
        "actions": {
          "rollover": {
            "max_age": "7d",
            "max_docs": 10000000,
            "max_size": "50gb"
          }
        }
      },
      "warm": {
        "min_age": "30d",
        "actions": {
          "forcemerge": {
            "max_num_segments": 1
          },
          "shrink": {
            "number_of_shards": 1
          }
        }
      },
      "cold": {
        "min_age": "90d",
        "actions": {
          "freeze": {}
        }
      },
      "delete": {
        "min_age": "180d",
        "actions": {
          "delete": {}
        }
      }
    }
  }
}

Hot-Warm-Cold Architecture

A Hot-Warm-Cold architecture optimizes resource utilization and cost for time-series data by using different types of nodes for different ILM phases.

  • Hot nodes: Use fast CPUs, SSDs for high indexing and query throughput.
  • Warm nodes: Use less powerful CPUs, HDDs (or slower SSDs) for older, read-only data.
  • Cold nodes: Use very cheap, high-capacity storage, potentially even object storage, for rarely accessed data.

Assigning Node Roles

You can configure nodes to serve specific roles (hot, warm, cold) by setting node.attr in their elasticsearch.yml configuration file.

ILM policies then use these attributes to automatically move indices to the appropriate node types as they transition through phases.

# For a hot node
node.roles: [ data ]
node.attr.data: hot

# For a warm node
node.roles: [ data ]
node.attr.data: warm

# For a cold node
node.roles: [ data ]
node.attr.data: cold

ILM & Data Stream Check

Which of the following statements about Elasticsearch's time-series features are TRUE?

Time-Series Recap

Great job! You've explored how Elasticsearch is optimized for time-series data.

We covered:

  • Data Streams: Simplifying index management and rollovers.
  • ILM (Index Lifecycle Management): Automating index transitions through hot, warm, cold, and delete phases.
  • Hot-Warm-Cold Architectures: Optimizing cluster resources and costs by assigning different node types to specific ILM phases.

These features are essential for building scalable and cost-effective time-series solutions with Elasticsearch.

Preguntas frecuentes

¿La lección «Gestión de datos de series temporales» es gratis?

Sí — el texto completo de «Gestión de datos de series temporales» es gratis para leer aquí en la web. Para practicarla de forma interactiva (editor de código integrado y tutor de IA 24/7) y desbloquear el resto del curso de Elasticsearch & Full Text Search Systems, actualiza a CoddyKit PRO. El curso de Elasticsearch & Full Text Search Systems incluye 4 lecciones en total.

¿Qué aprenderé en «Gestión de datos de series temporales»?

Optimice Elasticsearch para datos de series temporales, incluidos los flujos de datos, ILM (Index Lifecycle Management) y las arquitecturas hot-warm-cold. Practicas Elasticsearch & Full Text Search Systems con código real que ejecutas directamente en el navegador, y un tutor de IA 24/7 responde tus preguntas mientras trabajas en la lección.

¿Necesito experiencia previa para empezar Elasticsearch & Full Text Search Systems?

No se requiere experiencia previa. Elasticsearch & Full Text Search Systems en CoddyKit está estructurado para principiantes hasta estudiantes avanzados, así que puedes empezar aquí o desde el inicio y avanzar a tu ritmo. Esta es la lección 2 de 4.

¿Cuánto tiempo toma la lección «Gestión de datos de series temporales»?

La mayoría de las lecciones de CoddyKit toman alrededor de 5–10 minutos. Cada una es compacta e interactiva, así que avanzas constantemente y retomas exactamente por donde dejaste en la web y la app.

¿Puedo escribir y ejecutar código en esta lección de Elasticsearch & Full Text Search Systems?

Sí. Cada lección de Elasticsearch & Full Text Search Systems incluye un editor de código integrado, así que escribes y ejecutas código real directamente en tu navegador y obtienes retroalimentación instantánea de IA — sin configuración local necesaria.

Todas las lecciones de este curso

  1. Capacidades de búsqueda geoespacial
  2. Gestión de datos de series temporales
  3. Estrategias de despliegue en producción
  4. Gestión del ciclo de vida de los índices
← Volver a Elasticsearch & Full Text Search Systems