0Pricing
Apache Kafka & Stream Processing Fundamentals · Lección

Esquemas Avro y Protobuf

Aprenda sobre formatos de serialización populares como Avro y Protobuf y cómo se utilizan con Schema Registry.

Esquemas Avro y Protobuf es una lección gratuita de Apache Kafka & Stream Processing Fundamentals en CoddyKit. Esta es la lección 2 de 4. Puedes leer la lección completa abajo gratuitamente — luego la practicas en el navegador con un editor de código integrado y un tutor de IA 24/7. Forma parte de la ruta de aprendizaje de Apache Kafka & Stream Processing Fundamentals, y tu progreso se sincroniza en la web y la app de CoddyKit. El curso de Apache Kafka & Stream Processing Fundamentals incluye 4 lecciones en total.

Partes de esta lección aún no han sido traducidas y se muestran en inglés.

Data Formats for Kafka

When sending data through Kafka, it's just bytes. To make sense of these bytes, both the sender (producer) and receiver (consumer) need to agree on a common structure.

This is where data serialization formats and schemas come in, providing a blueprint for your data.

Meet Apache Avro

Apache Avro is a popular, language-agnostic data serialization system. It relies heavily on schemas to define the structure of data.

  • Compactness: Data is serialized into a compact binary format.
  • Schema Evolution: Avro handles schema changes gracefully, allowing producers and consumers with different schema versions to communicate.
  • Language Agnostic: Tools exist for many programming languages.

Avro Schemas: JSON Power

Avro schemas are defined using JSON. This makes them human-readable and easy to manage. A schema describes the data's fields, their types, and any default values.

Basic Avro types include string, int, long, boolean, float, double, bytes, and null.

Building an Avro Schema

Let's define a simple Avro schema for a "User" record. Notice the type, name, and fields with their own name and type.

{
  "type": "record",
  "name": "User",
  "namespace": "com.coddykit.avro",
  "fields": [
    {"name": "name", "type": "string"},
    {"name": "age", "type": ["int", "null"], "default": 0}
  ]
}

How Avro Uses Schemas

With Avro, the schema travels with the data (or is known by the consumer via Schema Registry). This means the actual data payload is very small, as field names and types aren't repeated for every message.

The schema acts as a contract, ensuring data consistency and enabling efficient serialization and deserialization.

Meet Protocol Buffers (Protobuf)

Protocol Buffers (Protobuf) is Google's language-neutral, platform-neutral, extensible mechanism for serializing structured data. It's designed to be smaller and faster than XML.

  • Efficiency: Very compact binary format.
  • Code Generation: Compilers generate code for various languages based on .proto definitions.
  • Backward/Forward Compatibility: Supports schema evolution through careful field numbering.

Protobuf Schemas: .proto Files

Protobuf schemas are defined in .proto files using a special syntax. You define message types, which are like classes, and specify fields within them.

Each field requires a type, a name, and a unique field number. These numbers are crucial for backward and forward compatibility.

Building a Protobuf Schema

Here's a simple .proto definition for a "Product" message. Notice the syntax, message keyword, and the assigned field numbers (e.g., 1, 2).

syntax = "proto3";

package com.coddykit.protobuf;

message Product {
  string id = 1;
  string name = 2;
  double price = 3;
}

Avro vs. Protobuf: A Quick Comparison

Both Avro and Protobuf are excellent for data serialization, but they have different characteristics:

  • Schema Format: Avro uses JSON; Protobuf uses its own .proto syntax.
  • Code Generation: Avro is schema-first (can generate code); Protobuf is often code-first (generates code from .proto).
  • Data Size: Both are very compact, often outperforming JSON/XML.
  • Schema Evolution: Both support it, but with different strategies (Avro relies on Schema Registry, Protobuf on field numbers).

Quick Check

You've learned about Avro and Protobuf. Which statement correctly identifies the primary way Avro and Protobuf schemas are defined?

Recap: Structured Data for Kafka

Great job! In this lesson, we explored two powerful data serialization formats: Apache Avro and Protocol Buffers (Protobuf).

  • Avro uses JSON for schema definitions and is excellent for schema evolution.
  • Protobuf uses a .proto syntax with field numbers for compact, efficient data.

Both help ensure data consistency and efficiency when used with Kafka and Schema Registry, which we'll explore further next!

Preguntas frecuentes

¿La lección «Esquemas Avro y Protobuf» es gratis?

Sí — el texto completo de «Esquemas Avro y Protobuf» es gratis para leer aquí en la web. Para practicarla de forma interactiva (editor de código integrado y tutor de IA 24/7) y desbloquear el resto del curso de Apache Kafka & Stream Processing Fundamentals, actualiza a CoddyKit PRO. El curso de Apache Kafka & Stream Processing Fundamentals incluye 4 lecciones en total.

¿Qué aprenderé en «Esquemas Avro y Protobuf»?

Aprenda sobre formatos de serialización populares como Avro y Protobuf y cómo se utilizan con Schema Registry. Practicas Apache Kafka & Stream Processing Fundamentals con código real que ejecutas directamente en el navegador, y un tutor de IA 24/7 responde tus preguntas mientras trabajas en la lección.

¿Necesito experiencia previa para empezar Apache Kafka & Stream Processing Fundamentals?

No se requiere experiencia previa. Apache Kafka & Stream Processing Fundamentals en CoddyKit está estructurado para principiantes hasta estudiantes avanzados, así que puedes empezar aquí o desde el inicio y avanzar a tu ritmo. Esta es la lección 2 de 4.

¿Cuánto tiempo toma la lección «Esquemas Avro y Protobuf»?

La mayoría de las lecciones de CoddyKit toman alrededor de 5–10 minutos. Cada una es compacta e interactiva, así que avanzas constantemente y retomas exactamente por donde dejaste en la web y la app.

¿Puedo escribir y ejecutar código en esta lección de Apache Kafka & Stream Processing Fundamentals?

Sí. Cada lección de Apache Kafka & Stream Processing Fundamentals incluye un editor de código integrado, así que escribes y ejecutas código real directamente en tu navegador y obtienes retroalimentación instantánea de IA — sin configuración local necesaria.

Todas las lecciones de este curso

  1. ¿Por qué gestionar esquemas?
  2. Esquemas Avro y Protobuf
  3. Integración de Schema Registry con Kafka
  4. Evolución de esquemas y modos de compatibilidad
← Volver a Apache Kafka & Stream Processing Fundamentals