0Pricing
Apache Kafka & Stream Processing Fundamentals · Урок

Схемы Avro и Protobuf

Изучите популярные форматы сериализации, такие как Avro и Protobuf, и способы их использования с Schema Registry.

«Схемы Avro и Protobuf» — бесплатный урок Apache Kafka & Stream Processing Fundamentals на CoddyKit. Это урок 2 из 4. Ты можешь прочитать весь урок бесплатно ниже — а потом практиковать его прямо в браузере с встроенным редактором кода и ИИ-репетитором 24/7. Это часть пути обучения Apache Kafka & Stream Processing Fundamentals, и твой прогресс синхронизируется между веб-версией и приложением CoddyKit. Курс Apache Kafka & Stream Processing Fundamentals содержит 4 уроков всего.

Части этого урока еще не переведены и отображаются на английском.

Data Formats for Kafka

When sending data through Kafka, it's just bytes. To make sense of these bytes, both the sender (producer) and receiver (consumer) need to agree on a common structure.

This is where data serialization formats and schemas come in, providing a blueprint for your data.

Meet Apache Avro

Apache Avro is a popular, language-agnostic data serialization system. It relies heavily on schemas to define the structure of data.

  • Compactness: Data is serialized into a compact binary format.
  • Schema Evolution: Avro handles schema changes gracefully, allowing producers and consumers with different schema versions to communicate.
  • Language Agnostic: Tools exist for many programming languages.

Avro Schemas: JSON Power

Avro schemas are defined using JSON. This makes them human-readable and easy to manage. A schema describes the data's fields, their types, and any default values.

Basic Avro types include string, int, long, boolean, float, double, bytes, and null.

Building an Avro Schema

Let's define a simple Avro schema for a "User" record. Notice the type, name, and fields with their own name and type.

{
  "type": "record",
  "name": "User",
  "namespace": "com.coddykit.avro",
  "fields": [
    {"name": "name", "type": "string"},
    {"name": "age", "type": ["int", "null"], "default": 0}
  ]
}

How Avro Uses Schemas

With Avro, the schema travels with the data (or is known by the consumer via Schema Registry). This means the actual data payload is very small, as field names and types aren't repeated for every message.

The schema acts as a contract, ensuring data consistency and enabling efficient serialization and deserialization.

Meet Protocol Buffers (Protobuf)

Protocol Buffers (Protobuf) is Google's language-neutral, platform-neutral, extensible mechanism for serializing structured data. It's designed to be smaller and faster than XML.

  • Efficiency: Very compact binary format.
  • Code Generation: Compilers generate code for various languages based on .proto definitions.
  • Backward/Forward Compatibility: Supports schema evolution through careful field numbering.

Protobuf Schemas: .proto Files

Protobuf schemas are defined in .proto files using a special syntax. You define message types, which are like classes, and specify fields within them.

Each field requires a type, a name, and a unique field number. These numbers are crucial for backward and forward compatibility.

Building a Protobuf Schema

Here's a simple .proto definition for a "Product" message. Notice the syntax, message keyword, and the assigned field numbers (e.g., 1, 2).

syntax = "proto3";

package com.coddykit.protobuf;

message Product {
  string id = 1;
  string name = 2;
  double price = 3;
}

Avro vs. Protobuf: A Quick Comparison

Both Avro and Protobuf are excellent for data serialization, but they have different characteristics:

  • Schema Format: Avro uses JSON; Protobuf uses its own .proto syntax.
  • Code Generation: Avro is schema-first (can generate code); Protobuf is often code-first (generates code from .proto).
  • Data Size: Both are very compact, often outperforming JSON/XML.
  • Schema Evolution: Both support it, but with different strategies (Avro relies on Schema Registry, Protobuf on field numbers).

Quick Check

You've learned about Avro and Protobuf. Which statement correctly identifies the primary way Avro and Protobuf schemas are defined?

Recap: Structured Data for Kafka

Great job! In this lesson, we explored two powerful data serialization formats: Apache Avro and Protocol Buffers (Protobuf).

  • Avro uses JSON for schema definitions and is excellent for schema evolution.
  • Protobuf uses a .proto syntax with field numbers for compact, efficient data.

Both help ensure data consistency and efficiency when used with Kafka and Schema Registry, which we'll explore further next!

Часто задаваемые вопросы

Урок «Схемы Avro и Protobuf» бесплатный?

Да — полный текст урока «Схемы Avro и Protobuf» бесплатно доступен здесь в веб-версии. Чтобы практиковать его интерактивно (встроенный редактор кода и ИИ-репетитор 24/7) и разблокировать остальной курс Apache Kafka & Stream Processing Fundamentals, подпишись на CoddyKit PRO. Курс Apache Kafka & Stream Processing Fundamentals содержит 4 уроков всего.

Чему я научусь в уроке «Схемы Avro и Protobuf»?

Изучите популярные форматы сериализации, такие как Avro и Protobuf, и способы их использования с Schema Registry. Ты практикуешь Apache Kafka & Stream Processing Fundamentals с помощью реального кода, который запускаешь прямо в браузере, и ИИ-репетитор 24/7 отвечает на твои вопросы во время урока.

Нужен ли мне опыт, чтобы начать Apache Kafka & Stream Processing Fundamentals?

Предыдущий опыт не требуется. Apache Kafka & Stream Processing Fundamentals на CoddyKit структурирован для всех уровней — от новичков до продвинутых, поэтому ты можешь начать отсюда или с самого начала и учиться в своем темпе. Это урок 2 из 4.

Сколько времени занимает урок «Схемы Avro и Protobuf»?

Большинство уроков CoddyKit занимают около 5–10 минут. Каждый из них компактный и интерактивный, поэтому ты постоянно делаешь прогресс и продолжаешь с того же места в веб-версии и приложении.

Можно ли писать и запускать код в этом уроке Apache Kafka & Stream Processing Fundamentals?

Да. Каждый урок Apache Kafka & Stream Processing Fundamentals включает встроенный редактор кода, поэтому ты пишешь и запускаешь реальный код прямо в браузере и получаешь моментальную обратную связь от AI — локальная установка не требуется.

Все уроки этого курса

  1. Зачем нужно управление схемами
  2. Схемы Avro и Protobuf
  3. Интеграция Schema Registry с Kafka
  4. Эволюция схемы и режимы совместимости
← Назад к Apache Kafka & Stream Processing Fundamentals