0Pricing
Apache Kafka & Stream Processing Fundamentals · レッスン

AvroとProtobufのスキーマ

AvroやProtobufなどの代表的なシリアライズ形式と、Schema Registryとの連携方法を学習します。

「AvroとProtobufのスキーマ」はCoddyKit上の無料Apache Kafka & Stream Processing Fundamentalsレッスンです。 これはレッスン2/4です。 下記で完全なレッスンを無料で読むことができます。その後、ブラウザ内の組み込みコードエディタと24時間対応のAIチューターでハンズオン演習できます。 これはApache Kafka & Stream Processing Fundamentals学習パスの一部であり、ウェブとCoddyKitアプリ全体で進捗が同期されます。 Apache Kafka & Stream Processing Fundamentalsコースには全4レッスンが含まれています。

このレッスンの一部はまだ翻訳されておらず、英語で表示されています。

Data Formats for Kafka

When sending data through Kafka, it's just bytes. To make sense of these bytes, both the sender (producer) and receiver (consumer) need to agree on a common structure.

This is where data serialization formats and schemas come in, providing a blueprint for your data.

Meet Apache Avro

Apache Avro is a popular, language-agnostic data serialization system. It relies heavily on schemas to define the structure of data.

  • Compactness: Data is serialized into a compact binary format.
  • Schema Evolution: Avro handles schema changes gracefully, allowing producers and consumers with different schema versions to communicate.
  • Language Agnostic: Tools exist for many programming languages.

Avro Schemas: JSON Power

Avro schemas are defined using JSON. This makes them human-readable and easy to manage. A schema describes the data's fields, their types, and any default values.

Basic Avro types include string, int, long, boolean, float, double, bytes, and null.

Building an Avro Schema

Let's define a simple Avro schema for a "User" record. Notice the type, name, and fields with their own name and type.

{
  "type": "record",
  "name": "User",
  "namespace": "com.coddykit.avro",
  "fields": [
    {"name": "name", "type": "string"},
    {"name": "age", "type": ["int", "null"], "default": 0}
  ]
}

How Avro Uses Schemas

With Avro, the schema travels with the data (or is known by the consumer via Schema Registry). This means the actual data payload is very small, as field names and types aren't repeated for every message.

The schema acts as a contract, ensuring data consistency and enabling efficient serialization and deserialization.

Meet Protocol Buffers (Protobuf)

Protocol Buffers (Protobuf) is Google's language-neutral, platform-neutral, extensible mechanism for serializing structured data. It's designed to be smaller and faster than XML.

  • Efficiency: Very compact binary format.
  • Code Generation: Compilers generate code for various languages based on .proto definitions.
  • Backward/Forward Compatibility: Supports schema evolution through careful field numbering.

Protobuf Schemas: .proto Files

Protobuf schemas are defined in .proto files using a special syntax. You define message types, which are like classes, and specify fields within them.

Each field requires a type, a name, and a unique field number. These numbers are crucial for backward and forward compatibility.

Building a Protobuf Schema

Here's a simple .proto definition for a "Product" message. Notice the syntax, message keyword, and the assigned field numbers (e.g., 1, 2).

syntax = "proto3";

package com.coddykit.protobuf;

message Product {
  string id = 1;
  string name = 2;
  double price = 3;
}

Avro vs. Protobuf: A Quick Comparison

Both Avro and Protobuf are excellent for data serialization, but they have different characteristics:

  • Schema Format: Avro uses JSON; Protobuf uses its own .proto syntax.
  • Code Generation: Avro is schema-first (can generate code); Protobuf is often code-first (generates code from .proto).
  • Data Size: Both are very compact, often outperforming JSON/XML.
  • Schema Evolution: Both support it, but with different strategies (Avro relies on Schema Registry, Protobuf on field numbers).

Quick Check

You've learned about Avro and Protobuf. Which statement correctly identifies the primary way Avro and Protobuf schemas are defined?

Recap: Structured Data for Kafka

Great job! In this lesson, we explored two powerful data serialization formats: Apache Avro and Protocol Buffers (Protobuf).

  • Avro uses JSON for schema definitions and is excellent for schema evolution.
  • Protobuf uses a .proto syntax with field numbers for compact, efficient data.

Both help ensure data consistency and efficiency when used with Kafka and Schema Registry, which we'll explore further next!

よくある質問

「AvroとProtobufのスキーマ」レッスンは無料ですか?

はい。「AvroとProtobufのスキーマ」の完全なテキストはこのウェブで無料で読めます。インタラクティブに演習し(組み込みコードエディタと24時間対応のAIチューター)、Apache Kafka & Stream Processing Fundamentalsコースの残りをアンロックするには、CoddyKit PROにアップグレードしてください。 Apache Kafka & Stream Processing Fundamentalsコースには全4レッスンが含まれています。

「AvroとProtobufのスキーマ」で何を学びますか?

AvroやProtobufなどの代表的なシリアライズ形式と、Schema Registryとの連携方法を学習します。 ブラウザで直接実行するハンズオンコードでApache Kafka & Stream Processing Fundamentalsを演習し、24時間対応のAIチューターがレッスンを進める中での質問に答えます。

Apache Kafka & Stream Processing Fundamentalsを始めるのに経験は必要ですか?

事前経験は必要ありません。CoddyKitのApache Kafka & Stream Processing Fundamentalsは初級者から上級者向けに構成されているため、ここから始めるか最初から始めて、自分のペースで進むことができます。 これはレッスン2/4です。

「AvroとProtobufのスキーマ」レッスンにはどのくらい時間がかかりますか?

ほとんどのCoddyKitレッスンは約5~10分かかります。各レッスンはコンパクトでインタラクティブなので、着実に進歩し、ウェブとアプリ全体で正確に前回の場所から再開できます。

このApache Kafka & Stream Processing Fundamentalsレッスンでコードを書いて実行できますか?

はい。すべてのApache Kafka & Stream Processing Fundamentalsレッスンに組み込みコードエディタが含まれているため、ブラウザでリアルコードを書いて実行し、即座のAIフィードバックを取得できます。ローカル設定は不要です。

このコースのすべてのレッスン

  1. なぜスキーマ管理が必要なのか
  2. AvroとProtobufのスキーマ
  3. Schema RegistryとKafkaの統合
  4. スキーマ進化と互換性モード
← Apache Kafka & Stream Processing Fundamentalsに戻る