Apache Kafka & Stream Processing Fundamentals · Lezione

Schemi Avro e Protobuf

Impari a conoscere formati di serializzazione diffusi come Avro e Protobuf e il loro utilizzo con Schema Registry.

Lezione 2 di 411 passaggi

Schemi Avro e Protobuf è una lezione Apache Kafka & Stream Processing Fundamentals gratuita su CoddyKit. Questa è la lezione 2 di 4. Puoi leggere la lezione completa qui gratuitamente — poi esercitati direttamente nel browser con un editor di codice integrato e un tutor IA disponibile 24/7. Fa parte del percorso di apprendimento Apache Kafka & Stream Processing Fundamentals, e i tuoi progressi si sincronizzano tra il web e l'app CoddyKit. Il corso Apache Kafka & Stream Processing Fundamentals include 4 lezioni in totale.

Parti di questa lezione non sono ancora state tradotte e vengono mostrate in inglese.

Data Formats for Kafka

When sending data through Kafka, it's just bytes. To make sense of these bytes, both the sender (producer) and receiver (consumer) need to agree on a common structure.

This is where data serialization formats and schemas come in, providing a blueprint for your data.

Meet Apache Avro

Apache Avro is a popular, language-agnostic data serialization system. It relies heavily on schemas to define the structure of data.

  • Compactness: Data is serialized into a compact binary format.
  • Schema Evolution: Avro handles schema changes gracefully, allowing producers and consumers with different schema versions to communicate.
  • Language Agnostic: Tools exist for many programming languages.

Avro Schemas: JSON Power

Avro schemas are defined using JSON. This makes them human-readable and easy to manage. A schema describes the data's fields, their types, and any default values.

Basic Avro types include string, int, long, boolean, float, double, bytes, and null.

Building an Avro Schema

Let's define a simple Avro schema for a "User" record. Notice the type, name, and fields with their own name and type.

{
  "type": "record",
  "name": "User",
  "namespace": "com.coddykit.avro",
  "fields": [
    {"name": "name", "type": "string"},
    {"name": "age", "type": ["int", "null"], "default": 0}
  ]
}

How Avro Uses Schemas

With Avro, the schema travels with the data (or is known by the consumer via Schema Registry). This means the actual data payload is very small, as field names and types aren't repeated for every message.

The schema acts as a contract, ensuring data consistency and enabling efficient serialization and deserialization.

Meet Protocol Buffers (Protobuf)

Protocol Buffers (Protobuf) is Google's language-neutral, platform-neutral, extensible mechanism for serializing structured data. It's designed to be smaller and faster than XML.

  • Efficiency: Very compact binary format.
  • Code Generation: Compilers generate code for various languages based on .proto definitions.
  • Backward/Forward Compatibility: Supports schema evolution through careful field numbering.

Protobuf Schemas: .proto Files

Protobuf schemas are defined in .proto files using a special syntax. You define message types, which are like classes, and specify fields within them.

Each field requires a type, a name, and a unique field number. These numbers are crucial for backward and forward compatibility.

Building a Protobuf Schema

Here's a simple .proto definition for a "Product" message. Notice the syntax, message keyword, and the assigned field numbers (e.g., 1, 2).

syntax = "proto3";

package com.coddykit.protobuf;

message Product {
  string id = 1;
  string name = 2;
  double price = 3;
}

Avro vs. Protobuf: A Quick Comparison

Both Avro and Protobuf are excellent for data serialization, but they have different characteristics:

  • Schema Format: Avro uses JSON; Protobuf uses its own .proto syntax.
  • Code Generation: Avro is schema-first (can generate code); Protobuf is often code-first (generates code from .proto).
  • Data Size: Both are very compact, often outperforming JSON/XML.
  • Schema Evolution: Both support it, but with different strategies (Avro relies on Schema Registry, Protobuf on field numbers).

Quick Check

You've learned about Avro and Protobuf. Which statement correctly identifies the primary way Avro and Protobuf schemas are defined?

Recap: Structured Data for Kafka

Great job! In this lesson, we explored two powerful data serialization formats: Apache Avro and Protocol Buffers (Protobuf).

  • Avro uses JSON for schema definitions and is excellent for schema evolution.
  • Protobuf uses a .proto syntax with field numbers for compact, efficient data.

Both help ensure data consistency and efficiency when used with Kafka and Schema Registry, which we'll explore further next!

Gratis per iniziare

Impara Apache Kafka & Stream Processing Fundamentals con un tutor IA — gratis

Scrivi ed esegui vero codice nel tuo browser, ricevi aiuto istantaneo da un tutor IA disponibile 24/7, e riprendi da dove hai lasciato sul web o nell'app.

Corsi
12
Lezioni
48

Domande Frequenti

La lezione «Schemi Avro e Protobuf» è gratuita?

Sì — il testo completo di «Schemi Avro e Protobuf» è gratuito qui sul web. Per esercitarvi in modo interattivo (un editor di codice integrato e un tutor IA 24/7) e sbloccare il resto del corso Apache Kafka & Stream Processing Fundamentals, passa a CoddyKit PRO. Il corso Apache Kafka & Stream Processing Fundamentals include 4 lezioni in totale.

Cosa imparerò in «Schemi Avro e Protobuf»?

Impari a conoscere formati di serializzazione diffusi come Avro e Protobuf e il loro utilizzo con Schema Registry. Eserciti Apache Kafka & Stream Processing Fundamentals con codice pratico che esegui direttamente nel browser, e un tutor IA 24/7 risponde alle tue domande mentre lavori sulla lezione.

Ho bisogno di esperienza per iniziare Apache Kafka & Stream Processing Fundamentals?

Non è richiesta alcuna esperienza precedente. Apache Kafka & Stream Processing Fundamentals su CoddyKit è strutturato per principianti e studenti avanzati, quindi puoi iniziare da qui o dall'inizio e procedere al tuo ritmo. Questa è la lezione 2 di 4.

Quanto tempo richiede la lezione «Schemi Avro e Protobuf»?

La maggior parte delle lezioni CoddyKit richiede circa 5–10 minuti. Ogni lezione è breve e interattiva, quindi fai progressi costanti e riprendi esattamente da dove hai lasciato su web e app.

Posso scrivere ed eseguire codice in questa lezione Apache Kafka & Stream Processing Fundamentals?

Sì. Ogni lezione Apache Kafka & Stream Processing Fundamentals include un editor di codice integrato, quindi scrivi ed esegui codice reale direttamente nel tuo browser e ricevi feedback istantaneo dall'IA — nessuna configurazione locale necessaria.

Tutte le lezioni di questo corso

  1. Perché gestire gli schemi?
  2. Schemi Avro e Protobuf
  3. Integrazione di Schema Registry con Kafka
  4. Evoluzione dello schema e modalità di compatibilità
← Torna a Apache Kafka & Stream Processing Fundamentals