0Pricing
Apache Kafka & Stream Processing Fundamentals · レッスン

パーティションとオフセットの理解

スケーラビリティと並列処理におけるパーティションの重要性と、オフセットによるコンシューマーの進捗管理を理解します。

「パーティションとオフセットの理解」はCoddyKit上の無料Apache Kafka & Stream Processing Fundamentalsレッスンです。 これはレッスン3/4です。 下記で完全なレッスンを無料で読むことができます。その後、ブラウザ内の組み込みコードエディタと24時間対応のAIチューターでハンズオン演習できます。 これはApache Kafka & Stream Processing Fundamentals学習パスの一部であり、ウェブとCoddyKitアプリ全体で進捗が同期されます。 Apache Kafka & Stream Processing Fundamentalsコースには全4レッスンが含まれています。

このレッスンの一部はまだ翻訳されておらず、英語で表示されています。

What are Kafka Partitions?

Imagine a Kafka topic as a category for messages. To handle lots of messages efficiently, Kafka divides a topic into smaller, ordered segments called partitions.

Think of each partition as its own mini-log. Messages are appended to the end of a partition in the order they arrive. Once written, messages in a partition are immutable.

Partitions: Ordered & Immutable

It's crucial to understand that while messages within a single partition are strictly ordered, there's no guaranteed order across different partitions of the same topic.

  • Ordered: Messages in one partition always have a clear sequence.
  • Immutable: Once a message is written to a partition, it cannot be changed.
  • Append-only: New messages are always added to the end.

Scalability Through Partitions

Partitions are the backbone of Kafka's scalability and parallelism. Here's why they matter:

  • Parallel Processing: Multiple consumers can read from different partitions of the same topic simultaneously.
  • Distributed Storage: Partitions can be spread across different Kafka brokers (servers) in a cluster. This allows topics to handle more data than a single server could.

How Messages Are Assigned

When a producer sends a message, Kafka needs to decide which partition it should go into. This is called partitioning strategy:

  • With a Key: If a message includes a key (e.g., a user ID), Kafka uses a hash of that key to consistently assign it to the same partition. This ensures all messages for a specific key are processed in order.
  • Without a Key: If no key is provided, Kafka typically uses a round-robin approach, distributing messages evenly across all partitions.

Producer with Message Keys

This Java example shows how a producer sends messages to a topic, explicitly providing a key. Messages with the same key will end up in the same partition.

import java.util.Properties;
import org.apache.kafka.clients.producer.KafkaProducer;
import org.apache.kafka.clients.producer.ProducerRecord;

public class KeyedProducer {
  public static void main(String[] args) {
    Properties props = new Properties();
    props.put("bootstrap.servers", "localhost:9092");
    props.put("key.serializer", "org.apache.kafka.common.serialization.StringSerializer");
    props.put("value.serializer", "org.apache.kafka.common.serialization.StringSerializer");

    try (KafkaProducer<String, String> producer = new KafkaProducer<>(props)) {
      String topic = "my_keyed_topic";
      for (int i = 0; i < 4; i++) {
        String key = "user-" + (i % 2); // user-0, user-1, user-0, user-1
        String value = "Message " + i + " for " + key;
        producer.send(new ProducerRecord<>(topic, key, value));
        System.out.println("Sent: Key=" + key + ", Value=" + value);
      }
    } catch (Exception e) {
      e.printStackTrace();
    }
  }
}

Introducing Message Offsets

Every message within a Kafka partition has a unique, sequential identifier called an offset. Think of it as an index number for messages within that specific partition.

  • The first message in a partition has offset 0.
  • The next message has offset 1, and so on.
  • Offsets are local to each partition.

Offsets for Consumer Progress

Offsets are critical for consumers to track their progress. A consumer keeps a record of the offset of the last message it successfully processed in each partition.

This allows consumers to:

  • Resume processing exactly where they left off if they stop or crash.
  • Know which messages they still need to read.

Committing Offsets

After processing messages, consumers need to inform Kafka about their progress by committing their offsets. This means saving the current offset to a special Kafka topic (__consumer_offsets).

Committing can be:

  • Automatic: Kafka commits offsets periodically in the background.
  • Manual: The application explicitly tells Kafka when to commit offsets, offering more control over processing guarantees.

Check Your Understanding

Let's test your knowledge about Kafka partitions and offsets.

Recap: Partitions & Offsets

Today, we explored two core Kafka concepts:

  • Partitions: These segments divide a topic, enabling parallel processing, distributed storage, and ordered messages within each partition.
  • Offsets: These sequential IDs track the position of messages within a partition, allowing consumers to precisely manage their progress and resume reliably.

Understanding these concepts is key to building scalable and robust Kafka applications!

よくある質問

「パーティションとオフセットの理解」レッスンは無料ですか?

はい。「パーティションとオフセットの理解」の完全なテキストはこのウェブで無料で読めます。インタラクティブに演習し(組み込みコードエディタと24時間対応のAIチューター)、Apache Kafka & Stream Processing Fundamentalsコースの残りをアンロックするには、CoddyKit PROにアップグレードしてください。 Apache Kafka & Stream Processing Fundamentalsコースには全4レッスンが含まれています。

「パーティションとオフセットの理解」で何を学びますか?

スケーラビリティと並列処理におけるパーティションの重要性と、オフセットによるコンシューマーの進捗管理を理解します。 ブラウザで直接実行するハンズオンコードでApache Kafka & Stream Processing Fundamentalsを演習し、24時間対応のAIチューターがレッスンを進める中での質問に答えます。

Apache Kafka & Stream Processing Fundamentalsを始めるのに経験は必要ですか?

事前経験は必要ありません。CoddyKitのApache Kafka & Stream Processing Fundamentalsは初級者から上級者向けに構成されているため、ここから始めるか最初から始めて、自分のペースで進むことができます。 これはレッスン3/4です。

「パーティションとオフセットの理解」レッスンにはどのくらい時間がかかりますか?

ほとんどのCoddyKitレッスンは約5~10分かかります。各レッスンはコンパクトでインタラクティブなので、着実に進歩し、ウェブとアプリ全体で正確に前回の場所から再開できます。

このApache Kafka & Stream Processing Fundamentalsレッスンでコードを書いて実行できますか?

はい。すべてのApache Kafka & Stream Processing Fundamentalsレッスンに組み込みコードエディタが含まれているため、ブラウザでリアルコードを書いて実行し、即座のAIフィードバックを取得できます。ローカル設定は不要です。

このコースのすべてのレッスン

  1. Kafkaへのメッセージ送信
  2. Kafkaからのメッセージ取得
  3. パーティションとオフセットの理解
  4. メッセージキーとパーティショニング戦略
← Apache Kafka & Stream Processing Fundamentalsに戻る