0Pricing
Apache Kafka & Stream Processing Fundamentals · 课时

提交日志:Kafka 如何存储数据

学习 Kafka 如何将消息存储为仅追加的提交日志,了解偏移量和数据保留的工作方式,以及这种设计为何让 Kafka 既快速又持久。

提交日志:Kafka 如何存储数据 是 CoddyKit 上的免费 Apache Kafka & Stream Processing Fundamentals 课时。 这是第 4 节课,共 4 节。 你可以在下方免费阅读本课时的完整内容 — 然后在浏览器中使用内置代码编辑器和全天候 AI 导师进行实践。 这是 Apache Kafka & Stream Processing Fundamentals 学习路径的一部分,你的进度在网页和 CoddyKit 应用中同步。 Apache Kafka & Stream Processing Fundamentals 课程共包含 4 节课。

本课时的部分内容尚未翻译,以英文显示。

Kafka Is a Log

At its core Kafka is a distributed, append-only commit log. Producers append to the end and data is never modified in place — that's the secret to its speed and durability.

Append-Only Writes

Every record is written sequentially to the end of a partition's log. Sequential disk writes are far faster than random ones, sustaining high throughput on cheap disks.

The Offset

Each record gets a monotonically increasing offset — its position in the partition. Offsets are unique within a partition and never reused.

Partition 0: [off 0][off 1][off 2][off 3] -> next write off 4

Reads Don't Delete

Unlike a queue, consuming a record does not delete it. Many consumers read the same log independently, each tracking its own offset.

Log Segments

A partition log is split into segment files on disk. Only the active segment is written to; older segments are immutable and can be deleted or compacted.

00000000000000000000.log
00000000000000010000.log

Retention by Time

Kafka keeps data for a configurable period regardless of whether it's consumed. The default retention is 7 days.

retention.ms=604800000

Retention by Size

You can also cap a partition by size. Whichever limit hits first — time or bytes — triggers deletion of the oldest segments.

retention.bytes=1073741824

Log Compaction

Instead of deleting by age, compaction keeps only the latest record per key — perfect for changelog topics where you just want each key's current value.

cleanup.policy=compact

Zero-Copy Reads

Kafka serves reads straight from the OS page cache to the network using zero-copy, skipping extra memory copies. Another big reason consumption is so fast.

Durability via Flush and Replicas

Records are durable because they're written to disk and replicated to other brokers. Even if the OS cache is lost, replicas preserve the data.

Putting It Together

The commit log explains Kafka's character: sequential writes for speed, offsets for replayable reads, retention and compaction for storage, replication for durability.

Quick Check

Test your understanding of the commit log.

Recap

You learned Kafka's storage: an append-only commit log with stable offsets, reads that don't delete, retention plus compaction, and replication for durability.

常见问题解答

「提交日志:Kafka 如何存储数据」课时是免费的吗?

是的 — 「提交日志:Kafka 如何存储数据」的完整文本可在网页上免费阅读。要进行交互式练习(内置代码编辑器和全天候 AI 导师)并解锁 Apache Kafka & Stream Processing Fundamentals 课程的其余内容,请升级到 CoddyKit PRO。 Apache Kafka & Stream Processing Fundamentals 课程共包含 4 节课。

「提交日志:Kafka 如何存储数据」这节课中我会学到什么?

学习 Kafka 如何将消息存储为仅追加的提交日志,了解偏移量和数据保留的工作方式,以及这种设计为何让 Kafka 既快速又持久。 你通过在浏览器中直接运行的动手代码来练习 Apache Kafka & Stream Processing Fundamentals,全天候 AI 导师会在你学习这节课的过程中回答你的问题。

学习 Apache Kafka & Stream Processing Fundamentals 需要有经验吗?

无需任何先前经验。CoddyKit 上的 Apache Kafka & Stream Processing Fundamentals 课程适合初学者到高级学习者,你可以从这里开始或从头开始,按照自己的节奏学习。 这是第 4 节课,共 4 节。

「提交日志:Kafka 如何存储数据」课时需要多长时间?

大多数 CoddyKit 课程大约需要 5–10 分钟。每节课都很精短且互动,所以你能稳步进步,并在网页和应用中从离开的地方继续。

我能在这节 Apache Kafka & Stream Processing Fundamentals 课中编写并运行代码吗?

能。每节 Apache Kafka & Stream Processing Fundamentals 课都包含内置代码编辑器,你可以在浏览器中直接编写并运行真实代码,并获得即时 AI 反馈 — 无需本地设置。

此课程中的所有课时

  1. 什么是 Apache Kafka
  2. Kafka 核心概念:代理与主题
  3. 在本地搭建 Kafka
  4. 提交日志:Kafka 如何存储数据
← 返回 Apache Kafka & Stream Processing Fundamentals