0Pricing
Apache Kafka & Stream Processing Fundamentals · 课时

Kafka Connect 简介

了解 Kafka Connect 的架构及其在 Kafka 与外部系统集成方面的优势

Kafka Connect 简介 是 CoddyKit 上的免费 Apache Kafka & Stream Processing Fundamentals 课时。 这是第 1 节课,共 4 节。 你可以在下方免费阅读本课时的完整内容 — 然后在浏览器中使用内置代码编辑器和全天候 AI 导师进行实践。 这是 Apache Kafka & Stream Processing Fundamentals 学习路径的一部分,你的进度在网页和 CoddyKit 应用中同步。 Apache Kafka & Stream Processing Fundamentals 课程共包含 4 节课。

本课时的部分内容尚未翻译,以英文显示。

Meet Kafka Connect

Kafka Connect is a powerful framework for streaming data between Apache Kafka and other data systems. Think of it as a bridge that automatically moves data for you.

It simplifies the process of getting data into Kafka (from databases, file systems, etc.) and out of Kafka (to data warehouses, search indexes, etc.).

Why Data Integration is Hard

Moving data between different systems can be tricky. You often need to write custom code for each integration, handle errors, ensure data consistency, and scale it as your data grows.

This takes a lot of time and effort! Kafka Connect aims to solve these headaches by providing a standardized, robust way to do it.

Connectors: The Data Bridges

At the heart of Kafka Connect are Connectors. A connector is a ready-to-use component that knows how to interact with a specific external data system.

You don't write custom code for each integration; you just configure a connector. There are two main types:

  • Source Connectors
  • Sink Connectors

Source Connectors: Into Kafka

Source connectors are responsible for importing data from an external system into Kafka topics.

Imagine you have a database. A database source connector would continuously read new changes or records from that database and publish them as messages to a Kafka topic.

Examples: JDBC Source Connector (databases), FileStreamSource Connector (files).

Sink Connectors: Out of Kafka

Sink connectors do the opposite: they export data from Kafka topics to an external system.

For instance, a data warehouse sink connector would read messages from a Kafka topic and write them into tables in your data warehouse for analysis.

Examples: JDBC Sink Connector (databases), S3 Sink Connector (cloud storage), Elasticsearch Sink Connector (search).

Why Use Kafka Connect?

Kafka Connect offers several powerful benefits:

  • No Code Required: Most integrations are configuration-driven.
  • Scalable: Easily scales to handle large data volumes.
  • Fault-Tolerant: Automatically recovers from failures.
  • Distributed: Can run across multiple servers for high availability.
  • Extensible: Many pre-built connectors, or you can write your own.

How Connect Works

Kafka Connect runs as a cluster of workers. Each worker is a JVM process. These workers host connector tasks.

When you start a connector, Connect distributes its tasks across the available workers. If a worker fails, its tasks are automatically reassigned to other active workers.

Deployment Modes

Kafka Connect can operate in two modes:

  • Standalone Mode: A single process for development or small-scale use. Not fault-tolerant.
  • Distributed Mode: Multiple worker processes form a cluster, providing scalability and fault tolerance. Ideal for production environments.

Most production deployments use the distributed mode for reliability.

Simulating Data Inflow

While Kafka Connect handles the heavy lifting, understanding the basic data flow helps. Here's a simple Java program that sends a message to a Kafka topic, similar to what a source connector might automate.

This example shows how data enters a Kafka topic, which Kafka Connect can then manage.

import org.apache.kafka.clients.producer.KafkaProducer;
import org.apache.kafka.clients.producer.ProducerRecord;
import java.util.Properties;

public class SimpleProducer {
  public static void main(String[] args) {
    // 1. Configure producer
    Properties props = new Properties();
    props.put("bootstrap.servers", "localhost:9092");
    props.put("key.serializer", "org.apache.kafka.common.serialization.StringSerializer");
    props.put("value.serializer", "org.apache.kafka.common.serialization.StringSerializer");

    // 2. Create producer
    try (KafkaProducer<String, String> producer = new KafkaProducer<>(props)) {
      // 3. Create a record
      ProducerRecord<String, String> record = new ProducerRecord<>("my_topic", "hello_key", "Hello from CoddyKit!");

      // 4. Send the record
      producer.send(record);
      System.out.println("Message sent: Hello from CoddyKit!");
    } catch (Exception e) {
      e.printStackTrace();
    }
  }
}

Identify the Connector

You want to move customer order data from a PostgreSQL database into a Kafka topic for real-time processing.

Which type of Kafka Connect connector would you use for this task?

Recap & Next Steps

Great job! In this lesson, we introduced Kafka Connect, a powerful framework for data integration with Kafka.

  • Kafka Connect simplifies moving data between Kafka and other systems.
  • Source Connectors bring data into Kafka.
  • Sink Connectors take data out of Kafka.
  • It offers scalability, fault tolerance, and reduces custom coding.

Next, we'll dive deeper into configuring and deploying Source Connectors!

常见问题解答

「Kafka Connect 简介」课时是免费的吗?

是的 — 「Kafka Connect 简介」的完整文本可在网页上免费阅读。要进行交互式练习(内置代码编辑器和全天候 AI 导师)并解锁 Apache Kafka & Stream Processing Fundamentals 课程的其余内容,请升级到 CoddyKit PRO。 Apache Kafka & Stream Processing Fundamentals 课程共包含 4 节课。

「Kafka Connect 简介」这节课中我会学到什么?

了解 Kafka Connect 的架构及其在 Kafka 与外部系统集成方面的优势 你通过在浏览器中直接运行的动手代码来练习 Apache Kafka & Stream Processing Fundamentals,全天候 AI 导师会在你学习这节课的过程中回答你的问题。

学习 Apache Kafka & Stream Processing Fundamentals 需要有经验吗?

无需任何先前经验。CoddyKit 上的 Apache Kafka & Stream Processing Fundamentals 课程适合初学者到高级学习者,你可以从这里开始或从头开始,按照自己的节奏学习。 这是第 1 节课,共 4 节。

「Kafka Connect 简介」课时需要多长时间?

大多数 CoddyKit 课程大约需要 5–10 分钟。每节课都很精短且互动,所以你能稳步进步,并在网页和应用中从离开的地方继续。

我能在这节 Apache Kafka & Stream Processing Fundamentals 课中编写并运行代码吗?

能。每节 Apache Kafka & Stream Processing Fundamentals 课都包含内置代码编辑器,你可以在浏览器中直接编写并运行真实代码,并获得即时 AI 反馈 — 无需本地设置。

此课程中的所有课时

  1. Kafka Connect 简介
  2. 用于数据导入的源连接器
  3. 用于数据导出的接收器连接器
  4. 单消息转换(SMT)
← 返回 Apache Kafka & Stream Processing Fundamentals