取り込み用Source Connector
データベース、ファイル、その他のソースからKafkaへデータを取り込むSource Connectorの設定とデプロイ方法を学習します。
「取り込み用Source Connector」はCoddyKit上の無料Apache Kafka & Stream Processing Fundamentalsレッスンです。 これはレッスン2/4です。 下記で完全なレッスンを無料で読むことができます。その後、ブラウザ内の組み込みコードエディタと24時間対応のAIチューターでハンズオン演習できます。 これはApache Kafka & Stream Processing Fundamentals学習パスの一部であり、ウェブとCoddyKitアプリ全体で進捗が同期されます。 Apache Kafka & Stream Processing Fundamentalsコースには全4レッスンが含まれています。
このレッスンの一部はまだ翻訳されておらず、英語で表示されています。
Data Ingestion with Connectors
Welcome! In this lesson, we'll dive into Source Connectors, a powerful feature of Kafka Connect.
Source connectors are like bridges. They help you bring data from external systems, such as databases or files, into Kafka topics.
This allows your data to flow seamlessly into your real-time data pipelines.
Why Use Source Connectors?
Imagine you need to move data from a database into Kafka. You could write custom code, but that takes time and effort.
Source Connectors simplify this process by:
- Reducing boilerplate code: No need to write custom producers.
- Providing fault tolerance: They handle failures and resume data transfer.
- Scaling easily: Distribute work across multiple Kafka Connect workers.
- Offering pre-built solutions: Many common connectors are already available.
How Source Connectors Work
At a high level, a source connector operates by:
- Polling Data: It continuously checks the external system for new or updated data.
- Converting Data: It transforms the external data format into Kafka records.
- Producing to Kafka: These records are then sent to a specified Kafka topic.
The Kafka Connect framework manages the connector's lifecycle and distributes its tasks.
Common Source Connector Types
Kafka Connect has a rich ecosystem of connectors. Here are a few popular examples:
- FileStreamSourceConnector: Reads data from local files. Great for initial testing!
- JdbcSourceConnector: Connects to relational databases (like PostgreSQL, MySQL) to pull data.
- S3 Source Connector: Ingests data from Amazon S3 buckets.
- Cloud-specific connectors: For Google Cloud Storage, Azure Blob Storage, etc.
Each connector is designed for a specific data source.
Essential Connector Configuration
When you set up a connector, you provide a configuration. This is typically a JSON or properties file.
Key properties include:
name: A unique name for your connector instance.connector.class: The fully qualified class name of the connector to use (e.g.,FileStreamSourceConnector).tasks.max: The maximum number of tasks the connector can use to parallelize data ingestion.
Other properties are specific to the connector's type.
Example: FileStreamSource Connector
Let's use the FileStreamSourceConnector for a practical example. It's simple and helps demonstrate the core concepts.
This connector reads new lines appended to a specified file and publishes each line as a message to a Kafka topic.
First, let's create a simple input file named test.txt with some initial content.
Configuring Our FileStreamSource
To tell the FileStreamSourceConnector what to do, we create a configuration file. Let's call it file-source-connector.json:
{
"name": "local-file-source",
"config": {
"connector.class": "org.apache.kafka.connect.file.FileStreamSourceConnector",
"tasks.max": "1",
"file": "/path/to/your/test.txt",
"topic": "file-input-topic"
}
}Remember to replace /path/to/your/test.txt with the actual path to your file.
Deploying the Connector via REST
Kafka Connect clusters expose a REST API to manage connectors. We can use curl to deploy our connector.
Assuming your Kafka Connect worker is running on localhost:8083, you'd send a POST request:
curl -X POST -H "Content-Type: application/json" \
--data @file-source-connector.json \
http://localhost:8083/connectorsThis command tells Kafka Connect to create a new connector using the configuration in our JSON file.
Checking Connector Status
After deploying, you'll want to ensure your connector started correctly. You can check its status using another REST API call:
curl http://localhost:8083/connectors/local-file-source/statusLook for the "state": "RUNNING" in the response. If it's FAILED, check the Kafka Connect worker logs for error details.
Once running, any new lines added to test.txt will appear in the file-input-topic in Kafka!
Quick Check: Source Connector Role
What is the primary function of a Kafka Connect Source Connector?
Recap: Ingesting Data with Connectors
Congratulations! You've learned the essentials of Kafka Connect Source Connectors.
- Source Connectors ingest data from external systems into Kafka.
- They offer a code-free, fault-tolerant, and scalable way to integrate data.
- You configure them with properties like
connector.class,name, andtasks.max. - Deployment and monitoring are done via the Kafka Connect REST API.
Next, we'll explore the other side of the coin: Sink Connectors, for exporting data from Kafka!
よくある質問
「取り込み用Source Connector」レッスンは無料ですか?
はい。「取り込み用Source Connector」の完全なテキストはこのウェブで無料で読めます。インタラクティブに演習し(組み込みコードエディタと24時間対応のAIチューター)、Apache Kafka & Stream Processing Fundamentalsコースの残りをアンロックするには、CoddyKit PROにアップグレードしてください。 Apache Kafka & Stream Processing Fundamentalsコースには全4レッスンが含まれています。
「取り込み用Source Connector」で何を学びますか?
データベース、ファイル、その他のソースからKafkaへデータを取り込むSource Connectorの設定とデプロイ方法を学習します。 ブラウザで直接実行するハンズオンコードでApache Kafka & Stream Processing Fundamentalsを演習し、24時間対応のAIチューターがレッスンを進める中での質問に答えます。
Apache Kafka & Stream Processing Fundamentalsを始めるのに経験は必要ですか?
事前経験は必要ありません。CoddyKitのApache Kafka & Stream Processing Fundamentalsは初級者から上級者向けに構成されているため、ここから始めるか最初から始めて、自分のペースで進むことができます。 これはレッスン2/4です。
「取り込み用Source Connector」レッスンにはどのくらい時間がかかりますか?
ほとんどのCoddyKitレッスンは約5~10分かかります。各レッスンはコンパクトでインタラクティブなので、着実に進歩し、ウェブとアプリ全体で正確に前回の場所から再開できます。
このApache Kafka & Stream Processing Fundamentalsレッスンでコードを書いて実行できますか?
はい。すべてのApache Kafka & Stream Processing Fundamentalsレッスンに組み込みコードエディタが含まれているため、ブラウザでリアルコードを書いて実行し、即座のAIフィードバックを取得できます。ローカル設定は不要です。
このコースのすべてのレッスン
- Kafka Connect入門
- 取り込み用Source Connector
- エクスポート用Sink Connector
- 単一メッセージ変換(SMT)