0Pricing
Apache Kafka & Stream Processing Fundamentals · Lesson

Source Connectors for Ingestion

Learn to configure and deploy source connectors to import data from databases, files, and other sources into Kafka.

Source Connectors for Ingestion is a free Apache Kafka & Stream Processing Fundamentals lesson on CoddyKit — lesson 2 of 4. You can read the complete lesson below for free — then practise it hands-on in the browser with a built-in code editor and a 24/7 AI tutor. It is part of the Apache Kafka & Stream Processing Fundamentals learning path, one of 4 lessons in the course, and your progress syncs across the web and the CoddyKit app.

Data Ingestion with Connectors

Welcome! In this lesson, we'll dive into Source Connectors, a powerful feature of Kafka Connect.

Source connectors are like bridges. They help you bring data from external systems, such as databases or files, into Kafka topics.

This allows your data to flow seamlessly into your real-time data pipelines.

Why Use Source Connectors?

Imagine you need to move data from a database into Kafka. You could write custom code, but that takes time and effort.

Source Connectors simplify this process by:

  • Reducing boilerplate code: No need to write custom producers.
  • Providing fault tolerance: They handle failures and resume data transfer.
  • Scaling easily: Distribute work across multiple Kafka Connect workers.
  • Offering pre-built solutions: Many common connectors are already available.

How Source Connectors Work

At a high level, a source connector operates by:

  • Polling Data: It continuously checks the external system for new or updated data.
  • Converting Data: It transforms the external data format into Kafka records.
  • Producing to Kafka: These records are then sent to a specified Kafka topic.

The Kafka Connect framework manages the connector's lifecycle and distributes its tasks.

Common Source Connector Types

Kafka Connect has a rich ecosystem of connectors. Here are a few popular examples:

  • FileStreamSourceConnector: Reads data from local files. Great for initial testing!
  • JdbcSourceConnector: Connects to relational databases (like PostgreSQL, MySQL) to pull data.
  • S3 Source Connector: Ingests data from Amazon S3 buckets.
  • Cloud-specific connectors: For Google Cloud Storage, Azure Blob Storage, etc.

Each connector is designed for a specific data source.

Essential Connector Configuration

When you set up a connector, you provide a configuration. This is typically a JSON or properties file.

Key properties include:

  • name: A unique name for your connector instance.
  • connector.class: The fully qualified class name of the connector to use (e.g., FileStreamSourceConnector).
  • tasks.max: The maximum number of tasks the connector can use to parallelize data ingestion.

Other properties are specific to the connector's type.

Example: FileStreamSource Connector

Let's use the FileStreamSourceConnector for a practical example. It's simple and helps demonstrate the core concepts.

This connector reads new lines appended to a specified file and publishes each line as a message to a Kafka topic.

First, let's create a simple input file named test.txt with some initial content.

Configuring Our FileStreamSource

To tell the FileStreamSourceConnector what to do, we create a configuration file. Let's call it file-source-connector.json:

{
  "name": "local-file-source",
  "config": {
    "connector.class": "org.apache.kafka.connect.file.FileStreamSourceConnector",
    "tasks.max": "1",
    "file": "/path/to/your/test.txt",
    "topic": "file-input-topic"
  }
}

Remember to replace /path/to/your/test.txt with the actual path to your file.

Deploying the Connector via REST

Kafka Connect clusters expose a REST API to manage connectors. We can use curl to deploy our connector.

Assuming your Kafka Connect worker is running on localhost:8083, you'd send a POST request:

curl -X POST -H "Content-Type: application/json" \
      --data @file-source-connector.json \
      http://localhost:8083/connectors

This command tells Kafka Connect to create a new connector using the configuration in our JSON file.

Checking Connector Status

After deploying, you'll want to ensure your connector started correctly. You can check its status using another REST API call:

curl http://localhost:8083/connectors/local-file-source/status

Look for the "state": "RUNNING" in the response. If it's FAILED, check the Kafka Connect worker logs for error details.

Once running, any new lines added to test.txt will appear in the file-input-topic in Kafka!

Quick Check: Source Connector Role

What is the primary function of a Kafka Connect Source Connector?

Recap: Ingesting Data with Connectors

Congratulations! You've learned the essentials of Kafka Connect Source Connectors.

  • Source Connectors ingest data from external systems into Kafka.
  • They offer a code-free, fault-tolerant, and scalable way to integrate data.
  • You configure them with properties like connector.class, name, and tasks.max.
  • Deployment and monitoring are done via the Kafka Connect REST API.

Next, we'll explore the other side of the coin: Sink Connectors, for exporting data from Kafka!

Frequently asked questions

Is the “Source Connectors for Ingestion” lesson free?

Yes — the full text of “Source Connectors for Ingestion” is free to read here on the web, and the Apache Kafka & Stream Processing Fundamentals course includes 4 lessons in total. To practise it interactively (a built-in code editor and a 24/7 AI tutor) and unlock the rest of the Apache Kafka & Stream Processing Fundamentals course, upgrade to CoddyKit PRO.

What will I learn in “Source Connectors for Ingestion”?

Learn to configure and deploy source connectors to import data from databases, files, and other sources into Kafka. You practise Apache Kafka & Stream Processing Fundamentals with hands-on code you run directly in the browser, and a 24/7 AI tutor answers your questions as you work through the lesson.

Do I need any experience to start Apache Kafka & Stream Processing Fundamentals?

No prior experience is required. Apache Kafka & Stream Processing Fundamentals on CoddyKit is structured for beginners through advanced learners; this is — lesson 2 of 4, so you can start here or from the beginning and move at your own pace.

How long does the “Source Connectors for Ingestion” lesson take?

Most CoddyKit lessons take about 5–10 minutes. Each one is bite-sized and interactive, so you make steady progress and pick up exactly where you left off across the web and the app.

Can I write and run code in this Apache Kafka & Stream Processing Fundamentals lesson?

Yes. Every Apache Kafka & Stream Processing Fundamentals lesson includes a built-in code editor, so you write and run real code right in your browser and get instant AI feedback — no local setup required.

All lessons in this course

  1. Introduction to Kafka Connect
  2. Source Connectors for Ingestion
  3. Sink Connectors for Export
  4. Single Message Transforms (SMTs)
← Back to Apache Kafka & Stream Processing Fundamentals