0Pricing
Apache Kafka & Stream Processing Fundamentals · Lección

Conectores de origen para la ingesta

Aprenda a configurar y desplegar conectores de origen para importar datos de bases de datos, archivos y otras fuentes a Kafka.

Conectores de origen para la ingesta es una lección gratuita de Apache Kafka & Stream Processing Fundamentals en CoddyKit. Esta es la lección 2 de 4. Puedes leer la lección completa abajo gratuitamente — luego la practicas en el navegador con un editor de código integrado y un tutor de IA 24/7. Forma parte de la ruta de aprendizaje de Apache Kafka & Stream Processing Fundamentals, y tu progreso se sincroniza en la web y la app de CoddyKit. El curso de Apache Kafka & Stream Processing Fundamentals incluye 4 lecciones en total.

Partes de esta lección aún no han sido traducidas y se muestran en inglés.

Data Ingestion with Connectors

Welcome! In this lesson, we'll dive into Source Connectors, a powerful feature of Kafka Connect.

Source connectors are like bridges. They help you bring data from external systems, such as databases or files, into Kafka topics.

This allows your data to flow seamlessly into your real-time data pipelines.

Why Use Source Connectors?

Imagine you need to move data from a database into Kafka. You could write custom code, but that takes time and effort.

Source Connectors simplify this process by:

  • Reducing boilerplate code: No need to write custom producers.
  • Providing fault tolerance: They handle failures and resume data transfer.
  • Scaling easily: Distribute work across multiple Kafka Connect workers.
  • Offering pre-built solutions: Many common connectors are already available.

How Source Connectors Work

At a high level, a source connector operates by:

  • Polling Data: It continuously checks the external system for new or updated data.
  • Converting Data: It transforms the external data format into Kafka records.
  • Producing to Kafka: These records are then sent to a specified Kafka topic.

The Kafka Connect framework manages the connector's lifecycle and distributes its tasks.

Common Source Connector Types

Kafka Connect has a rich ecosystem of connectors. Here are a few popular examples:

  • FileStreamSourceConnector: Reads data from local files. Great for initial testing!
  • JdbcSourceConnector: Connects to relational databases (like PostgreSQL, MySQL) to pull data.
  • S3 Source Connector: Ingests data from Amazon S3 buckets.
  • Cloud-specific connectors: For Google Cloud Storage, Azure Blob Storage, etc.

Each connector is designed for a specific data source.

Essential Connector Configuration

When you set up a connector, you provide a configuration. This is typically a JSON or properties file.

Key properties include:

  • name: A unique name for your connector instance.
  • connector.class: The fully qualified class name of the connector to use (e.g., FileStreamSourceConnector).
  • tasks.max: The maximum number of tasks the connector can use to parallelize data ingestion.

Other properties are specific to the connector's type.

Example: FileStreamSource Connector

Let's use the FileStreamSourceConnector for a practical example. It's simple and helps demonstrate the core concepts.

This connector reads new lines appended to a specified file and publishes each line as a message to a Kafka topic.

First, let's create a simple input file named test.txt with some initial content.

Configuring Our FileStreamSource

To tell the FileStreamSourceConnector what to do, we create a configuration file. Let's call it file-source-connector.json:

{
  "name": "local-file-source",
  "config": {
    "connector.class": "org.apache.kafka.connect.file.FileStreamSourceConnector",
    "tasks.max": "1",
    "file": "/path/to/your/test.txt",
    "topic": "file-input-topic"
  }
}

Remember to replace /path/to/your/test.txt with the actual path to your file.

Deploying the Connector via REST

Kafka Connect clusters expose a REST API to manage connectors. We can use curl to deploy our connector.

Assuming your Kafka Connect worker is running on localhost:8083, you'd send a POST request:

curl -X POST -H "Content-Type: application/json" \
      --data @file-source-connector.json \
      http://localhost:8083/connectors

This command tells Kafka Connect to create a new connector using the configuration in our JSON file.

Checking Connector Status

After deploying, you'll want to ensure your connector started correctly. You can check its status using another REST API call:

curl http://localhost:8083/connectors/local-file-source/status

Look for the "state": "RUNNING" in the response. If it's FAILED, check the Kafka Connect worker logs for error details.

Once running, any new lines added to test.txt will appear in the file-input-topic in Kafka!

Quick Check: Source Connector Role

What is the primary function of a Kafka Connect Source Connector?

Recap: Ingesting Data with Connectors

Congratulations! You've learned the essentials of Kafka Connect Source Connectors.

  • Source Connectors ingest data from external systems into Kafka.
  • They offer a code-free, fault-tolerant, and scalable way to integrate data.
  • You configure them with properties like connector.class, name, and tasks.max.
  • Deployment and monitoring are done via the Kafka Connect REST API.

Next, we'll explore the other side of the coin: Sink Connectors, for exporting data from Kafka!

Preguntas frecuentes

¿La lección «Conectores de origen para la ingesta» es gratis?

Sí — el texto completo de «Conectores de origen para la ingesta» es gratis para leer aquí en la web. Para practicarla de forma interactiva (editor de código integrado y tutor de IA 24/7) y desbloquear el resto del curso de Apache Kafka & Stream Processing Fundamentals, actualiza a CoddyKit PRO. El curso de Apache Kafka & Stream Processing Fundamentals incluye 4 lecciones en total.

¿Qué aprenderé en «Conectores de origen para la ingesta»?

Aprenda a configurar y desplegar conectores de origen para importar datos de bases de datos, archivos y otras fuentes a Kafka. Practicas Apache Kafka & Stream Processing Fundamentals con código real que ejecutas directamente en el navegador, y un tutor de IA 24/7 responde tus preguntas mientras trabajas en la lección.

¿Necesito experiencia previa para empezar Apache Kafka & Stream Processing Fundamentals?

No se requiere experiencia previa. Apache Kafka & Stream Processing Fundamentals en CoddyKit está estructurado para principiantes hasta estudiantes avanzados, así que puedes empezar aquí o desde el inicio y avanzar a tu ritmo. Esta es la lección 2 de 4.

¿Cuánto tiempo toma la lección «Conectores de origen para la ingesta»?

La mayoría de las lecciones de CoddyKit toman alrededor de 5–10 minutos. Cada una es compacta e interactiva, así que avanzas constantemente y retomas exactamente por donde dejaste en la web y la app.

¿Puedo escribir y ejecutar código en esta lección de Apache Kafka & Stream Processing Fundamentals?

Sí. Cada lección de Apache Kafka & Stream Processing Fundamentals incluye un editor de código integrado, así que escribes y ejecutas código real directamente en tu navegador y obtienes retroalimentación instantánea de IA — sin configuración local necesaria.

Todas las lecciones de este curso

  1. Introducción a Kafka Connect
  2. Conectores de origen para la ingesta
  3. Conectores de destino para la exportación
  4. Transformaciones de un solo mensaje (SMT)
← Volver a Apache Kafka & Stream Processing Fundamentals