0Pricing
Spring Boot 4 Complete Guide · 课时

面向块的读取器—处理器—写入器流程

连接 ItemReader、ItemProcessor 和 ItemWriter 组件,实现可扩展的分块处理。

面向块的读取器—处理器—写入器流程 是 CoddyKit 上的免费 Spring Boot 4 Complete Guide 课时。 这是第 2 节课,共 4 节。 你可以在下方免费阅读本课时的完整内容 — 然后在浏览器中使用内置代码编辑器和全天候 AI 导师进行实践。 这是 Spring Boot 4 Complete Guide 学习路径的一部分,你的进度在网页和 CoddyKit 应用中同步。 Spring Boot 4 Complete Guide 课程共包含 4 节课。

本课时的部分内容尚未翻译,以英文显示。

Why Chunk-Oriented Processing?

Spring Batch reads, processes, and writes data in chunks instead of one record at a time. A chunk is a configurable number of items (the commit-interval) handled inside a single transaction.

  • Read N items one by one with an ItemReader.
  • Process each item with an ItemProcessor.
  • Write the whole batch of N at once with an ItemWriter.

Writing in bulk and committing per chunk is what makes batch jobs scale to millions of rows without exhausting memory.

The Three Core Interfaces

Every chunk step is built from three single-method interfaces. Knowing their contracts is the foundation of the whole pattern.

  • ItemReader<I> → I read() returns the next item, or null when the input is exhausted.
  • ItemProcessor<I, O> → O process(I item) transforms an input into an output; returning null filters the item out.
  • ItemWriter<O> → void write(Chunk<? extends O> chunk) persists the accumulated chunk.
public interface ItemReader<I> {
    I read() throws Exception; // null = end of input
}

public interface ItemProcessor<I, O> {
    O process(I item) throws Exception; // null = filter
}

public interface ItemWriter<O> {
    void write(Chunk<? extends O> chunk) throws Exception;
}

Defining a Domain Type

A chunk step flows typed data from reader to processor to writer. Let's model a simple input and output. The reader emits raw Customer records and the writer stores normalized ones.

Using a Java record keeps these immutable value types concise. This snippet is plain Java with no framework, so it runs standalone.

public class DomainDemo {
    record Customer(String name, String email) {}

    public static void main(String[] args) {
        Customer c = new Customer("  Ada ", "ADA@MAIL.COM");
        Customer normalized = new Customer(
                c.name().trim(),
                c.email().toLowerCase());
        System.out.println(normalized);
    }
}

Building the ItemReader

For database input, JdbcCursorItemReader streams rows one at a time so memory stays flat. You give it a DataSource, a SQL query, and a RowMapper to turn each row into a domain object.

Spring Batch calls read() repeatedly until it returns null, advancing the cursor each time.

@Bean
public JdbcCursorItemReader<Customer> reader(DataSource dataSource) {
    return new JdbcCursorItemReaderBuilder<Customer>()
            .name("customerReader")
            .dataSource(dataSource)
            .sql("SELECT name, email FROM customers WHERE active = true")
            .rowMapper((rs, rowNum) ->
                    new Customer(rs.getString("name"), rs.getString("email")))
            .build();
}

Building the ItemProcessor

The processor is where business logic lives: validation, enrichment, transformation, or filtering. Its input and output types may differ.

  • Return a transformed object to pass it downstream.
  • Return null to skip the item entirely — it never reaches the writer.

Keep processors stateless and idempotent so they behave correctly when chunks are retried.

@Bean
public ItemProcessor<Customer, Customer> processor() {
    return customer -> {
        if (customer.email() == null || !customer.email().contains("@")) {
            return null; // filter invalid records out of the chunk
        }
        return new Customer(
                customer.name().trim(),
                customer.email().toLowerCase());
    };
}

Building the ItemWriter

The writer receives the whole processed chunk at once via a Chunk<O>. Writing in bulk — one batched INSERT per chunk instead of one per row — is the key performance win.

JdbcBatchItemWriter uses a parameterized SQL statement and a bean-property parameter source to map fields automatically.

@Bean
public JdbcBatchItemWriter<Customer> writer(DataSource dataSource) {
    return new JdbcBatchItemWriterBuilder<Customer>()
            .dataSource(dataSource)
            .sql("INSERT INTO customers_clean (name, email) VALUES (:name, :email)")
            .beanMapped()
            .build();
}

Wiring the Chunk Step

A Step ties the three components together. The generic types <Customer, Customer> declare the input and output of the chunk, and the integer is the commit-interval.

With a chunk size of 100, Spring Batch reads 100 items, processes each, then writes all survivors in one transaction before committing.

@Bean
public Step chunkStep(JobRepository jobRepository,
                      PlatformTransactionManager txManager,
                      ItemReader<Customer> reader,
                      ItemProcessor<Customer, Customer> processor,
                      ItemWriter<Customer> writer) {
    return new StepBuilder("chunkStep", jobRepository)
            .<Customer, Customer>chunk(100, txManager)
            .reader(reader)
            .processor(processor)
            .writer(writer)
            .build();
}

Assembling the Job

A Job is an ordered set of steps. For a single chunk step, the job simply starts with it. Spring Boot auto-detects the Job bean and runs it on startup.

The JobRepository records execution metadata — status, read/write counts, and the last committed position — enabling restartability.

@Bean
public Job importCustomersJob(JobRepository jobRepository, Step chunkStep) {
    return new JobBuilder("importCustomersJob", jobRepository)
            .start(chunkStep)
            .build();
}

Choosing the Chunk Size

The commit-interval is a tuning lever, not a magic number.

  • Too small (e.g. 1): one transaction per row — high commit overhead, slow.
  • Too large (e.g. 100,000): bigger transactions, more memory and rollback cost if a chunk fails.
  • Typical sweet spot: 100–1000, tuned by measuring throughput against your database and row size.

Remember: a failed item rolls back the entire chunk, so larger chunks mean more work redone on failure.

Fault Tolerance: Skip and Retry

Real input is messy. Wrap the step with .faultTolerant() to keep processing despite isolated failures.

  • .skip(...) — tolerate up to skipLimit bad items instead of failing the job.
  • .retry(...) — re-attempt transient errors (e.g. deadlocks) up to retryLimit before giving up.

On a skip or retry, Spring Batch re-scans the chunk item by item, isolating the offending record.

@Bean
public Step faultTolerantStep(JobRepository jobRepository,
                              PlatformTransactionManager txManager,
                              ItemReader<Customer> reader,
                              ItemProcessor<Customer, Customer> processor,
                              ItemWriter<Customer> writer) {
    return new StepBuilder("ftStep", jobRepository)
            .<Customer, Customer>chunk(100, txManager)
            .reader(reader)
            .processor(processor)
            .writer(writer)
            .faultTolerant()
            .skip(FlatFileParseException.class)
            .skipLimit(20)
            .retry(DeadlockLoserDataAccessException.class)
            .retryLimit(3)
            .build();
}

The Chunk Lifecycle in Order

Putting it together, each chunk follows the same loop inside one transaction:

  • 1. Read: call read() repeatedly until chunk-size items are buffered (or null ends input).
  • 2. Process: call process() on each item; nulls are filtered out.
  • 3. Write: pass the surviving items as one Chunk to write().
  • 4. Commit: commit the transaction and record progress in the JobRepository.

Then the loop repeats for the next chunk until the reader is exhausted.

Quick Check

Test your understanding of the chunk pipeline.

Recap

You wired a complete chunk-oriented flow in Spring Batch:

  • ItemReader streams input one item at a time, returning null at the end.
  • ItemProcessor transforms or filters items; null drops an item.
  • ItemWriter persists the whole chunk in bulk for performance.
  • The Step binds them with a chunk(size, txManager) commit-interval, and the Job orchestrates the steps.
  • Tune chunk size for throughput, and add .faultTolerant() with skip/retry for resilient pipelines.

This read-process-write loop, committed per chunk, is the backbone of scalable batch processing.

常见问题解答

「面向块的读取器—处理器—写入器流程」课时是免费的吗?

是的 — 「面向块的读取器—处理器—写入器流程」的完整文本可在网页上免费阅读。要进行交互式练习(内置代码编辑器和全天候 AI 导师)并解锁 Spring Boot 4 Complete Guide 课程的其余内容,请升级到 CoddyKit PRO。 Spring Boot 4 Complete Guide 课程共包含 4 节课。

「面向块的读取器—处理器—写入器流程」这节课中我会学到什么?

连接 ItemReader、ItemProcessor 和 ItemWriter 组件,实现可扩展的分块处理。 你通过在浏览器中直接运行的动手代码来练习 Spring Boot 4 Complete Guide,全天候 AI 导师会在你学习这节课的过程中回答你的问题。

学习 Spring Boot 4 Complete Guide 需要有经验吗?

无需任何先前经验。CoddyKit 上的 Spring Boot 4 Complete Guide 课程适合初学者到高级学习者,你可以从这里开始或从头开始,按照自己的节奏学习。 这是第 2 节课,共 4 节。

「面向块的读取器—处理器—写入器流程」课时需要多长时间?

大多数 CoddyKit 课程大约需要 5–10 分钟。每节课都很精短且互动,所以你能稳步进步,并在网页和应用中从离开的地方继续。

我能在这节 Spring Boot 4 Complete Guide 课中编写并运行代码吗?

能。每节 Spring Boot 4 Complete Guide 课都包含内置代码编辑器,你可以在浏览器中直接编写并运行真实代码,并获得即时 AI 反馈 — 无需本地设置。

此课程中的所有课时

  1. 作业、步骤与 JobRepository 模型
  2. 面向块的读取器—处理器—写入器流程
  3. 容错、跳过与重试策略
  4. 分区与并行步骤执行
← 返回 Spring Boot 4 Complete Guide