0Pricing
Spring Boot 4 Complete Guide · レッスン

フォールトトレランス、スキップ、リトライポリシー

一時的なデータエラーに強いジョブにするため、スキップ、リトライ、再起動のセマンティクスを設定します。

「フォールトトレランス、スキップ、リトライポリシー」はCoddyKit上の無料Spring Boot 4 Complete Guideレッスンです。 これはレッスン3/4です。 下記で完全なレッスンを無料で読むことができます。その後、ブラウザ内の組み込みコードエディタと24時間対応のAIチューターでハンズオン演習できます。 これはSpring Boot 4 Complete Guide学習パスの一部であり、ウェブとCoddyKitアプリ全体で進捗が同期されます。 Spring Boot 4 Complete Guideコースには全4レッスンが含まれています。

このレッスンの一部はまだ翻訳されておらず、英語で表示されています。

Why Fault Tolerance Matters

Batch jobs process huge volumes of data, and real-world data is messy. A single malformed row, a momentary database deadlock, or a flaky downstream call can blow up a job that has already processed millions of records.

Spring Batch gives a chunk-oriented step fault tolerance so it can survive these issues without aborting the whole run. The three core tools are:

  • Skip — discard records that cause unrecoverable errors (e.g. bad data) and keep going.
  • Retry — re-attempt an operation that failed due to a transient error (e.g. a lock timeout).
  • Restart — resume a failed job instance from where it stopped instead of starting over.

Used together, these turn a brittle job into a resilient one.

Enabling Fault Tolerance on a Step

Fault tolerance is opt-in. When you build a chunk step, call .faultTolerant() on the step builder to switch to the fault-tolerant variant. Only then can you declare skip and retry rules.

Without .faultTolerant(), any exception thrown by a reader, processor, or writer rolls back the chunk and fails the step immediately.

@Bean
public Step importStep(JobRepository jobRepository,
                        PlatformTransactionManager txManager,
                        ItemReader<Customer> reader,
                        ItemProcessor<Customer, Customer> processor,
                        ItemWriter<Customer> writer) {
    return new StepBuilder("importStep", jobRepository)
            .<Customer, Customer>chunk(100, txManager)
            .reader(reader)
            .processor(processor)
            .writer(writer)
            .faultTolerant()   // unlocks skip & retry configuration
            .build();
}

Configuring Skip Policies

Skipping lets the step throw away an individual item that can never succeed — typically a parsing or validation failure — and continue with the next one.

You declare which exceptions are skippable and a global limit:

  • .skip(Exception.class) — mark an exception type as skippable.
  • .noSkip(Exception.class) — explicitly exclude a subtype from skipping.
  • .skipLimit(n) — total number of skips allowed before the step fails.

Once the cumulative skip count exceeds skipLimit, the step aborts. This prevents a job from silently swallowing thousands of bad records.

return new StepBuilder("importStep", jobRepository)
        .<Customer, Customer>chunk(100, txManager)
        .reader(reader)
        .processor(processor)
        .writer(writer)
        .faultTolerant()
        .skip(FlatFileParseException.class)   // bad CSV line
        .skip(ValidationException.class)      // failed bean validation
        .noSkip(FileNotFoundException.class)  // never skip this
        .skipLimit(50)                        // fail after 50 skips
        .build();

How Skip Interacts with Chunks

Skip behaviour depends on where the exception is thrown:

  • Reader skip: the bad item is dropped and reading continues — cheap, no rollback.
  • Processor skip: the chunk transaction rolls back, then Spring Batch re-processes the chunk item-by-item, skipping only the offending item.
  • Writer skip: same scan-and-retry — the chunk rolls back and items are re-written one at a time so the single bad item can be isolated and skipped.

Because processor/writer skips trigger a rollback and a single-item replay, they are far more expensive than reader skips. Keep validation in the reader/processor where possible so failures are caught early.

A Custom SkipPolicy

The declarative .skip()/.skipLimit() API covers most cases, but you can implement SkipPolicy for fully custom logic — for example, allow more skips for one exception type than another, or inspect the exception message.

The shouldSkip method receives the thrown Throwable and the current skip count; return true to skip, or throw SkipLimitExceededException to fail the step.

public class CustomSkipPolicy implements SkipPolicy {

    @Override
    public boolean shouldSkip(Throwable t, long skipCount)
            throws SkipLimitExceededException {
        if (t instanceof FileNotFoundException) {
            return false; // fatal: never skip
        }
        if (t instanceof ValidationException && skipCount < 100) {
            return true;  // tolerate up to 100 bad records
        }
        if (t instanceof FlatFileParseException && skipCount < 20) {
            return true;
        }
        return false;
    }
}

When to Retry vs Skip

The decision between skip and retry comes down to the nature of the error:

  • Retry a transient error that may succeed if attempted again: deadlock victim, lock timeout, optimistic locking conflict, brief network blip.
  • Skip a deterministic error that will always fail: malformed input, a failed business validation, a constraint that the data itself violates.

Retrying a deterministic error just wastes attempts before failing; skipping a transient error throws away data that would have succeeded. Classify your exceptions correctly — this is the key design decision of the lesson.

Configuring Retry Policies

Retry re-attempts the failing operation up to a configured number of times before giving up. On a fault-tolerant step you declare:

  • .retry(Exception.class) — exception types that are retryable.
  • .noRetry(Exception.class) — exclude a subtype.
  • .retryLimit(n) — maximum attempts per item (including the first try).

When an item fails with a retryable exception, the chunk transaction rolls back and the item is replayed up to retryLimit times. If it still fails, the exception propagates — at which point it may be skipped if it is also declared skippable.

return new StepBuilder("importStep", jobRepository)
        .<Customer, Customer>chunk(100, txManager)
        .reader(reader)
        .processor(processor)
        .writer(writer)
        .faultTolerant()
        .retry(DeadlockLoserDataAccessException.class)
        .retry(OptimisticLockingFailureException.class)
        .retryLimit(3)            // up to 3 attempts per item
        .skip(ValidationException.class)
        .skipLimit(50)
        .build();

Backoff Between Retries

Hammering a contended resource with immediate retries often makes contention worse. A backoff policy inserts a delay between attempts. ExponentialBackOffPolicy grows the wait multiplicatively, spreading out load.

You attach a custom RetryPolicy or BackOffPolicy via .retryPolicy(...) / by configuring a RetryTemplate. Below, the wait starts at 200ms and doubles each attempt up to 5s.

@Bean
public RetryTemplate retryTemplate() {
    ExponentialBackOffPolicy backOff = new ExponentialBackOffPolicy();
    backOff.setInitialInterval(200);   // 200 ms
    backOff.setMultiplier(2.0);        // 200, 400, 800, ...
    backOff.setMaxInterval(5000);      // cap at 5 s

    SimpleRetryPolicy retryPolicy = new SimpleRetryPolicy(3,
            Map.of(DeadlockLoserDataAccessException.class, true));

    RetryTemplate template = new RetryTemplate();
    template.setBackOffPolicy(backOff);
    template.setRetryPolicy(retryPolicy);
    return template;
}

Listeners: Observing Skips and Retries

Silently skipping records is dangerous — you need an audit trail. SkipListener callbacks fire for each skipped item so you can log it, write it to a dead-letter table, or alert.

  • onSkipInRead — a read failure was skipped.
  • onSkipInProcess — gives you the item and the exception.
  • onSkipInWrite — the item that could not be written.

Register it with .listener(skipListener) on the step builder. There is also RetryListener for observing retry attempts.

public class LoggingSkipListener implements SkipListener<Customer, Customer> {

    private static final Logger log =
            LoggerFactory.getLogger(LoggingSkipListener.class);

    @Override
    public void onSkipInRead(Throwable t) {
        log.warn("Skipped unreadable record: {}", t.getMessage());
    }

    @Override
    public void onSkipInProcess(Customer item, Throwable t) {
        log.warn("Skipped {} in process: {}", item.getId(), t.getMessage());
    }

    @Override
    public void onSkipInWrite(Customer item, Throwable t) {
        log.warn("Skipped {} in write: {}", item.getId(), t.getMessage());
    }
}

Restartability and the Job Repository

Skip and retry handle errors within a run; restart handles a run that failed completely. Because Spring Batch persists each step's ExecutionContext and read/write counts in the job repository, relaunching the same JobInstance resumes from the last committed chunk rather than reprocessing everything.

Key rules:

  • A JobInstance is identified by its identifying job parameters; reuse them to restart, change them to start a fresh instance.
  • Only jobs in a non-COMPLETED state (e.g. FAILED, STOPPED) can be restarted.
  • Mark a step .allowStartIfComplete(true) to force already-completed steps to re-run on restart.
  • Cap retries with .startLimit(n) so a broken step is not relaunched forever.

Putting It All Together

A production-grade resilient step combines all three concerns: retry transient failures with backoff, skip deterministic bad data within a bounded limit, audit every skip, and rely on the job repository for restart.

Note the layering: an item that fails is first retried; if it still fails and the exception is skippable, it is skipped (and the listener records it). Exceptions can be both retryable and skippable — retry exhausts first, then skip applies.

return new StepBuilder("resilientImport", jobRepository)
        .<Customer, Customer>chunk(100, txManager)
        .reader(reader)
        .processor(processor)
        .writer(writer)
        .faultTolerant()
        // transient -> retry with backoff
        .retry(DeadlockLoserDataAccessException.class)
        .retryLimit(3)
        // deterministic bad data -> skip
        .skip(FlatFileParseException.class)
        .skip(ValidationException.class)
        .skipLimit(100)
        // audit + restart safety
        .listener(new LoggingSkipListener())
        .startLimit(3)
        .build();

Quick Check

Test your understanding of skip vs. retry semantics.

Recap

You made a Spring Batch step resilient to transient and deterministic failures:

  • Enable tolerance with .faultTolerant() before declaring any skip/retry rules.
  • Skip deterministic bad data with .skip() + .skipLimit(); processor/writer skips cost a rollback and single-item replay, so validate early.
  • Retry transient errors with .retry() + .retryLimit(), adding an exponential backoff to ease contention.
  • Classify carefully: retry transient (deadlock, lock timeout), skip deterministic (parse/validation). When an exception is both, retry runs first, then skip.
  • Audit every skip with a SkipListener so nothing disappears silently.
  • Restart failed instances from the last committed chunk via the job repository; control re-runs with allowStartIfComplete and startLimit.

よくある質問

「フォールトトレランス、スキップ、リトライポリシー」レッスンは無料ですか?

はい。「フォールトトレランス、スキップ、リトライポリシー」の完全なテキストはこのウェブで無料で読めます。インタラクティブに演習し(組み込みコードエディタと24時間対応のAIチューター)、Spring Boot 4 Complete Guideコースの残りをアンロックするには、CoddyKit PROにアップグレードしてください。 Spring Boot 4 Complete Guideコースには全4レッスンが含まれています。

「フォールトトレランス、スキップ、リトライポリシー」で何を学びますか?

一時的なデータエラーに強いジョブにするため、スキップ、リトライ、再起動のセマンティクスを設定します。 ブラウザで直接実行するハンズオンコードでSpring Boot 4 Complete Guideを演習し、24時間対応のAIチューターがレッスンを進める中での質問に答えます。

Spring Boot 4 Complete Guideを始めるのに経験は必要ですか?

事前経験は必要ありません。CoddyKitのSpring Boot 4 Complete Guideは初級者から上級者向けに構成されているため、ここから始めるか最初から始めて、自分のペースで進むことができます。 これはレッスン3/4です。

「フォールトトレランス、スキップ、リトライポリシー」レッスンにはどのくらい時間がかかりますか?

ほとんどのCoddyKitレッスンは約5~10分かかります。各レッスンはコンパクトでインタラクティブなので、着実に進歩し、ウェブとアプリ全体で正確に前回の場所から再開できます。

このSpring Boot 4 Complete Guideレッスンでコードを書いて実行できますか?

はい。すべてのSpring Boot 4 Complete Guideレッスンに組み込みコードエディタが含まれているため、ブラウザでリアルコードを書いて実行し、即座のAIフィードバックを取得できます。ローカル設定は不要です。

このコースのすべてのレッスン

  1. ジョブ、ステップ、JobRepositoryモデル
  2. チャンク指向のReader-Processor-Writerフロー
  3. フォールトトレランス、スキップ、リトライポリシー
  4. パーティショニングとステップの並列実行
← Spring Boot 4 Complete Guideに戻る