Spring Boot 4 Complete Guide · Lezione

Tolleranza ai guasti, skip e policy di retry

Configuri semantiche di skip, retry e riavvio per rendere i job resilienti agli errori transitori dei dati.

Lezione 3 di 413 passaggi

Tolleranza ai guasti, skip e policy di retry è una lezione Spring Boot 4 Complete Guide gratuita su CoddyKit. Questa è la lezione 3 di 4. Puoi leggere la lezione completa qui gratuitamente — poi esercitati direttamente nel browser con un editor di codice integrato e un tutor IA disponibile 24/7. Fa parte del percorso di apprendimento Spring Boot 4 Complete Guide, e i tuoi progressi si sincronizzano tra il web e l'app CoddyKit. Il corso Spring Boot 4 Complete Guide include 4 lezioni in totale.

Parti di questa lezione non sono ancora state tradotte e vengono mostrate in inglese.

Why Fault Tolerance Matters

Batch jobs process huge volumes of data, and real-world data is messy. A single malformed row, a momentary database deadlock, or a flaky downstream call can blow up a job that has already processed millions of records.

Spring Batch gives a chunk-oriented step fault tolerance so it can survive these issues without aborting the whole run. The three core tools are:

  • Skip — discard records that cause unrecoverable errors (e.g. bad data) and keep going.
  • Retry — re-attempt an operation that failed due to a transient error (e.g. a lock timeout).
  • Restart — resume a failed job instance from where it stopped instead of starting over.

Used together, these turn a brittle job into a resilient one.

Enabling Fault Tolerance on a Step

Fault tolerance is opt-in. When you build a chunk step, call .faultTolerant() on the step builder to switch to the fault-tolerant variant. Only then can you declare skip and retry rules.

Without .faultTolerant(), any exception thrown by a reader, processor, or writer rolls back the chunk and fails the step immediately.

@Bean
public Step importStep(JobRepository jobRepository,
                        PlatformTransactionManager txManager,
                        ItemReader<Customer> reader,
                        ItemProcessor<Customer, Customer> processor,
                        ItemWriter<Customer> writer) {
    return new StepBuilder("importStep", jobRepository)
            .<Customer, Customer>chunk(100, txManager)
            .reader(reader)
            .processor(processor)
            .writer(writer)
            .faultTolerant()   // unlocks skip & retry configuration
            .build();
}

Configuring Skip Policies

Skipping lets the step throw away an individual item that can never succeed — typically a parsing or validation failure — and continue with the next one.

You declare which exceptions are skippable and a global limit:

  • .skip(Exception.class) — mark an exception type as skippable.
  • .noSkip(Exception.class) — explicitly exclude a subtype from skipping.
  • .skipLimit(n) — total number of skips allowed before the step fails.

Once the cumulative skip count exceeds skipLimit, the step aborts. This prevents a job from silently swallowing thousands of bad records.

return new StepBuilder("importStep", jobRepository)
        .<Customer, Customer>chunk(100, txManager)
        .reader(reader)
        .processor(processor)
        .writer(writer)
        .faultTolerant()
        .skip(FlatFileParseException.class)   // bad CSV line
        .skip(ValidationException.class)      // failed bean validation
        .noSkip(FileNotFoundException.class)  // never skip this
        .skipLimit(50)                        // fail after 50 skips
        .build();

How Skip Interacts with Chunks

Skip behaviour depends on where the exception is thrown:

  • Reader skip: the bad item is dropped and reading continues — cheap, no rollback.
  • Processor skip: the chunk transaction rolls back, then Spring Batch re-processes the chunk item-by-item, skipping only the offending item.
  • Writer skip: same scan-and-retry — the chunk rolls back and items are re-written one at a time so the single bad item can be isolated and skipped.

Because processor/writer skips trigger a rollback and a single-item replay, they are far more expensive than reader skips. Keep validation in the reader/processor where possible so failures are caught early.

A Custom SkipPolicy

The declarative .skip()/.skipLimit() API covers most cases, but you can implement SkipPolicy for fully custom logic — for example, allow more skips for one exception type than another, or inspect the exception message.

The shouldSkip method receives the thrown Throwable and the current skip count; return true to skip, or throw SkipLimitExceededException to fail the step.

public class CustomSkipPolicy implements SkipPolicy {

    @Override
    public boolean shouldSkip(Throwable t, long skipCount)
            throws SkipLimitExceededException {
        if (t instanceof FileNotFoundException) {
            return false; // fatal: never skip
        }
        if (t instanceof ValidationException && skipCount < 100) {
            return true;  // tolerate up to 100 bad records
        }
        if (t instanceof FlatFileParseException && skipCount < 20) {
            return true;
        }
        return false;
    }
}

When to Retry vs Skip

The decision between skip and retry comes down to the nature of the error:

  • Retry a transient error that may succeed if attempted again: deadlock victim, lock timeout, optimistic locking conflict, brief network blip.
  • Skip a deterministic error that will always fail: malformed input, a failed business validation, a constraint that the data itself violates.

Retrying a deterministic error just wastes attempts before failing; skipping a transient error throws away data that would have succeeded. Classify your exceptions correctly — this is the key design decision of the lesson.

Configuring Retry Policies

Retry re-attempts the failing operation up to a configured number of times before giving up. On a fault-tolerant step you declare:

  • .retry(Exception.class) — exception types that are retryable.
  • .noRetry(Exception.class) — exclude a subtype.
  • .retryLimit(n) — maximum attempts per item (including the first try).

When an item fails with a retryable exception, the chunk transaction rolls back and the item is replayed up to retryLimit times. If it still fails, the exception propagates — at which point it may be skipped if it is also declared skippable.

return new StepBuilder("importStep", jobRepository)
        .<Customer, Customer>chunk(100, txManager)
        .reader(reader)
        .processor(processor)
        .writer(writer)
        .faultTolerant()
        .retry(DeadlockLoserDataAccessException.class)
        .retry(OptimisticLockingFailureException.class)
        .retryLimit(3)            // up to 3 attempts per item
        .skip(ValidationException.class)
        .skipLimit(50)
        .build();

Backoff Between Retries

Hammering a contended resource with immediate retries often makes contention worse. A backoff policy inserts a delay between attempts. ExponentialBackOffPolicy grows the wait multiplicatively, spreading out load.

You attach a custom RetryPolicy or BackOffPolicy via .retryPolicy(...) / by configuring a RetryTemplate. Below, the wait starts at 200ms and doubles each attempt up to 5s.

@Bean
public RetryTemplate retryTemplate() {
    ExponentialBackOffPolicy backOff = new ExponentialBackOffPolicy();
    backOff.setInitialInterval(200);   // 200 ms
    backOff.setMultiplier(2.0);        // 200, 400, 800, ...
    backOff.setMaxInterval(5000);      // cap at 5 s

    SimpleRetryPolicy retryPolicy = new SimpleRetryPolicy(3,
            Map.of(DeadlockLoserDataAccessException.class, true));

    RetryTemplate template = new RetryTemplate();
    template.setBackOffPolicy(backOff);
    template.setRetryPolicy(retryPolicy);
    return template;
}

Listeners: Observing Skips and Retries

Silently skipping records is dangerous — you need an audit trail. SkipListener callbacks fire for each skipped item so you can log it, write it to a dead-letter table, or alert.

  • onSkipInRead — a read failure was skipped.
  • onSkipInProcess — gives you the item and the exception.
  • onSkipInWrite — the item that could not be written.

Register it with .listener(skipListener) on the step builder. There is also RetryListener for observing retry attempts.

public class LoggingSkipListener implements SkipListener<Customer, Customer> {

    private static final Logger log =
            LoggerFactory.getLogger(LoggingSkipListener.class);

    @Override
    public void onSkipInRead(Throwable t) {
        log.warn("Skipped unreadable record: {}", t.getMessage());
    }

    @Override
    public void onSkipInProcess(Customer item, Throwable t) {
        log.warn("Skipped {} in process: {}", item.getId(), t.getMessage());
    }

    @Override
    public void onSkipInWrite(Customer item, Throwable t) {
        log.warn("Skipped {} in write: {}", item.getId(), t.getMessage());
    }
}

Restartability and the Job Repository

Skip and retry handle errors within a run; restart handles a run that failed completely. Because Spring Batch persists each step's ExecutionContext and read/write counts in the job repository, relaunching the same JobInstance resumes from the last committed chunk rather than reprocessing everything.

Key rules:

  • A JobInstance is identified by its identifying job parameters; reuse them to restart, change them to start a fresh instance.
  • Only jobs in a non-COMPLETED state (e.g. FAILED, STOPPED) can be restarted.
  • Mark a step .allowStartIfComplete(true) to force already-completed steps to re-run on restart.
  • Cap retries with .startLimit(n) so a broken step is not relaunched forever.

Putting It All Together

A production-grade resilient step combines all three concerns: retry transient failures with backoff, skip deterministic bad data within a bounded limit, audit every skip, and rely on the job repository for restart.

Note the layering: an item that fails is first retried; if it still fails and the exception is skippable, it is skipped (and the listener records it). Exceptions can be both retryable and skippable — retry exhausts first, then skip applies.

return new StepBuilder("resilientImport", jobRepository)
        .<Customer, Customer>chunk(100, txManager)
        .reader(reader)
        .processor(processor)
        .writer(writer)
        .faultTolerant()
        // transient -> retry with backoff
        .retry(DeadlockLoserDataAccessException.class)
        .retryLimit(3)
        // deterministic bad data -> skip
        .skip(FlatFileParseException.class)
        .skip(ValidationException.class)
        .skipLimit(100)
        // audit + restart safety
        .listener(new LoggingSkipListener())
        .startLimit(3)
        .build();

Quick Check

Test your understanding of skip vs. retry semantics.

Recap

You made a Spring Batch step resilient to transient and deterministic failures:

  • Enable tolerance with .faultTolerant() before declaring any skip/retry rules.
  • Skip deterministic bad data with .skip() + .skipLimit(); processor/writer skips cost a rollback and single-item replay, so validate early.
  • Retry transient errors with .retry() + .retryLimit(), adding an exponential backoff to ease contention.
  • Classify carefully: retry transient (deadlock, lock timeout), skip deterministic (parse/validation). When an exception is both, retry runs first, then skip.
  • Audit every skip with a SkipListener so nothing disappears silently.
  • Restart failed instances from the last committed chunk via the job repository; control re-runs with allowStartIfComplete and startLimit.
Gratis per iniziare

Impara Java con un tutor IA — gratis

Scrivi ed esegui vero codice nel tuo browser, ricevi aiuto istantaneo da un tutor IA disponibile 24/7, e riprendi da dove hai lasciato sul web o nell'app.

Corsi
21
Lezioni
84

Domande Frequenti

La lezione «Tolleranza ai guasti, skip e policy di retry» è gratuita?

Sì — il testo completo di «Tolleranza ai guasti, skip e policy di retry» è gratuito qui sul web. Per esercitarvi in modo interattivo (un editor di codice integrato e un tutor IA 24/7) e sbloccare il resto del corso Spring Boot 4 Complete Guide, passa a CoddyKit PRO. Il corso Spring Boot 4 Complete Guide include 4 lezioni in totale.

Cosa imparerò in «Tolleranza ai guasti, skip e policy di retry»?

Configuri semantiche di skip, retry e riavvio per rendere i job resilienti agli errori transitori dei dati. Eserciti Spring Boot 4 Complete Guide con codice pratico che esegui direttamente nel browser, e un tutor IA 24/7 risponde alle tue domande mentre lavori sulla lezione.

Ho bisogno di esperienza per iniziare Spring Boot 4 Complete Guide?

Non è richiesta alcuna esperienza precedente. Spring Boot 4 Complete Guide su CoddyKit è strutturato per principianti e studenti avanzati, quindi puoi iniziare da qui o dall'inizio e procedere al tuo ritmo. Questa è la lezione 3 di 4.

Quanto tempo richiede la lezione «Tolleranza ai guasti, skip e policy di retry»?

La maggior parte delle lezioni CoddyKit richiede circa 5–10 minuti. Ogni lezione è breve e interattiva, quindi fai progressi costanti e riprendi esattamente da dove hai lasciato su web e app.

Posso scrivere ed eseguire codice in questa lezione Spring Boot 4 Complete Guide?

Sì. Ogni lezione Spring Boot 4 Complete Guide include un editor di codice integrato, quindi scrivi ed esegui codice reale direttamente nel tuo browser e ricevi feedback istantaneo dall'IA — nessuna configurazione locale necessaria.

Tutte le lezioni di questo corso

  1. Job, step e modello JobRepository
  2. Flussi Reader-Processor-Writer orientati ai chunk
  3. Tolleranza ai guasti, skip e policy di retry
  4. Partizionamento ed esecuzione parallela degli step
← Torna a Spring Boot 4 Complete Guide