容错、跳过与重试策略
配置跳过、重试和重启语义,使作业能够应对暂时性数据错误。
容错、跳过与重试策略 是 CoddyKit 上的免费 Spring Boot 4 Complete Guide 课时。 这是第 3 节课,共 4 节。 你可以在下方免费阅读本课时的完整内容 — 然后在浏览器中使用内置代码编辑器和全天候 AI 导师进行实践。 这是 Spring Boot 4 Complete Guide 学习路径的一部分,你的进度在网页和 CoddyKit 应用中同步。 Spring Boot 4 Complete Guide 课程共包含 4 节课。
本课时的部分内容尚未翻译,以英文显示。
Why Fault Tolerance Matters
Batch jobs process huge volumes of data, and real-world data is messy. A single malformed row, a momentary database deadlock, or a flaky downstream call can blow up a job that has already processed millions of records.
Spring Batch gives a chunk-oriented step fault tolerance so it can survive these issues without aborting the whole run. The three core tools are:
- Skip — discard records that cause unrecoverable errors (e.g. bad data) and keep going.
- Retry — re-attempt an operation that failed due to a transient error (e.g. a lock timeout).
- Restart — resume a failed job instance from where it stopped instead of starting over.
Used together, these turn a brittle job into a resilient one.
Enabling Fault Tolerance on a Step
Fault tolerance is opt-in. When you build a chunk step, call .faultTolerant() on the step builder to switch to the fault-tolerant variant. Only then can you declare skip and retry rules.
Without .faultTolerant(), any exception thrown by a reader, processor, or writer rolls back the chunk and fails the step immediately.
@Bean
public Step importStep(JobRepository jobRepository,
PlatformTransactionManager txManager,
ItemReader<Customer> reader,
ItemProcessor<Customer, Customer> processor,
ItemWriter<Customer> writer) {
return new StepBuilder("importStep", jobRepository)
.<Customer, Customer>chunk(100, txManager)
.reader(reader)
.processor(processor)
.writer(writer)
.faultTolerant() // unlocks skip & retry configuration
.build();
}Configuring Skip Policies
Skipping lets the step throw away an individual item that can never succeed — typically a parsing or validation failure — and continue with the next one.
You declare which exceptions are skippable and a global limit:
.skip(Exception.class)— mark an exception type as skippable..noSkip(Exception.class)— explicitly exclude a subtype from skipping..skipLimit(n)— total number of skips allowed before the step fails.
Once the cumulative skip count exceeds skipLimit, the step aborts. This prevents a job from silently swallowing thousands of bad records.
return new StepBuilder("importStep", jobRepository)
.<Customer, Customer>chunk(100, txManager)
.reader(reader)
.processor(processor)
.writer(writer)
.faultTolerant()
.skip(FlatFileParseException.class) // bad CSV line
.skip(ValidationException.class) // failed bean validation
.noSkip(FileNotFoundException.class) // never skip this
.skipLimit(50) // fail after 50 skips
.build();How Skip Interacts with Chunks
Skip behaviour depends on where the exception is thrown:
- Reader skip: the bad item is dropped and reading continues — cheap, no rollback.
- Processor skip: the chunk transaction rolls back, then Spring Batch re-processes the chunk item-by-item, skipping only the offending item.
- Writer skip: same scan-and-retry — the chunk rolls back and items are re-written one at a time so the single bad item can be isolated and skipped.
Because processor/writer skips trigger a rollback and a single-item replay, they are far more expensive than reader skips. Keep validation in the reader/processor where possible so failures are caught early.
A Custom SkipPolicy
The declarative .skip()/.skipLimit() API covers most cases, but you can implement SkipPolicy for fully custom logic — for example, allow more skips for one exception type than another, or inspect the exception message.
The shouldSkip method receives the thrown Throwable and the current skip count; return true to skip, or throw SkipLimitExceededException to fail the step.
public class CustomSkipPolicy implements SkipPolicy {
@Override
public boolean shouldSkip(Throwable t, long skipCount)
throws SkipLimitExceededException {
if (t instanceof FileNotFoundException) {
return false; // fatal: never skip
}
if (t instanceof ValidationException && skipCount < 100) {
return true; // tolerate up to 100 bad records
}
if (t instanceof FlatFileParseException && skipCount < 20) {
return true;
}
return false;
}
}When to Retry vs Skip
The decision between skip and retry comes down to the nature of the error:
- Retry a transient error that may succeed if attempted again: deadlock victim, lock timeout, optimistic locking conflict, brief network blip.
- Skip a deterministic error that will always fail: malformed input, a failed business validation, a constraint that the data itself violates.
Retrying a deterministic error just wastes attempts before failing; skipping a transient error throws away data that would have succeeded. Classify your exceptions correctly — this is the key design decision of the lesson.
Configuring Retry Policies
Retry re-attempts the failing operation up to a configured number of times before giving up. On a fault-tolerant step you declare:
.retry(Exception.class)— exception types that are retryable..noRetry(Exception.class)— exclude a subtype..retryLimit(n)— maximum attempts per item (including the first try).
When an item fails with a retryable exception, the chunk transaction rolls back and the item is replayed up to retryLimit times. If it still fails, the exception propagates — at which point it may be skipped if it is also declared skippable.
return new StepBuilder("importStep", jobRepository)
.<Customer, Customer>chunk(100, txManager)
.reader(reader)
.processor(processor)
.writer(writer)
.faultTolerant()
.retry(DeadlockLoserDataAccessException.class)
.retry(OptimisticLockingFailureException.class)
.retryLimit(3) // up to 3 attempts per item
.skip(ValidationException.class)
.skipLimit(50)
.build();Backoff Between Retries
Hammering a contended resource with immediate retries often makes contention worse. A backoff policy inserts a delay between attempts. ExponentialBackOffPolicy grows the wait multiplicatively, spreading out load.
You attach a custom RetryPolicy or BackOffPolicy via .retryPolicy(...) / by configuring a RetryTemplate. Below, the wait starts at 200ms and doubles each attempt up to 5s.
@Bean
public RetryTemplate retryTemplate() {
ExponentialBackOffPolicy backOff = new ExponentialBackOffPolicy();
backOff.setInitialInterval(200); // 200 ms
backOff.setMultiplier(2.0); // 200, 400, 800, ...
backOff.setMaxInterval(5000); // cap at 5 s
SimpleRetryPolicy retryPolicy = new SimpleRetryPolicy(3,
Map.of(DeadlockLoserDataAccessException.class, true));
RetryTemplate template = new RetryTemplate();
template.setBackOffPolicy(backOff);
template.setRetryPolicy(retryPolicy);
return template;
}Listeners: Observing Skips and Retries
Silently skipping records is dangerous — you need an audit trail. SkipListener callbacks fire for each skipped item so you can log it, write it to a dead-letter table, or alert.
onSkipInRead— a read failure was skipped.onSkipInProcess— gives you the item and the exception.onSkipInWrite— the item that could not be written.
Register it with .listener(skipListener) on the step builder. There is also RetryListener for observing retry attempts.
public class LoggingSkipListener implements SkipListener<Customer, Customer> {
private static final Logger log =
LoggerFactory.getLogger(LoggingSkipListener.class);
@Override
public void onSkipInRead(Throwable t) {
log.warn("Skipped unreadable record: {}", t.getMessage());
}
@Override
public void onSkipInProcess(Customer item, Throwable t) {
log.warn("Skipped {} in process: {}", item.getId(), t.getMessage());
}
@Override
public void onSkipInWrite(Customer item, Throwable t) {
log.warn("Skipped {} in write: {}", item.getId(), t.getMessage());
}
}Restartability and the Job Repository
Skip and retry handle errors within a run; restart handles a run that failed completely. Because Spring Batch persists each step's ExecutionContext and read/write counts in the job repository, relaunching the same JobInstance resumes from the last committed chunk rather than reprocessing everything.
Key rules:
- A
JobInstanceis identified by its identifying job parameters; reuse them to restart, change them to start a fresh instance. - Only jobs in a non-
COMPLETEDstate (e.g.FAILED,STOPPED) can be restarted. - Mark a step
.allowStartIfComplete(true)to force already-completed steps to re-run on restart. - Cap retries with
.startLimit(n)so a broken step is not relaunched forever.
Putting It All Together
A production-grade resilient step combines all three concerns: retry transient failures with backoff, skip deterministic bad data within a bounded limit, audit every skip, and rely on the job repository for restart.
Note the layering: an item that fails is first retried; if it still fails and the exception is skippable, it is skipped (and the listener records it). Exceptions can be both retryable and skippable — retry exhausts first, then skip applies.
return new StepBuilder("resilientImport", jobRepository)
.<Customer, Customer>chunk(100, txManager)
.reader(reader)
.processor(processor)
.writer(writer)
.faultTolerant()
// transient -> retry with backoff
.retry(DeadlockLoserDataAccessException.class)
.retryLimit(3)
// deterministic bad data -> skip
.skip(FlatFileParseException.class)
.skip(ValidationException.class)
.skipLimit(100)
// audit + restart safety
.listener(new LoggingSkipListener())
.startLimit(3)
.build();Quick Check
Test your understanding of skip vs. retry semantics.
Recap
You made a Spring Batch step resilient to transient and deterministic failures:
- Enable tolerance with
.faultTolerant()before declaring any skip/retry rules. - Skip deterministic bad data with
.skip()+.skipLimit(); processor/writer skips cost a rollback and single-item replay, so validate early. - Retry transient errors with
.retry()+.retryLimit(), adding an exponential backoff to ease contention. - Classify carefully: retry transient (deadlock, lock timeout), skip deterministic (parse/validation). When an exception is both, retry runs first, then skip.
- Audit every skip with a
SkipListenerso nothing disappears silently. - Restart failed instances from the last committed chunk via the job repository; control re-runs with
allowStartIfCompleteandstartLimit.
常见问题解答
「容错、跳过与重试策略」课时是免费的吗?
是的 — 「容错、跳过与重试策略」的完整文本可在网页上免费阅读。要进行交互式练习(内置代码编辑器和全天候 AI 导师)并解锁 Spring Boot 4 Complete Guide 课程的其余内容,请升级到 CoddyKit PRO。 Spring Boot 4 Complete Guide 课程共包含 4 节课。
「容错、跳过与重试策略」这节课中我会学到什么?
配置跳过、重试和重启语义,使作业能够应对暂时性数据错误。 你通过在浏览器中直接运行的动手代码来练习 Spring Boot 4 Complete Guide,全天候 AI 导师会在你学习这节课的过程中回答你的问题。
学习 Spring Boot 4 Complete Guide 需要有经验吗?
无需任何先前经验。CoddyKit 上的 Spring Boot 4 Complete Guide 课程适合初学者到高级学习者,你可以从这里开始或从头开始,按照自己的节奏学习。 这是第 3 节课,共 4 节。
「容错、跳过与重试策略」课时需要多长时间?
大多数 CoddyKit 课程大约需要 5–10 分钟。每节课都很精短且互动,所以你能稳步进步,并在网页和应用中从离开的地方继续。
我能在这节 Spring Boot 4 Complete Guide 课中编写并运行代码吗?
能。每节 Spring Boot 4 Complete Guide 课都包含内置代码编辑器,你可以在浏览器中直接编写并运行真实代码,并获得即时 AI 反馈 — 无需本地设置。