LLM Apps in Production (RAG + Vector DB + Caching) · 课时

错误处理与弹性模式

设计健壮的错误处理、重试机制和断路器,让您的 LLM 应用更能抵御故障。

第 3 / 4 课11 个步骤

错误处理与弹性模式 是 CoddyKit 上的免费 LLM Apps in Production (RAG + Vector DB + Caching) 课时。 这是第 3 节课,共 4 节。 你可以在下方免费阅读本课时的完整内容 — 然后在浏览器中使用内置代码编辑器和全天候 AI 导师进行实践。 这是 LLM Apps in Production (RAG + Vector DB + Caching) 学习路径的一部分,你的进度在网页和 CoddyKit 应用中同步。 LLM Apps in Production (RAG + Vector DB + Caching) 课程共包含 4 节课。

本课时的部分内容尚未翻译,以英文显示。

Build Robust LLM Apps

LLM applications, especially those interacting with external APIs, need to be tough!

Resilience is about designing systems that can recover from failures gracefully, without crashing or providing a bad user experience.

In this lesson, we'll learn patterns to make your LLM apps more fault-tolerant.

Typical Failures

What kind of errors can an LLM application face?

  • API Rate Limits: Too many requests at once.
  • Network Issues: Temporary connection drops.
  • LLM Service Unavailability: The LLM provider is down.
  • Bad LLM Responses: Model returns invalid JSON or hallucinates.
  • Dependency Failures: Vector DB or other services fail.

Standard Try-Catch

The first line of defense is standard error handling using try-catch blocks. This prevents your entire application from crashing when an expected error occurs.

It allows you to log the error, inform the user, or attempt a fallback.

public class Main {
  public static void main(String[] args) {
    try {
      // Simulate an LLM API call that might fail
      callLlmApi();
      System.out.println("API call successful.");
    } catch (Exception e) {
      System.out.println("Error: " + e.getMessage());
      // Log the error, notify user, etc.
    }
  }

  public static void callLlmApi() throws Exception {
    // In a real app, this would make an actual API call
    if (Math.random() < 0.5) { // 50% chance of failure
      throw new RuntimeException("LLM service unavailable.");
    }
  }
}

Why Just Catching Isn't Enough

Some errors are transient, meaning they're temporary and might resolve if you just try again. Think of a brief network glitch or a momentary rate limit.

A simple try-catch just fails immediately. For transient errors, a retry mechanism can significantly improve reliability without user intervention.

Simple Retry Logic

We can implement a basic retry loop. If an error occurs, we wait a bit and try again, up to a maximum number of attempts.

public class Main {
  public static void main(String[] args) {
    int maxRetries = 3;
    int currentRetry = 0;
    boolean success = false;

    while (currentRetry < maxRetries && !success) {
      try {
        System.out.println("Attempt " + (currentRetry + 1));
        callLlmApi();
        System.out.println("API call successful.");
        success = true;
      } catch (Exception e) {
        System.out.println("Error: " + e.getMessage());
        currentRetry++;
        if (currentRetry < maxRetries) {
          System.out.println("Retrying in 1 second...");
          try { Thread.sleep(1000); } catch (InterruptedException ie) {}
        }
      }
    }
    if (!success) {
      System.out.println("All retries failed.");
    }
  }

  public static void callLlmApi() throws Exception {
    // Simulate an LLM API call with 70% chance of failure
    if (Math.random() < 0.7) {
      throw new RuntimeException("Transient network error.");
    }
  }
}

Smart Retries: Exponential Backoff

Constant retry delays can overwhelm a struggling service. Exponential backoff is a strategy where the delay between retries increases exponentially.

This gives the remote service more time to recover and prevents your app from hammering it with requests.

  • Initial delay: 1s
  • Second delay: 2s
  • Third delay: 4s
  • And so on...

Circuit Breaker Pattern

What if a service is truly down, not just experiencing transient errors? Retrying repeatedly only wastes resources and delays failure detection.

The Circuit Breaker pattern prevents an application from repeatedly trying to invoke a service that is likely to fail, saving resources and allowing the service time to recover.

Circuit Breaker States

A circuit breaker has three main states:

  • Closed: Operations proceed normally. If errors exceed a threshold, it trips to Open.
  • Open: All requests fail immediately without trying the service. After a timeout, it transitions to Half-Open.
  • Half-Open: A limited number of requests are allowed to pass through to test if the service has recovered. If successful, it goes back to Closed; otherwise, back to Open.

Preventing Hung Requests with Timeouts

LLM API calls can sometimes hang indefinitely, waiting for a response that never comes. This can exhaust resources and degrade user experience.

Always configure timeouts for your API calls. This sets a maximum duration your application will wait for a response before giving up and throwing an error.

import java.util.concurrent.TimeUnit;

public class Main {
  public static void main(String[] args) {
    long startTime = System.nanoTime();
    long timeoutMillis = 2000; // 2 seconds timeout

    try {
      System.out.println("Calling LLM API with a timeout...");
      callLlmApiWithTimeout(timeoutMillis);
      System.out.println("API call completed successfully.");
    } catch (Exception e) {
      System.out.println("API call failed: " + e.getMessage());
    }

    long endTime = System.nanoTime();
    long duration = TimeUnit.NANOSECONDS.toMillis(endTime - startTime);
    System.out.println("Total duration: " + duration + "ms");
  }

  public static void callLlmApiWithTimeout(long timeoutMillis) throws Exception {
    // Simulate a long-running/hung API call
    long processingTime = 2500; // 2.5 seconds
    if (processingTime > timeoutMillis) {
      throw new RuntimeException("Operation timed out after " + timeoutMillis + "ms");
    }
    try {
      Thread.sleep(processingTime);
    } catch (InterruptedException e) {
      Thread.currentThread().interrupt();
      throw new RuntimeException("API call interrupted.", e);
    }
  }
}

Resilience Check

When should you use a Circuit Breaker pattern instead of just a Retry mechanism?

Recap: Building Resilient LLM Apps

We've covered key patterns for making your LLM applications fault-tolerant:

  • Basic Error Handling: Using try-catch for immediate failure management.
  • Retry Mechanisms: For handling transient errors, often with exponential backoff.
  • Circuit Breakers: To prevent overwhelming consistently failing services.
  • Timeouts: Essential for preventing hung API calls and resource exhaustion.

These patterns are crucial for robust production LLM systems!

免费开始

用 AI 导师学习 LLM Apps in Production (RAG + Vector DB + Caching) — 免费

在浏览器中编写并运行真实代码,获得全天候 AI 导师的即时帮助,并在网页或应用中继续学习。

课程
12
课程
48

常见问题解答

「错误处理与弹性模式」课时是免费的吗?

是的 — 「错误处理与弹性模式」的完整文本可在网页上免费阅读。要进行交互式练习(内置代码编辑器和全天候 AI 导师)并解锁 LLM Apps in Production (RAG + Vector DB + Caching) 课程的其余内容,请升级到 CoddyKit PRO。 LLM Apps in Production (RAG + Vector DB + Caching) 课程共包含 4 节课。

「错误处理与弹性模式」这节课中我会学到什么?

设计健壮的错误处理、重试机制和断路器,让您的 LLM 应用更能抵御故障。 你通过在浏览器中直接运行的动手代码来练习 LLM Apps in Production (RAG + Vector DB + Caching),全天候 AI 导师会在你学习这节课的过程中回答你的问题。

学习 LLM Apps in Production (RAG + Vector DB + Caching) 需要有经验吗?

无需任何先前经验。CoddyKit 上的 LLM Apps in Production (RAG + Vector DB + Caching) 课程适合初学者到高级学习者,你可以从这里开始或从头开始,按照自己的节奏学习。 这是第 3 节课,共 4 节。

「错误处理与弹性模式」课时需要多长时间?

大多数 CoddyKit 课程大约需要 5–10 分钟。每节课都很精短且互动,所以你能稳步进步,并在网页和应用中从离开的地方继续。

我能在这节 LLM Apps in Production (RAG + Vector DB + Caching) 课中编写并运行代码吗?

能。每节 LLM Apps in Production (RAG + Vector DB + Caching) 课都包含内置代码编辑器,你可以在浏览器中直接编写并运行真实代码,并获得即时 AI 反馈 — 无需本地设置。

此课程中的所有课时

  1. 保护 LLM 应用程序接口密钥与敏感数据
  2. 速率限制与滥用防护
  3. 错误处理与弹性模式
  4. 防御提示注入
← 返回 LLM Apps in Production (RAG + Vector DB + Caching)