0Pricing
AI Agents with LangChain & Autonomous Workflows · Leçon

Gestion des erreurs et résilience

Élaborez des stratégies robustes pour anticiper, intercepter et gérer élégamment les erreurs dans les flux de travail autonomes des agents.

Gestion des erreurs et résilience est une leçon AI Agents with LangChain & Autonomous Workflows gratuite sur CoddyKit. Ceci est la leçon 3 sur 4. Tu peux lire la leçon complète ci-dessous gratuitement — puis la pratiquer en direct dans le navigateur avec un éditeur de code intégré et un tuteur IA 24/7. Elle fait partie du parcours d'apprentissage AI Agents with LangChain & Autonomous Workflows, et ta progression se synchronise sur le web et l'application CoddyKit. Le cours AI Agents with LangChain & Autonomous Workflows comprend 4 leçons au total.

Certaines parties de cette leçon n'ont pas encore été traduites et s'affichent en anglais.

Why Error Handling Matters

Autonomous agents perform complex tasks, often interacting with external services or making decisions based on potentially unreliable information. What happens when things go wrong?

Error handling is crucial for agents to be reliable and robust. It ensures your agent can recover from unexpected issues, prevent crashes, and maintain a consistent user experience.

Common Agent Workflow Errors

Agents can encounter various types of errors during their operation:

  • API Failures: Large Language Model (LLM) providers or external tools might experience downtime, rate limits, or authentication issues.
  • Tool Execution Issues: A custom or pre-built tool might receive bad input, fail to execute correctly, or return an unexpected format.
  • LLM Misinterpretations: The LLM might generate unparseable output, hallucinate, or respond in a way the agent's logic cannot handle.
  • Network Issues: Connectivity problems to external services can prevent agents from fetching data or calling APIs.

Catching Errors with Try-Except

In Python, the try-except block is your fundamental mechanism to catch errors. It allows you to attempt an operation and gracefully handle specific exceptions if they occur.

This prevents your entire agent workflow from crashing due to a single failure point.

def perform_risky_operation(value):
    try:
        # Attempt a potentially failing operation
        result = 100 / value
        print(f"Operation successful! Result: {result}")
    except ZeroDivisionError:
        # Handle specific error: division by zero
        print("Error: Cannot divide by zero!")
    except TypeError as e:
        # Handle specific error: incorrect type
        print(f"Error: Invalid input type - {e}")
    except Exception as e:
        # Catch any other unexpected errors
        print(f"An unexpected error occurred: {e}")

if __name__ == "__main__":
    perform_risky_operation(20) # Works fine
    perform_risky_operation(0)  # Catches ZeroDivisionError
    perform_risky_operation("abc") # Catches TypeError

Handling LLM API Errors

When your agent interacts with an LLM (e.g., OpenAI, Anthropic), API calls can fail. These failures could be due to rate limits, invalid API keys, or temporary service outages.

It's vital to catch these specific API errors to implement recovery strategies or inform the user.

import random

# Simulate a custom API error for demonstration
class LLMAPIError(Exception):
    pass

def call_llm_service(prompt):
    # Simulate a 25% chance of API failure
    if random.random() < 0.25:
        raise LLMAPIError("LLM API call failed: Service unavailable.")
    return f"LLM response to '{prompt}': Here's your answer."

if __name__ == "__main__":
    print("--- Attempt 1 ---")
    try:
        response = call_llm_service("Summarize the news.")
        print(response)
    except LLMAPIError as e:
        print(f"Caught LLM API Error: {e}")
    except Exception as e:
        print(f"An unexpected error occurred: {e}")

    print("\n--- Attempt 2 ---")
    try:
        response = call_llm_service("Write a haiku.")
        print(response)
    except LLMAPIError as e:
        print(f"Caught LLM API Error: {e}")
    except Exception as e:
        print(f"An unexpected error occurred: {e}")

Robust Tool Execution Errors

Agents use tools to extend their capabilities (e.g., searching the web, executing code). A tool might fail if its external dependency is down, it receives invalid input, or encounters an internal error.

By anticipating and handling these tool-specific errors, your agent can decide on alternative actions or provide helpful feedback.

import random

# Simulate a custom tool execution error
class WebSearchToolError(Exception):
    pass

def perform_web_search(query):
    # Simulate a 30% chance of tool failure
    if random.random() < 0.3:
        raise WebSearchToolError(f"Web search for '{query}' failed due to network issues.")
    return f"Web search results for: {query}"

if __name__ == "__main__":
    print("--- Search 1 ---")
    try:
        result = perform_web_search("current weather")
        print(result)
    except WebSearchToolError as e:
        print(f"Caught Web Search Tool Error: {e}")
    except Exception as e:
        print(f"An unexpected error occurred: {e}")

    print("\n--- Search 2 ---")
    try:
        result = perform_web_search("AI agent frameworks")
        print(result)
    except WebSearchToolError as e:
        print(f"Caught Web Search Tool Error: {e}")
    except Exception as e:
        print(f"An unexpected error occurred: {e}")

Implementing Retries with Backoff

Many errors are transient, meaning they are temporary and might resolve themselves. For these, a simple retry mechanism can be highly effective. Exponential backoff is a common strategy where the delay between retries increases with each attempt.

This prevents overwhelming a failing service and gives it time to recover.

import time
import random

def retry_with_backoff(func, max_retries=3):
    for attempt in range(max_retries):
        try:
            return func() # Try to execute the function
        except Exception as e:
            print(f"Attempt {attempt + 1} failed: {e}")
            if attempt < max_retries - 1:
                # Calculate exponential backoff delay
                wait_time = 2 ** attempt
                print(f"Retrying in {wait_time} seconds...")
                time.sleep(wait_time)
            else:
                # Re-raise error if max retries reached
                raise ValueError("Operation failed after multiple retries.")

def unreliable_action():
    # Simulate an action that fails 60% of the time
    if random.random() < 0.6:
        raise ConnectionError("Temporary network issue.")
    return "Action completed successfully!"

if __name__ == "__main__":
    try:
        result = retry_with_backoff(unreliable_action)
        print(result)
    except ValueError as e:
        print(f"Final result: {e}")

LangChain Callbacks for Errors

LangChain's Callback system provides a powerful way to inject custom logic into various stages of an agent or chain's execution, including error handling.

  • You can define functions that run specifically when an error occurs (e.g., on_tool_error, on_chain_error).
  • This allows for centralized logging, monitoring, or triggering alerts when issues arise.
  • Callbacks can capture detailed context about the error, aiding in debugging complex agent workflows.

Graceful Degradation Strategies

Not all errors are recoverable. Sometimes, an agent needs to degrade gracefully rather than completely failing. This means providing a reduced but still functional experience.

  • Fallback Mechanisms: If a primary, complex tool fails, switch to a simpler, more reliable alternative (e.g., if a specialized database search fails, fall back to a general web search).
  • Partial Completion: Complete as much of the task as possible and inform the user about the limitations or incomplete parts.
  • Informative User Messages: Clearly communicate to the user when a specific feature or capability is temporarily unavailable due to an underlying issue.

Logging Errors for Observability

Effective logging is crucial for understanding why an autonomous agent failed, especially in production environments. Good logs provide observability into your agent's internal workings.

  • What to Log: Include error messages, stack traces, relevant input parameters, the agent's current state, and timestamps.
  • Where to Log: Send logs to a centralized logging system (e.g., ELK stack, Splunk, cloud logging services) for easy analysis and alerting.
  • Why it's Important: Helps identify recurring issues, debug complex interactions, and monitor the overall health and reliability of your agent system.

Error Handling Check

Let's test your understanding of error handling and resilience in autonomous agent workflows.

Recap: Building Resilient Agents

We've explored how to make autonomous agent workflows more robust by handling errors effectively.

  • We covered using try-except blocks for basic error catching and managing specific types of exceptions.
  • We discussed specific strategies for handling LLM API and tool execution errors.
  • We learned about implementing retries with exponential backoff to overcome transient issues.
  • Finally, we touched upon graceful degradation for unrecoverable errors and the importance of logging for observability and debugging.

By applying these techniques, your agents can better withstand unexpected issues and provide a more reliable and stable user experience.

Questions Fréquemment Posées

La leçon « Gestion des erreurs et résilience » est-elle gratuite ?

Oui — le texte complet de « Gestion des erreurs et résilience » est gratuit à lire ici sur le web. Pour la pratiquer de manière interactive (un éditeur de code intégré et un tuteur IA 24/7) et déverrouiller le reste du cours AI Agents with LangChain & Autonomous Workflows, passe à CoddyKit PRO. Le cours AI Agents with LangChain & Autonomous Workflows comprend 4 leçons au total.

Qu'est-ce que j'apprendrai dans « Gestion des erreurs et résilience » ?

Élaborez des stratégies robustes pour anticiper, intercepter et gérer élégamment les erreurs dans les flux de travail autonomes des agents. Tu pratiques AI Agents with LangChain & Autonomous Workflows avec du code pratique que tu exécutes directement dans le navigateur, et un tuteur IA 24/7 répond à tes questions au fur et à mesure que tu avances dans la leçon.

Dois-je avoir de l'expérience pour commencer AI Agents with LangChain & Autonomous Workflows ?

Aucune expérience préalable n'est requise. AI Agents with LangChain & Autonomous Workflows sur CoddyKit est structuré pour les débutants jusqu'aux apprenants avancés, donc tu peux commencer ici ou depuis le début et avancer à ton rythme. Ceci est la leçon 3 sur 4.

Combien de temps prend la leçon « Gestion des erreurs et résilience » ?

La plupart des leçons CoddyKit prennent environ 5–10 minutes. Chacune est courte et interactive, tu progresses régulièrement et tu repiques exactement où tu t'es arrêté sur le web et l'app.

Peux-tu écrire et exécuter du code dans cette leçon AI Agents with LangChain & Autonomous Workflows ?

Oui. Chaque leçon AI Agents with LangChain & Autonomous Workflows inclut un éditeur de code intégré, tu écris et exécutes du vrai code directement dans ton navigateur et tu reçois des retours IA instantanés — aucune configuration locale requise.

Toutes les leçons de ce cours

  1. Concevoir des flux de travail complexes
  2. Exécution asynchrone des agents
  3. Gestion des erreurs et résilience
  4. Approbations avec intervention humaine
← Retour à AI Agents with LangChain & Autonomous Workflows