복원력 있는 서버리스 시스템 구축
함수 전반에 회로 차단기, 재시도, 멱등성과 같은 패턴을 적용하여 가용성이 높고 장애에 강한 서버리스 아키텍처를 설계합니다.
복원력 있는 서버리스 시스템 구축은(는) CoddyKit의 무료 Serverless AWS Lambda Development 강의입니다. 이것은 4개 중 2번째 강의입니다. 아래에서 전체 강의를 무료로 읽을 수 있으며, 내장 코드 에디터와 24/7 AI 튜터와 함께 브라우저에서 직접 실습할 수 있습니다. 이 강의는 Serverless AWS Lambda Development 학습 경로의 일부이며, 진행 상황이 웹과 CoddyKit 앱에 동기화됩니다. Serverless AWS Lambda Development 강의에는 총 4개의 강의가 포함되어 있습니다.
이 강의의 일부는 아직 번역되지 않았으며 영어로 표시됩니다.
Building Robust Serverless Systems
Welcome! In this lesson, we'll dive into designing highly resilient and fault-tolerant serverless applications. Even though AWS manages much of the infrastructure, your functions still need to handle failures gracefully.
We'll explore key architectural patterns to ensure your applications remain stable and performant, even when things go wrong.
The Reality of Distributed Systems
In a serverless world, your functions often interact with many other services: databases, APIs, message queues. These interactions happen over a network, and networks can be unreliable.
- Transient Failures: Brief network glitches or service slowdowns.
- Downstream Service Issues: A service your Lambda calls might be temporarily unavailable.
- Unexpected Data: Malformed input can cause your function to crash.
Designing for these "failures" is crucial for a stable system.
Lambda's Built-in Retry Logic
For certain invocation types, AWS Lambda automatically retries your function if it fails. This is a powerful built-in resilience mechanism for asynchronous invocations.
For example, if an SQS queue triggers your Lambda and your function errors, Lambda (or SQS) will retry the invocation a few times. This helps overcome transient issues without any code changes.
However, retries aren't a silver bullet; they can lead to duplicate processing if not handled carefully.
Making Operations Idempotent
When retries happen, your function might execute the same operation multiple times. This is where idempotency comes in.
An idempotent operation is one that can be applied multiple times without changing the result beyond the initial application.
- Example: Setting a value (
x = 5) is idempotent. - Non-Example: Incrementing a value (
x++) is NOT idempotent, as each retry would change the value.
For resilient systems, many operations should strive to be idempotent.
Keys to Idempotent Functions
To make your Lambda functions idempotent, you often need to track the state of a request. This typically involves:
- Generating a unique Idempotency Key for each request (e.g., from request ID, event source ID).
- Checking if this key has already been processed before performing the core logic.
- Storing the result or status of the operation associated with the key.
This ensures that even if a function is retried, the core side-effect only occurs once.
import hashlib
import json
# Imagine a database or cache for storing processed requests
# For simplicity, using a global dict here. A real app uses persistent storage.
processed_requests = {}
def is_idempotent(event_payload):
# Create a unique key from the event payload
# For a real app, use a proper hashing/unique ID strategy
event_hash = hashlib.md5(json.dumps(event_payload, sort_keys=True).encode('utf-8')).hexdigest()
if event_hash in processed_requests:
print(f"Request with hash {event_hash} already processed.")
return True
processed_requests[event_hash] = "processing" # Mark as processing
return False
def lambda_handler(event, context):
if is_idempotent(event):
return {
'statusCode': 200,
'body': json.dumps('Request already processed or is being processed.')
}
# Simulate actual work (e.g., writing to a database)
print(f"Processing new request: {event}")
# In a real scenario, update processed_requests[event_hash] = "completed"
# after successful processing and store the result in persistent storage.
return {
'statusCode': 200,
'body': json.dumps('Request processed successfully!')
}
Preventing Cascading Failures
The Circuit Breaker pattern is a powerful way to prevent a failing service from causing cascading failures throughout your application.
Imagine a call to an external API that starts failing. Continuously retrying it will just waste resources and slow down your function. A circuit breaker detects this and "opens" the circuit, stopping calls to the failing service temporarily.
This gives the failing service time to recover and prevents your application from getting bogged down.
How a Circuit Breaker Works
A circuit breaker typically has three states:
- Closed: Operations proceed as normal. If failures exceed a threshold, it transitions to Open.
- Open: All calls to the protected service immediately fail (or return a fallback). After a timeout, it transitions to Half-Open.
- Half-Open: A limited number of test calls are allowed through. If these succeed, it transitions back to Closed. If they fail, it returns to Open.
This intelligent behavior allows for self-healing.
Implementing a Simple Circuit Breaker
Implementing a full circuit breaker involves managing state (failures, success counts, last failure time). For serverless, this state might be stored in a shared cache (like ElastiCache) or a database.
While complex to implement from scratch in a simple Lambda, understanding the logic is key. Libraries exist for various languages to help, or you can leverage AWS services like Step Functions to orchestrate retry logic with delays.
import time
class CircuitBreaker:
def __init__(self, failure_threshold=3, reset_timeout=5):
self.state = "CLOSED"
self.failure_count = 0
self.last_failure_time = 0
self.failure_threshold = failure_threshold
self.reset_timeout = reset_timeout # seconds
def call(self, func, *args, **kwargs):
if self.state == "OPEN":
if time.time() - self.last_failure_time > self.reset_timeout:
self.state = "HALF-OPEN"
# In a real app, log this state change
else:
raise Exception("Circuit is open, service unavailable.")
try:
result = func(*args, **kwargs)
if self.state == "HALF-OPEN":
self.state = "CLOSED"
self.failure_count = 0
# In a real app, log this state change
return result
except Exception as e:
self.failure_count += 1
self.last_failure_time = time.time()
if self.failure_count >= self.failure_threshold:
self.state = "OPEN"
# In a real app, log this state change
raise e
Timeouts Prevent Hanging
Another crucial resilience pattern is using timeouts for external calls. If your Lambda function calls another service (e.g., a database, an HTTP API), that call could hang indefinitely if the service is unresponsive.
Configuring a timeout ensures your function doesn't wait forever, freeing up resources and allowing for retry logic to kick in faster. AWS Lambda itself has a configurable timeout, but you should also set timeouts within your code for specific external requests.
Resilient Design Challenge
Consider a Lambda function that processes incoming orders. If the function fails after successfully deducting payment but before updating the order status in a database, and then retries, what problem could arise if the payment deduction is NOT idempotent?
Summary: Building for Failure
We've covered essential patterns for building resilient serverless applications:
- Retries: Lambda's built-in mechanism for transient errors.
- Idempotency: Ensuring operations can be safely retried without unintended side-effects (e.g., duplicate charges).
- Circuit Breakers: Preventing cascading failures by intelligently stopping calls to failing services.
- Timeouts: Protecting against unresponsive external services.
By applying these principles, you can create serverless systems that gracefully handle inevitable failures.
자주 묻는 질문
“복원력 있는 서버리스 시스템 구축” 강의는 무료인가요?
네 — “복원력 있는 서버리스 시스템 구축” 전체 내용을 이 웹사이트에서 무료로 읽을 수 있습니다. 인터랙티브하게 실습하려면(내장 코드 에디터와 24/7 AI 튜터), CoddyKit PRO로 업그레이드하면 Serverless AWS Lambda Development 강의 전체를 잠금 해제할 수 있습니다. Serverless AWS Lambda Development 강의에는 총 4개의 강의가 포함되어 있습니다.
“복원력 있는 서버리스 시스템 구축”에서 뭘 배우나요?
함수 전반에 회로 차단기, 재시도, 멱등성과 같은 패턴을 적용하여 가용성이 높고 장애에 강한 서버리스 아키텍처를 설계합니다. 브라우저에서 직접 실행하는 실습 코드로 Serverless AWS Lambda Development을(를) 배우며, 24/7 AI 튜터가 강의를 진행하면서 질문에 답변해줍니다.
Serverless AWS Lambda Development을(를) 시작하는 데 경험이 필요한가요?
사전 경험은 필요하지 않습니다. CoddyKit의 Serverless AWS Lambda Development은(는) 초급자부터 고급 학습자까지를 위해 구성되어 있으므로, 여기서 시작하거나 처음부터 시작할 수 있으며 자신의 속도대로 진행할 수 있습니다. 이것은 4개 중 2번째 강의입니다.
“복원력 있는 서버리스 시스템 구축” 강의는 얼마나 걸리나요?
대부분의 CoddyKit 강의는 약 5~10분이 소요됩니다. 각 강의는 간결하고 인터랙티브하여 꾸준한 진행이 가능하며, 웹과 앱에서 중단한 부분부터 바로 시작할 수 있습니다.
이 Serverless AWS Lambda Development 강의에서 코드를 작성하고 실행할 수 있나요?
네. 모든 Serverless AWS Lambda Development 강의에는 내장 코드 에디터가 포함되어 있으므로, 브라우저에서 바로 실제 코드를 작성하고 실행한 후 즉시 AI 피드백을 받을 수 있습니다 — 로컬 설정이 필요 없습니다.
이 강의의 모든 강의
- 카나리 및 블루/그린 배포
- 복원력 있는 서버리스 시스템 구축
- 서버리스 아키텍처 패턴
- 서버리스 아키텍처의 비용 최적화