0Pricing
AI Engineering Academy · 강의

Redis를 활용한 정확한 캐싱

전체 프롬프트를 해시해 LLM 응답을 캐시하고, TTL과 함께 Redis에 결과를 저장하여 API 호출 없이 동일한 요청에 즉시 응답합니다.

Redis를 활용한 정확한 캐싱은(는) CoddyKit의 무료 AI Engineering Academy 강의입니다. 이것은 4개 중 1번째 강의입니다. 아래에서 전체 강의를 무료로 읽을 수 있으며, 내장 코드 에디터와 24/7 AI 튜터와 함께 브라우저에서 직접 실습할 수 있습니다. 이 강의는 AI Engineering Academy 학습 경로의 일부이며, 진행 상황이 웹과 CoddyKit 앱에 동기화됩니다. AI Engineering Academy 강의에는 총 4개의 강의가 포함되어 있습니다.

이 강의의 일부는 아직 번역되지 않았으며 영어로 표시됩니다.

Why Cache LLM Responses?

LLM API calls are expensive: a single GPT-4o request can cost $0.005-$0.15 depending on token count. In many applications, a significant fraction of incoming queries are identical or near-identical to previous ones — think FAQ bots, customer support systems, or code review tools where users ask the same questions repeatedly. Caching can eliminate 20-50 percent of API calls in these use cases, directly cutting costs and reducing latency.

Exact Cache: Cache Key Design

An exact cache stores LLM responses keyed by a deterministic hash of the input. The cache key must capture every input that affects the output: the messages array, the model name, temperature, and any other parameters that change the response. Missing any of these from the key causes cache collisions where a cached response is served for a different effective request.

import hashlib
import json

def make_cache_key(messages: list[dict], model: str, temperature: float) -> str:
    # Create a canonical, order-stable representation
    key_data = {
        'model': model,
        'temperature': temperature,
        'messages': messages,  # list order matters
    }
    # Serialize to JSON with sorted keys for determinism
    serialized = json.dumps(key_data, sort_keys=True, ensure_ascii=False)
    # Hash to a fixed-length key safe for Redis
    return 'llm_cache:' + hashlib.sha256(serialized.encode()).hexdigest()

Connecting to Redis

Redis is the standard choice for LLM response caching due to its sub-millisecond read latency and built-in TTL support. Use the redis-py library for synchronous access or aioredis (now merged into redis-py as redis.asyncio) for async access in FastAPI applications. Store the Redis connection as a singleton to avoid connection pool exhaustion.

import redis
import redis.asyncio as aioredis

# Synchronous Redis client
r = redis.Redis(
    host='localhost',
    port=6379,
    db=0,
    decode_responses=True,  # return str instead of bytes
)

# Async Redis client (for FastAPI)
async_r = aioredis.Redis(
    host='localhost',
    port=6379,
    db=0,
    decode_responses=True,
)

# Test connection
print(r.ping())  # True if Redis is running

Cache-Aside Pattern Implementation

The cache-aside pattern is the standard caching strategy for LLM APIs. On each request: (1) compute the cache key, (2) check Redis for a cached response, (3) if found (cache hit) return it immediately, (4) if not found (cache miss) call the LLM API, (5) store the response in Redis with a TTL, (6) return the response. This pattern keeps caching logic external to the LLM call itself.

import json
from openai import OpenAI

client = OpenAI()

def cached_completion(
    messages: list[dict],
    model: str = 'gpt-4o-mini',
    temperature: float = 0.7,
    ttl_seconds: int = 3600,
) -> str:
    cache_key = make_cache_key(messages, model, temperature)

    # Cache hit?
    cached = r.get(cache_key)
    if cached is not None:
        print('[CACHE HIT]')
        return json.loads(cached)

    # Cache miss: call API
    print('[CACHE MISS]')
    response = client.chat.completions.create(
        model=model,
        messages=messages,
        temperature=temperature,
    )
    result = response.choices[0].message.content

    # Store in cache with TTL
    r.setex(cache_key, ttl_seconds, json.dumps(result))
    return result

Async Cache-Aside for FastAPI

In an async FastAPI application, use the async Redis client so cache lookups do not block the event loop. The pattern is identical to the synchronous version but uses await for all Redis operations. This keeps the caching layer fully non-blocking and compatible with the async LLM client.

from openai import AsyncOpenAI
import redis.asyncio as aioredis
import json

async_client = AsyncOpenAI()
async_r = aioredis.Redis(host='localhost', port=6379, decode_responses=True)

async def async_cached_completion(
    messages: list[dict],
    model: str = 'gpt-4o-mini',
    temperature: float = 0.0,
    ttl: int = 86400,
) -> str:
    key = make_cache_key(messages, model, temperature)

    cached = await async_r.get(key)
    if cached:
        return json.loads(cached)

    response = await async_client.chat.completions.create(
        model=model, messages=messages, temperature=temperature
    )
    result = response.choices[0].message.content
    await async_r.setex(key, ttl, json.dumps(result))
    return result

Choosing the Right TTL

The TTL (Time To Live) controls how long cached responses remain valid. For factual Q&A with stable knowledge bases, long TTLs (24-72 hours) maximize cache hit rates. For responses that should reflect the latest data (news summarization, live prices), short TTLs (5-15 minutes) or no caching at all are appropriate. For creative tasks with non-zero temperature, caching may produce stale responses — consider caching only for temperature=0.

# TTL strategy by use case
TTL_STRATEGY = {
    'faq_answering':          86400 * 7,   # 7 days — stable facts
    'code_explanation':       86400,        # 1 day — code rarely changes
    'document_summarization': 3600 * 6,    # 6 hours
    'news_analysis':          300,          # 5 minutes — stale quickly
    'creative_writing':       0,            # 0 = don't cache (non-deterministic)
}

def get_ttl_for_use_case(use_case: str) -> int:
    return TTL_STRATEGY.get(use_case, 3600)  # default 1 hour

Caching Metrics and Monitoring

Track cache hit rate as a primary cost-reduction metric. A cache hit rate of 30 percent means 30 percent of API calls are avoided. Store hit and miss counts in Redis itself using INCR commands on separate counters. Expose a /metrics endpoint in your FastAPI app that reports current hit rate, total requests, and estimated cost savings to quantify the ROI of caching.

CACHE_HITS_KEY = 'llm_cache_metrics:hits'
CACHE_MISSES_KEY = 'llm_cache_metrics:misses'

async def async_cached_completion_instrumented(messages, model, temperature=0.0):
    key = make_cache_key(messages, model, temperature)
    cached = await async_r.get(key)

    if cached:
        await async_r.incr(CACHE_HITS_KEY)
        return json.loads(cached)

    await async_r.incr(CACHE_MISSES_KEY)
    response = await async_client.chat.completions.create(
        model=model, messages=messages, temperature=temperature
    )
    result = response.choices[0].message.content
    await async_r.setex(key, 3600, json.dumps(result))
    return result

async def get_cache_stats():
    hits = int(await async_r.get(CACHE_HITS_KEY) or 0)
    misses = int(await async_r.get(CACHE_MISSES_KEY) or 0)
    total = hits + misses
    return {'hit_rate': hits / total if total > 0 else 0, 'total': total}

Cache Invalidation Strategies

Exact cache invalidation is straightforward because keys are deterministic. To invalidate a specific entry, recompute its key and call r.delete(key). To invalidate all entries for a specific prompt pattern, use Redis key prefixes with a wildcard scan. To invalidate the entire cache on a major knowledge base update, call r.flushdb() (use with caution — this deletes all keys in the database).

async def invalidate_cache_entry(messages, model, temperature):
    key = make_cache_key(messages, model, temperature)
    deleted = await async_r.delete(key)
    print(f'Deleted {deleted} cache entries')

async def invalidate_all_llm_cache():
    # Scan for all keys with prefix 'llm_cache:'
    keys_to_delete = []
    async for key in async_r.scan_iter(match='llm_cache:*'):
        keys_to_delete.append(key)
    if keys_to_delete:
        await async_r.delete(*keys_to_delete)
    print(f'Invalidated {len(keys_to_delete)} cache entries')

Serializing Complex Responses

If your application caches entire API response objects (not just the text content), serialize them carefully. The full ChatCompletion object includes token usage, model version, and finish reason — useful for logging and cost tracking. Use the SDK's .model_dump_json() method to serialize Pydantic response objects to JSON strings, and reconstruct them with ChatCompletion.model_validate_json() on cache retrieval.

from openai.types.chat import ChatCompletion

async def cached_completion_full_response(
    messages, model='gpt-4o-mini', temperature=0.0
):
    key = make_cache_key(messages, model, temperature) + ':full'
    cached = await async_r.get(key)

    if cached:
        return ChatCompletion.model_validate_json(cached)  # reconstruct object

    response = await async_client.chat.completions.create(
        model=model, messages=messages, temperature=temperature
    )
    # Serialize Pydantic model to JSON
    await async_r.setex(key, 3600, response.model_dump_json())
    return response

Caching and Non-Determinism

Exact caching only makes sense for deterministic or near-deterministic requests. At temperature=0 and top_p=1.0, most LLMs produce the same output for the same input (though not guaranteed due to floating-point non-determinism). At higher temperatures, cached responses become stale as the model would have produced different outputs. Always cache at temperature=0 or document clearly in your cache key that responses may vary.

Redis Cluster and Production Setup

For production deployments with high cache volumes, use Redis Cluster for horizontal sharding across multiple nodes, or a managed Redis service like AWS ElastiCache or Redis Cloud. Set a maxmemory policy (typically allkeys-lru to evict the least recently used entries when memory is full) to prevent Redis from running out of memory and automatically manage cache size.

# Redis configuration for production LLM caching
# In redis.conf:
# maxmemory 2gb
# maxmemory-policy allkeys-lru

# Connection with retry and connection pool
import redis
from redis.retry import Retry
from redis.backoff import ExponentialBackoff

retry = Retry(ExponentialBackoff(base=0.1), 3)
production_redis = redis.Redis(
    host='your-redis-host.cache.amazonaws.com',
    port=6379,
    ssl=True,
    decode_responses=True,
    max_connections=50,
    retry=retry,
    retry_on_error=[redis.ConnectionError, redis.TimeoutError],
)

Quick Check

Test your understanding of exact LLM response caching with Redis from this lesson.

Lesson Recap

In this lesson you learned: exact caching hashes all LLM inputs to produce a deterministic cache key, the cache-aside pattern checks Redis before calling the API and stores results after a miss, and TTL selection should reflect how frequently your content changes — longer for stable knowledge, shorter for dynamic data. Monitor cache hit rate as a primary cost-reduction metric. Next up we build semantic caching for similar but non-identical queries.

자주 묻는 질문

“Redis를 활용한 정확한 캐싱” 강의는 무료인가요?

네 — “Redis를 활용한 정확한 캐싱” 전체 내용을 이 웹사이트에서 무료로 읽을 수 있습니다. 인터랙티브하게 실습하려면(내장 코드 에디터와 24/7 AI 튜터), CoddyKit PRO로 업그레이드하면 AI Engineering Academy 강의 전체를 잠금 해제할 수 있습니다. AI Engineering Academy 강의에는 총 4개의 강의가 포함되어 있습니다.

“Redis를 활용한 정확한 캐싱”에서 뭘 배우나요?

전체 프롬프트를 해시해 LLM 응답을 캐시하고, TTL과 함께 Redis에 결과를 저장하여 API 호출 없이 동일한 요청에 즉시 응답합니다. 브라우저에서 직접 실행하는 실습 코드로 AI Engineering Academy을(를) 배우며, 24/7 AI 튜터가 강의를 진행하면서 질문에 답변해줍니다.

AI Engineering Academy을(를) 시작하는 데 경험이 필요한가요?

사전 경험은 필요하지 않습니다. CoddyKit의 AI Engineering Academy은(는) 초급자부터 고급 학습자까지를 위해 구성되어 있으므로, 여기서 시작하거나 처음부터 시작할 수 있으며 자신의 속도대로 진행할 수 있습니다. 이것은 4개 중 1번째 강의입니다.

“Redis를 활용한 정확한 캐싱” 강의는 얼마나 걸리나요?

대부분의 CoddyKit 강의는 약 5~10분이 소요됩니다. 각 강의는 간결하고 인터랙티브하여 꾸준한 진행이 가능하며, 웹과 앱에서 중단한 부분부터 바로 시작할 수 있습니다.

이 AI Engineering Academy 강의에서 코드를 작성하고 실행할 수 있나요?

네. 모든 AI Engineering Academy 강의에는 내장 코드 에디터가 포함되어 있으므로, 브라우저에서 바로 실제 코드를 작성하고 실행한 후 즉시 AI 피드백을 받을 수 있습니다 — 로컬 설정이 필요 없습니다.

이 강의의 모든 강의

  1. Redis를 활용한 정확한 캐싱
  2. 임베딩을 활용한 의미 기반 캐싱
  3. OpenAI 프롬프트 접두사 캐싱
  4. 일괄 처리, 모델 라우팅, 비용 대시보드
← AI Engineering Academy(으)로 돌아가기