0Pricing
AI Prompt Engineering · Lesson

Caching Strategies for Prompts

Semantic caching, exact-match caching, and Anthropic prompt caching.

Caching Strategies for Prompts is a free AI Prompt Engineering lesson on CoddyKit — lesson 1 of 4. You can read the complete lesson below for free — then practise it hands-on in the browser with a built-in code editor and a 24/7 AI tutor. It is part of the AI Prompt Engineering learning path, one of 4 lessons in the course, and your progress syncs across the web and the CoddyKit app.

Why Cache Prompt Results?

LLM API calls are expensive and slow. Many production applications send the same (or very similar) prompts repeatedly. Caching returns stored results for repeated queries, eliminating redundant API calls and reducing both cost and latency dramatically.

Exact-Match Caching with Hash Keys

The simplest cache: hash the exact prompt string and store the result. If the same prompt string appears again, return the cached result without calling the API.

import hashlib
import json
from functools import lru_cache

class ExactMatchCache:
    def __init__(self, backend=None):
        # backend: a dict (in-memory) or Redis client
        self.store = backend or {}

    def _key(self, messages, model, max_tokens):
        content = json.dumps({'messages': messages, 'model': model,
                               'max_tokens': max_tokens}, sort_keys=True)
        return 'llm:' + hashlib.sha256(content.encode()).hexdigest()

    def get(self, messages, model, max_tokens):
        key = self._key(messages, model, max_tokens)
        return self.store.get(key)

    def set(self, messages, model, max_tokens, result, ttl_seconds=3600):
        key = self._key(messages, model, max_tokens)
        self.store[key] = result
        # In Redis: self.store.setex(key, ttl_seconds, json.dumps(result))

cache = ExactMatchCache()

# Usage
messages = [{'role': 'user', 'content': 'What is the capital of France?'}]
cached = cache.get(messages, 'gpt-4o-mini', 100)
if cached:
    print('Cache HIT:', cached[:50])
else:
    print('Cache MISS — calling API...')

Cache-Wrapped LLM Client

Wrap the LLM API call with a caching decorator so all callers get caching transparently without changing their code.

import openai
from typing import Optional

client = openai.OpenAI(api_key='YOUR_API_KEY')
cache = ExactMatchCache()

def cached_completion(messages, model='gpt-4o-mini', max_tokens=500,
                       temperature=0.0, use_cache=True) -> str:
    if use_cache and temperature == 0.0:
        # Only cache deterministic requests (temperature=0)
        cached = cache.get(messages, model, max_tokens)
        if cached:
            return cached

    response = client.chat.completions.create(
        model=model,
        messages=messages,
        max_tokens=max_tokens,
        temperature=temperature
    )
    result = response.choices[0].message.content

    if use_cache and temperature == 0.0:
        cache.set(messages, model, max_tokens, result)

    return result

# Important: only cache temperature=0 responses
# Non-deterministic responses (temp>0) may return stale results
print('Cache wrapping: only deterministic (temp=0) calls are cached.')

Semantic Caching with Embeddings

Semantic caching returns cached results for queries that are similar in meaning, not just identical strings. This uses embedding vectors and cosine similarity to find near-duplicate queries.

import numpy as np
from sklearn.metrics.pairwise import cosine_similarity

class SemanticCache:
    def __init__(self, similarity_threshold=0.95):
        self.entries = []  # [(embedding, query, result)]
        self.threshold = similarity_threshold

    def embed(self, text):
        '''Get embedding for text using OpenAI embeddings API.'''
        response = client.embeddings.create(
            model='text-embedding-3-small',
            input=text
        )
        return np.array(response.data[0].embedding)

    def get(self, query):
        if not self.entries:
            return None
        query_emb = self.embed(query)
        for emb, stored_query, result in self.entries:
            sim = cosine_similarity([query_emb], [emb])[0][0]
            if sim >= self.threshold:
                print(f'Semantic cache HIT (similarity={sim:.3f}): {stored_query[:40]}...')
                return result
        return None

    def set(self, query, result):
        emb = self.embed(query)
        self.entries.append((emb, query, result))

sem_cache = SemanticCache(similarity_threshold=0.95)
print('Semantic cache ready. Threshold: 0.95 cosine similarity.')

GPTCache Library

GPTCache is an open-source semantic caching library that supports multiple embedding models, similarity backends (FAISS, Redis), and eviction strategies. It integrates directly with OpenAI and LangChain clients.

# pip install gptcache
# GPTCache integration example

# from gptcache import cache
# from gptcache.adapter import openai
# from gptcache.embedding import Onnx
# from gptcache.manager import CacheBase, VectorBase, get_data_manager
# from gptcache.similarity_evaluation.distance import SearchDistanceEvaluation

# Initialize GPTCache
# onnx = Onnx()
# data_manager = get_data_manager(
#     CacheBase('sqlite'),
#     VectorBase('faiss', dimension=onnx.dimension)
# )
# cache.init(
#     embedding_func=onnx.to_embeddings,
#     data_manager=data_manager,
#     similarity_evaluation=SearchDistanceEvaluation(),
# )

# After init, use openai from gptcache.adapter instead of standard openai
# response = openai.ChatCompletion.create(
#     model='gpt-4o-mini',
#     messages=[{'role': 'user', 'content': 'What is Python?'}]
# )
# Same API, but cache is checked first

print('GPTCache: drop-in semantic cache for OpenAI API calls.')
print('Supports: FAISS, Redis, SQLite, Milvus as vector backends.')

Anthropic Prompt Caching (Native)

Anthropic offers native prompt caching that caches the system prompt processing on their servers. On a cache hit, you pay only 10% of the normal input token price. This is separate from application-level response caching.

import anthropic

client = anthropic.Anthropic(api_key='YOUR_API_KEY')

LONG_SYSTEM_PROMPT = '''You are an expert financial analyst with 20 years of experience.
''' + 'Domain knowledge: ' + 'analysis context...' * 500  # large system prompt

# Enable prompt caching with cache_control
response = client.messages.create(
    model='claude-opus-4-5',
    max_tokens=1024,
    system=[
        {
            'type': 'text',
            'text': LONG_SYSTEM_PROMPT,
            'cache_control': {'type': 'ephemeral'}  # cache this prefix
        }
    ],
    messages=[{'role': 'user', 'content': 'Analyze Q3 2024 earnings.'}]
)

print('Cache write tokens:', response.usage.cache_creation_input_tokens)
print('Cache read tokens: ', response.usage.cache_read_input_tokens)
print('Regular input tokens:', response.usage.input_tokens)
# On cache HIT: cache_read_input_tokens shows the cached tokens
# Cost: cached tokens charged at 10% of normal rate

Cache TTL and Eviction Strategies

Cached results become stale when the underlying knowledge changes or the model updates. TTL (Time-to-Live) and eviction strategies manage freshness.

import time
from collections import OrderedDict

class TTLCache:
    def __init__(self, max_size=1000, default_ttl=3600):
        self.store = OrderedDict()  # key: (value, expire_at)
        self.max_size = max_size
        self.default_ttl = default_ttl

    def set(self, key, value, ttl=None):
        ttl = ttl or self.default_ttl
        expire_at = time.time() + ttl
        if key in self.store:
            del self.store[key]
        self.store[key] = (value, expire_at)
        # LRU eviction: remove oldest if over capacity
        if len(self.store) > self.max_size:
            self.store.popitem(last=False)

    def get(self, key):
        if key not in self.store:
            return None
        value, expire_at = self.store[key]
        if time.time() > expire_at:
            del self.store[key]
            return None  # expired
        # Move to end (LRU update)
        self.store.move_to_end(key)
        return value

# TTL strategy guidelines
ttl_guidelines = {
    'Static knowledge': 86400,  # 24h (facts, definitions)
    'Semi-static': 3600,        # 1h (product info, FAQs)
    'Dynamic content': 300,     # 5min (news, prices)
    'Personalized': 0           # no cache (user-specific)
}
for k, v in ttl_guidelines.items():
    print(f'{k}: {v}s TTL')

Cache Invalidation Patterns

Cache invalidation — knowing when to clear stale data — is one of the hardest problems in computing. For LLM caches, these patterns handle the most common invalidation needs.

class InvalidationAwareCache(TTLCache):
    def __init__(self, *args, **kwargs):
        super().__init__(*args, **kwargs)
        self.tags = {}  # key: set of tags
        self.tag_index = {}  # tag: set of keys

    def set_with_tags(self, key, value, tags, ttl=None):
        self.set(key, value, ttl)
        self.tags[key] = set(tags)
        for tag in tags:
            self.tag_index.setdefault(tag, set()).add(key)

    def invalidate_by_tag(self, tag):
        keys_to_delete = self.tag_index.pop(tag, set())
        for key in keys_to_delete:
            self.store.pop(key, None)
            self.tags.pop(key, None)
        print(f'Invalidated {len(keys_to_delete)} entries with tag={tag}')

# Usage: tag cache entries by data source
cache = InvalidationAwareCache()
cache.set_with_tags('product_faq_123', 'Product FAQs...', tags=['product:123', 'faqs'])
cache.set_with_tags('product_spec_123', 'Spec sheet...', tags=['product:123', 'specs'])

# When product 123 is updated, invalidate all its cache entries
cache.invalidate_by_tag('product:123')  # Invalidated 2 entries

Measuring Cache Performance

Track cache performance metrics to understand the impact of caching on cost and latency. A well-tuned cache should achieve >50% hit rate for most production use cases.

class CacheMetrics:
    def __init__(self):
        self.hits = 0
        self.misses = 0
        self.total_latency_saved_ms = 0
        self.total_cost_saved_usd = 0
        self.avg_api_latency_ms = 1500  # typical LLM call latency
        self.avg_api_cost_usd = 0.002   # typical cost per call

    def record_hit(self):
        self.hits += 1
        self.total_latency_saved_ms += self.avg_api_latency_ms
        self.total_cost_saved_usd += self.avg_api_cost_usd

    def record_miss(self):
        self.misses += 1

    def report(self):
        total = self.hits + self.misses
        hit_rate = self.hits / total if total else 0
        return {
            'hit_rate': f'{hit_rate:.1%}',
            'total_requests': total,
            'cache_hits': self.hits,
            'latency_saved_sec': round(self.total_latency_saved_ms / 1000, 1),
            'cost_saved_usd': round(self.total_cost_saved_usd, 2)
        }

metrics = CacheMetrics()
for i in range(100):
    if i % 3 == 0:  # simulate 33% hit rate
        metrics.record_hit()
    else:
        metrics.record_miss()
print(metrics.report())

Redis-Backed Cache for Production

In-memory caches are lost on restart and cannot be shared across server instances. Redis provides a persistent, shared cache that works across multiple API servers in a production deployment.

import redis
import json
import hashlib

class RedisLLMCache:
    def __init__(self, host='localhost', port=6379, db=0, default_ttl=3600):
        self.client = redis.Redis(host=host, port=port, db=db,
                                   decode_responses=True)
        self.default_ttl = default_ttl

    def _key(self, messages, model):
        content = json.dumps({'messages': messages, 'model': model},
                              sort_keys=True)
        return 'llmcache:' + hashlib.sha256(content.encode()).hexdigest()

    def get(self, messages, model):
        key = self._key(messages, model)
        value = self.client.get(key)
        if value:
            self.client.expire(key, self.default_ttl)  # refresh TTL on hit
            return json.loads(value)
        return None

    def set(self, messages, model, result, ttl=None):
        key = self._key(messages, model)
        self.client.setex(key, ttl or self.default_ttl, json.dumps(result))

    def stats(self):
        keys = self.client.keys('llmcache:*')
        return {'cached_entries': len(keys),
                'memory_bytes': self.client.memory_usage('llmcache:') or 0}

# Usage: drop-in replacement for in-memory cache
# cache = RedisLLMCache(host='redis.internal', port=6379)
print('RedisLLMCache: shared across all server instances, survives restarts.')

When Not to Cache

Caching is not appropriate for all LLM calls. Understanding when to skip caching prevents serving stale or incorrect results.

DONT_CACHE_WHEN = {
    'High temperature': (
        'temperature > 0 produces different outputs for the same input. '
        'Caching would always return the first generation, defeating the purpose.'
    ),
    'Real-time data required': (
        'Queries about current prices, live news, or real-time status '
        'must always hit the API and live data source.'
    ),
    'Personalized responses': (
        'Responses that depend on user_id, session context, or personal data '
        'should not be shared across users.'
    ),
    'Safety-critical': (
        'Medical, legal, or financial responses where staleness could cause harm '
        'require fresh responses with the most current model version.'
    ),
    'Non-deterministic tools': (
        'If the prompt includes a current timestamp or random seed, '
        'the response is by design non-repeatable.'
    )
}

for condition, reason in DONT_CACHE_WHEN.items():
    print(f'Skip cache: {condition}')
    print(f'  Reason: {reason[:60]}...')
    print()

Quick Check

What is the key difference between exact-match caching and semantic caching for LLM responses?

Caching Strategies Summary

Effective prompt caching combines multiple strategies:

  • Exact-match: hash-based, zero overhead on hits, low hit rate for varied phrasing
  • Semantic caching: embedding similarity finds paraphrase matches, higher hit rate
  • GPTCache: open-source library combining both strategies with FAISS/Redis backends
  • Anthropic native caching: server-side system prompt caching at 10% token cost
  • TTL + LRU eviction: time-based freshness + capacity management
  • Tag-based invalidation: invalidate related entries when source data changes
  • When not to cache: non-zero temperature, real-time data, personalized, safety-critical

Frequently asked questions

Is the “Caching Strategies for Prompts” lesson free?

Yes — the full text of “Caching Strategies for Prompts” is free to read here on the web, and the AI Prompt Engineering course includes 4 lessons in total. To practise it interactively (a built-in code editor and a 24/7 AI tutor) and unlock the rest of the AI Prompt Engineering course, upgrade to CoddyKit PRO.

What will I learn in “Caching Strategies for Prompts”?

Semantic caching, exact-match caching, and Anthropic prompt caching. You practise AI Prompt Engineering with hands-on code you run directly in the browser, and a 24/7 AI tutor answers your questions as you work through the lesson.

Do I need any experience to start AI Prompt Engineering?

No prior experience is required. AI Prompt Engineering on CoddyKit is structured for beginners through advanced learners; this is — lesson 1 of 4, so you can start here or from the beginning and move at your own pace.

How long does the “Caching Strategies for Prompts” lesson take?

Most CoddyKit lessons take about 5–10 minutes. Each one is bite-sized and interactive, so you make steady progress and pick up exactly where you left off across the web and the app.

Can I write and run code in this AI Prompt Engineering lesson?

Yes. Every AI Prompt Engineering lesson includes a built-in code editor, so you write and run real code right in your browser and get instant AI feedback — no local setup required.

All lessons in this course

  1. Caching Strategies for Prompts
  2. Batch Processing and Async Execution
  3. Load Balancing Across Models
  4. Monitoring and Alerting for Prompt Pipelines
← Back to AI Prompt Engineering