Strategi Caching untuk Prompt
Caching semantik, caching pencocokan persis, dan caching prompt Anthropic.
Strategi Caching untuk Prompt adalah pelajaran AI Prompt Engineering gratis di CoddyKit. Ini adalah pelajaran 1 dari 4. Kamu bisa membaca pelajaran lengkapnya di bawah secara gratis — lalu praktikkan langsung di browser dengan editor kode bawaan dan tutor AI 24/7. Ini adalah bagian dari jalur belajar AI Prompt Engineering, dan progresmu tersinkronisasi di web dan aplikasi CoddyKit. Kursus AI Prompt Engineering mencakup 4 pelajaran total.
Mengapa Hasil Prompt Perlu Di-cache?
Panggilan API LLM mahal dan lambat. Banyak aplikasi produksi mengirim prompt yang sama (atau sangat mirip) berulang kali. Peng-cache-an mengembalikan hasil tersimpan untuk kueri yang berulang, menghilangkan panggilan API yang tidak perlu dan mengurangi biaya serta latensi secara drastis.
Peng-cache-an Pencocokan Tepat dengan Kunci Hash
Cache paling sederhana: lakukan hashing pada string prompt yang tepat lalu simpan hasilnya. Jika string prompt yang sama muncul lagi, kembalikan hasil dari cache tanpa memanggil API.
import hashlib
import json
from functools import lru_cache
class ExactMatchCache:
def __init__(self, backend=None):
# backend: a dict (in-memory) or Redis client
self.store = backend or {}
def _key(self, messages, model, max_tokens):
content = json.dumps({'messages': messages, 'model': model,
'max_tokens': max_tokens}, sort_keys=True)
return 'llm:' + hashlib.sha256(content.encode()).hexdigest()
def get(self, messages, model, max_tokens):
key = self._key(messages, model, max_tokens)
return self.store.get(key)
def set(self, messages, model, max_tokens, result, ttl_seconds=3600):
key = self._key(messages, model, max_tokens)
self.store[key] = result
# In Redis: self.store.setex(key, ttl_seconds, json.dumps(result))
cache = ExactMatchCache()
# Usage
messages = [{'role': 'user', 'content': 'What is the capital of France?'}]
cached = cache.get(messages, 'gpt-4o-mini', 100)
if cached:
print('Cache HIT:', cached[:50])
else:
print('Cache MISS — calling API...')Klien LLM dengan Pembungkus Cache
Bungkus panggilan API LLM dengan dekorator cache agar semua pemanggil memperoleh fungsi cache secara transparan tanpa mengubah kode mereka.
import openai
from typing import Optional
client = openai.OpenAI(api_key='YOUR_API_KEY')
cache = ExactMatchCache()
def cached_completion(messages, model='gpt-4o-mini', max_tokens=500,
temperature=0.0, use_cache=True) -> str:
if use_cache and temperature == 0.0:
# Only cache deterministic requests (temperature=0)
cached = cache.get(messages, model, max_tokens)
if cached:
return cached
response = client.chat.completions.create(
model=model,
messages=messages,
max_tokens=max_tokens,
temperature=temperature
)
result = response.choices[0].message.content
if use_cache and temperature == 0.0:
cache.set(messages, model, max_tokens, result)
return result
# Important: only cache temperature=0 responses
# Non-deterministic responses (temp>0) may return stale results
print('Cache wrapping: only deterministic (temp=0) calls are cached.')Peng-cache-an Semantik dengan Embedding
Peng-cache-an semantik mengembalikan hasil cache untuk kueri yang serupa maknanya, bukan hanya string yang identik. Cara ini menggunakan vektor embedding dan cosine similarity untuk menemukan kueri yang hampir duplikat.
import numpy as np
from sklearn.metrics.pairwise import cosine_similarity
class SemanticCache:
def __init__(self, similarity_threshold=0.95):
self.entries = [] # [(embedding, query, result)]
self.threshold = similarity_threshold
def embed(self, text):
'''Get embedding for text using OpenAI embeddings API.'''
response = client.embeddings.create(
model='text-embedding-3-small',
input=text
)
return np.array(response.data[0].embedding)
def get(self, query):
if not self.entries:
return None
query_emb = self.embed(query)
for emb, stored_query, result in self.entries:
sim = cosine_similarity([query_emb], [emb])[0][0]
if sim >= self.threshold:
print(f'Semantic cache HIT (similarity={sim:.3f}): {stored_query[:40]}...')
return result
return None
def set(self, query, result):
emb = self.embed(query)
self.entries.append((emb, query, result))
sem_cache = SemanticCache(similarity_threshold=0.95)
print('Semantic cache ready. Threshold: 0.95 cosine similarity.')Pustaka GPTCache
GPTCache adalah pustaka peng-cache-an semantik sumber terbuka yang mendukung beberapa model embedding, backend kemiripan (FAISS, Redis), dan strategi penghapusan. Pustaka ini terintegrasi langsung dengan klien OpenAI dan LangChain.
# pip install gptcache
# GPTCache integration example
# from gptcache import cache
# from gptcache.adapter import openai
# from gptcache.embedding import Onnx
# from gptcache.manager import CacheBase, VectorBase, get_data_manager
# from gptcache.similarity_evaluation.distance import SearchDistanceEvaluation
# Initialize GPTCache
# onnx = Onnx()
# data_manager = get_data_manager(
# CacheBase('sqlite'),
# VectorBase('faiss', dimension=onnx.dimension)
# )
# cache.init(
# embedding_func=onnx.to_embeddings,
# data_manager=data_manager,
# similarity_evaluation=SearchDistanceEvaluation(),
# )
# After init, use openai from gptcache.adapter instead of standard openai
# response = openai.ChatCompletion.create(
# model='gpt-4o-mini',
# messages=[{'role': 'user', 'content': 'What is Python?'}]
# )
# Same API, but cache is checked first
print('GPTCache: drop-in semantic cache for OpenAI API calls.')
print('Supports: FAISS, Redis, SQLite, Milvus as vector backends.')Peng-cache-an Prompt Anthropic (Native)
Anthropic menyediakan peng-cache-an prompt native yang menyimpan pemrosesan prompt sistem di server mereka. Saat cache terkena, Anda hanya membayar 10% dari harga token input normal. Ini terpisah dari peng-cache-an respons pada tingkat aplikasi.
import anthropic
client = anthropic.Anthropic(api_key='YOUR_API_KEY')
LONG_SYSTEM_PROMPT = '''You are an expert financial analyst with 20 years of experience.
''' + 'Domain knowledge: ' + 'analysis context...' * 500 # large system prompt
# Enable prompt caching with cache_control
response = client.messages.create(
model='claude-opus-4-5',
max_tokens=1024,
system=[
{
'type': 'text',
'text': LONG_SYSTEM_PROMPT,
'cache_control': {'type': 'ephemeral'} # cache this prefix
}
],
messages=[{'role': 'user', 'content': 'Analyze Q3 2024 earnings.'}]
)
print('Cache write tokens:', response.usage.cache_creation_input_tokens)
print('Cache read tokens: ', response.usage.cache_read_input_tokens)
print('Regular input tokens:', response.usage.input_tokens)
# On cache HIT: cache_read_input_tokens shows the cached tokens
# Cost: cached tokens charged at 10% of normal rateTTL Cache dan Strategi Penghapusan
Hasil cache menjadi usang ketika pengetahuan yang mendasarinya berubah atau model diperbarui. TTL (Masa Berlaku) dan strategi penghapusan mengelola kesegaran data.
import time
from collections import OrderedDict
class TTLCache:
def __init__(self, max_size=1000, default_ttl=3600):
self.store = OrderedDict() # key: (value, expire_at)
self.max_size = max_size
self.default_ttl = default_ttl
def set(self, key, value, ttl=None):
ttl = ttl or self.default_ttl
expire_at = time.time() + ttl
if key in self.store:
del self.store[key]
self.store[key] = (value, expire_at)
# LRU eviction: remove oldest if over capacity
if len(self.store) > self.max_size:
self.store.popitem(last=False)
def get(self, key):
if key not in self.store:
return None
value, expire_at = self.store[key]
if time.time() > expire_at:
del self.store[key]
return None # expired
# Move to end (LRU update)
self.store.move_to_end(key)
return value
# TTL strategy guidelines
ttl_guidelines = {
'Static knowledge': 86400, # 24h (facts, definitions)
'Semi-static': 3600, # 1h (product info, FAQs)
'Dynamic content': 300, # 5min (news, prices)
'Personalized': 0 # no cache (user-specific)
}
for k, v in ttl_guidelines.items():
print(f'{k}: {v}s TTL')Pola Invalidasi Cache
Invalidasi cache—mengetahui kapan harus menghapus data yang sudah usang—merupakan salah satu masalah tersulit dalam komputasi. Untuk cache LLM, pola-pola berikut menangani kebutuhan invalidasi yang paling umum.
class InvalidationAwareCache(TTLCache):
def __init__(self, *args, **kwargs):
super().__init__(*args, **kwargs)
self.tags = {} # key: set of tags
self.tag_index = {} # tag: set of keys
def set_with_tags(self, key, value, tags, ttl=None):
self.set(key, value, ttl)
self.tags[key] = set(tags)
for tag in tags:
self.tag_index.setdefault(tag, set()).add(key)
def invalidate_by_tag(self, tag):
keys_to_delete = self.tag_index.pop(tag, set())
for key in keys_to_delete:
self.store.pop(key, None)
self.tags.pop(key, None)
print(f'Invalidated {len(keys_to_delete)} entries with tag={tag}')
# Usage: tag cache entries by data source
cache = InvalidationAwareCache()
cache.set_with_tags('product_faq_123', 'Product FAQs...', tags=['product:123', 'faqs'])
cache.set_with_tags('product_spec_123', 'Spec sheet...', tags=['product:123', 'specs'])
# When product 123 is updated, invalidate all its cache entries
cache.invalidate_by_tag('product:123') # Invalidated 2 entriesMengukur Kinerja Cache
Lacak metrik kinerja cache untuk memahami dampak peng-cache-an terhadap biaya dan latensi. Cache yang disetel dengan baik seharusnya mencapai rasio cache terkena >50% untuk sebagian besar kasus penggunaan produksi.
class CacheMetrics:
def __init__(self):
self.hits = 0
self.misses = 0
self.total_latency_saved_ms = 0
self.total_cost_saved_usd = 0
self.avg_api_latency_ms = 1500 # typical LLM call latency
self.avg_api_cost_usd = 0.002 # typical cost per call
def record_hit(self):
self.hits += 1
self.total_latency_saved_ms += self.avg_api_latency_ms
self.total_cost_saved_usd += self.avg_api_cost_usd
def record_miss(self):
self.misses += 1
def report(self):
total = self.hits + self.misses
hit_rate = self.hits / total if total else 0
return {
'hit_rate': f'{hit_rate:.1%}',
'total_requests': total,
'cache_hits': self.hits,
'latency_saved_sec': round(self.total_latency_saved_ms / 1000, 1),
'cost_saved_usd': round(self.total_cost_saved_usd, 2)
}
metrics = CacheMetrics()
for i in range(100):
if i % 3 == 0: # simulate 33% hit rate
metrics.record_hit()
else:
metrics.record_miss()
print(metrics.report())Cache Berbasis Redis untuk Produksi
Cache dalam memori akan hilang saat server dimulai ulang dan tidak dapat digunakan bersama oleh beberapa instans server. Redis menyediakan cache persisten yang dapat digunakan bersama di beberapa server API dalam penerapan produksi.
import redis
import json
import hashlib
class RedisLLMCache:
def __init__(self, host='localhost', port=6379, db=0, default_ttl=3600):
self.client = redis.Redis(host=host, port=port, db=db,
decode_responses=True)
self.default_ttl = default_ttl
def _key(self, messages, model):
content = json.dumps({'messages': messages, 'model': model},
sort_keys=True)
return 'llmcache:' + hashlib.sha256(content.encode()).hexdigest()
def get(self, messages, model):
key = self._key(messages, model)
value = self.client.get(key)
if value:
self.client.expire(key, self.default_ttl) # refresh TTL on hit
return json.loads(value)
return None
def set(self, messages, model, result, ttl=None):
key = self._key(messages, model)
self.client.setex(key, ttl or self.default_ttl, json.dumps(result))
def stats(self):
keys = self.client.keys('llmcache:*')
return {'cached_entries': len(keys),
'memory_bytes': self.client.memory_usage('llmcache:') or 0}
# Usage: drop-in replacement for in-memory cache
# cache = RedisLLMCache(host='redis.internal', port=6379)
print('RedisLLMCache: shared across all server instances, survives restarts.')Kapan Tidak Perlu Menggunakan Cache
Penggunaan cache tidak sesuai untuk semua panggilan LLM. Memahami kapan cache harus dilewati akan mencegah penyajian hasil yang sudah kedaluwarsa atau keliru.
DONT_CACHE_WHEN = {
'High temperature': (
'temperature > 0 produces different outputs for the same input. '
'Caching would always return the first generation, defeating the purpose.'
),
'Real-time data required': (
'Queries about current prices, live news, or real-time status '
'must always hit the API and live data source.'
),
'Personalized responses': (
'Responses that depend on user_id, session context, or personal data '
'should not be shared across users.'
),
'Safety-critical': (
'Medical, legal, or financial responses where staleness could cause harm '
'require fresh responses with the most current model version.'
),
'Non-deterministic tools': (
'If the prompt includes a current timestamp or random seed, '
'the response is by design non-repeatable.'
)
}
for condition, reason in DONT_CACHE_WHEN.items():
print(f'Skip cache: {condition}')
print(f' Reason: {reason[:60]}...')
print()Pemeriksaan Singkat
Apa perbedaan utama antara cache pencocokan tepat dan cache semantik untuk respons LLM?
Ringkasan Strategi Caching
Cache prompt yang efektif menggabungkan beberapa strategi:
- Pencocokan tepat: berbasis hash, tanpa beban tambahan saat terjadi kecocokan, tetapi tingkat kecocokan rendah untuk variasi frasa
- Cache semantik: kemiripan embedding menemukan kecocokan parafrasa, dengan tingkat kecocokan lebih tinggi
- GPTCache: pustaka sumber terbuka yang menggabungkan kedua strategi dengan backend FAISS/Redis
- Cache bawaan Anthropic: caching prompt sistem di sisi server dengan biaya token 10%
- Penghapusan TTL + LRU: kesegaran berbasis waktu + pengelolaan kapasitas
- Invalidasi berbasis tag: membatalkan entri terkait saat data sumber berubah
- Kapan tidak perlu menggunakan cache: temperatur tidak nol, data waktu nyata, data yang dipersonalisasi, dan penggunaan yang kritis bagi keselamatan
Pertanyaan yang Sering Diajukan
Apakah pelajaran “Strategi Caching untuk Prompt” gratis?
Ya — teks lengkap “Strategi Caching untuk Prompt” gratis dibaca di sini di web. Untuk praktiknya secara interaktif (editor kode bawaan dan tutor AI 24/7) dan buka sisa kursus AI Prompt Engineering, upgrade ke CoddyKit PRO. Kursus AI Prompt Engineering mencakup 4 pelajaran total.
Apa yang akan aku pelajari di “Strategi Caching untuk Prompt”?
Caching semantik, caching pencocokan persis, dan caching prompt Anthropic. Kamu berlatih AI Prompt Engineering dengan kode praktik yang langsung kamu jalankan di browser, dan tutor AI 24/7 menjawab pertanyaanmu saat kamu mengerjakan pelajaran ini.
Apakah aku perlu pengalaman untuk memulai AI Prompt Engineering?
Tidak diperlukan pengalaman sebelumnya. AI Prompt Engineering di CoddyKit dirancang untuk pemula hingga pelajar tingkat lanjut, jadi kamu bisa memulai di sini atau dari awal dan belajar sesuai kecepatan kamu sendiri. Ini adalah pelajaran 1 dari 4.
Berapa lama pelajaran “Strategi Caching untuk Prompt” memakan waktu?
Sebagian besar pelajaran CoddyKit memakan waktu sekitar 5–10 menit. Setiap pelajaran ringkas dan interaktif, jadi kamu membuat kemajuan stabil dan melanjutkan dari tempat kamu tinggalkan di web dan aplikasi.
Bisakah aku menulis dan menjalankan kode dalam pelajaran AI Prompt Engineering ini?
Ya. Setiap pelajaran AI Prompt Engineering menyertakan editor kode bawaan, jadi kamu menulis dan menjalankan kode nyata langsung di browser dan mendapatkan umpan balik AI instan — tidak diperlukan penyiapan lokal.
Semua pelajaran dalam kursus ini
- Strategi Caching untuk Prompt
- Pemrosesan Batch dan Eksekusi Asinkron
- Penyeimbangan Beban di Seluruh Model
- Pemantauan dan Pemberitahuan untuk Pipeline Prompt