استراتيجيات التخزين المؤقت للمطالبات
التخزين المؤقت الدلالي، والتخزين المؤقت للتطابق التام، والتخزين المؤقت لمطالبات Anthropic
استراتيجيات التخزين المؤقت للمطالبات درس مجاني في AI Prompt Engineering على CoddyKit. هذا هو الدرس 1 من أصل 4. يمكنك قراءة الدرس كاملاً أدناه مجاناً — ثم تمرن عليه مباشرة في المتصفح باستخدام محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7. هذا الدرس جزء من مسار التعلم في AI Prompt Engineering، وتقدمك يتزامن عبر الويب وتطبيق CoddyKit. تتضمن دورة AI Prompt Engineering 4 دروس في المجموع.
لماذا نُخزّن نتائج المطالبات مؤقتًا؟
تُعد استدعاءات LLM مكلفة وبطيئة. وترسل العديد من تطبيقات الإنتاج المطالبات نفسها، أو مطالبات متشابهة جدًا، بشكل متكرر. ويعيد التخزين المؤقت النتائج المخزنة للاستعلامات المتكررة، مما يلغي استدعاءات API غير الضرورية ويقلل التكلفة ووقت الاستجابة بشكل كبير.
التخزين المؤقت للمطابقة التامة باستخدام مفاتيح التجزئة
أبسط ذاكرة تخزين مؤقت هي حساب تجزئة لسلسلة المطالبة نفسها وتخزين النتيجة. فإذا ظهرت سلسلة المطالبة نفسها مرة أخرى، تُعاد النتيجة المخزنة مؤقتًا من دون استدعاء API.
import hashlib
import json
from functools import lru_cache
class ExactMatchCache:
def __init__(self, backend=None):
# backend: a dict (in-memory) or Redis client
self.store = backend or {}
def _key(self, messages, model, max_tokens):
content = json.dumps({'messages': messages, 'model': model,
'max_tokens': max_tokens}, sort_keys=True)
return 'llm:' + hashlib.sha256(content.encode()).hexdigest()
def get(self, messages, model, max_tokens):
key = self._key(messages, model, max_tokens)
return self.store.get(key)
def set(self, messages, model, max_tokens, result, ttl_seconds=3600):
key = self._key(messages, model, max_tokens)
self.store[key] = result
# In Redis: self.store.setex(key, ttl_seconds, json.dumps(result))
cache = ExactMatchCache()
# Usage
messages = [{'role': 'user', 'content': 'What is the capital of France?'}]
cached = cache.get(messages, 'gpt-4o-mini', 100)
if cached:
print('Cache HIT:', cached[:50])
else:
print('Cache MISS — calling API...')عميل LLM مُغلّف بالتخزين المؤقت
غلّفوا استدعاء API الخاص بـ LLM بمزخرف للتخزين المؤقت، بحيث يحصل جميع المستدعين على التخزين المؤقت بشفافية من دون تغيير شيفرتهم.
import openai
from typing import Optional
client = openai.OpenAI(api_key='YOUR_API_KEY')
cache = ExactMatchCache()
def cached_completion(messages, model='gpt-4o-mini', max_tokens=500,
temperature=0.0, use_cache=True) -> str:
if use_cache and temperature == 0.0:
# Only cache deterministic requests (temperature=0)
cached = cache.get(messages, model, max_tokens)
if cached:
return cached
response = client.chat.completions.create(
model=model,
messages=messages,
max_tokens=max_tokens,
temperature=temperature
)
result = response.choices[0].message.content
if use_cache and temperature == 0.0:
cache.set(messages, model, max_tokens, result)
return result
# Important: only cache temperature=0 responses
# Non-deterministic responses (temp>0) may return stale results
print('Cache wrapping: only deterministic (temp=0) calls are cached.')التخزين المؤقت الدلالي باستخدام التضمينات
يعيد التخزين المؤقت الدلالي النتائج المخزنة مؤقتًا للاستعلامات المتشابهة في المعنى، وليس للسلاسل المتطابقة فقط. ويستخدم ذلك متجهات التضمين وتشابه جيب التمام للعثور على الاستعلامات شبه المكررة.
import numpy as np
from sklearn.metrics.pairwise import cosine_similarity
class SemanticCache:
def __init__(self, similarity_threshold=0.95):
self.entries = [] # [(embedding, query, result)]
self.threshold = similarity_threshold
def embed(self, text):
'''Get embedding for text using OpenAI embeddings API.'''
response = client.embeddings.create(
model='text-embedding-3-small',
input=text
)
return np.array(response.data[0].embedding)
def get(self, query):
if not self.entries:
return None
query_emb = self.embed(query)
for emb, stored_query, result in self.entries:
sim = cosine_similarity([query_emb], [emb])[0][0]
if sim >= self.threshold:
print(f'Semantic cache HIT (similarity={sim:.3f}): {stored_query[:40]}...')
return result
return None
def set(self, query, result):
emb = self.embed(query)
self.entries.append((emb, query, result))
sem_cache = SemanticCache(similarity_threshold=0.95)
print('Semantic cache ready. Threshold: 0.95 cosine similarity.')مكتبة GPTCache
تُعد GPTCache مكتبة مفتوحة المصدر للتخزين المؤقت الدلالي، وتدعم نماذج تضمين متعددة، وواجهات خلفية للتشابه (FAISS وRedis)، واستراتيجيات الإخلاء. كما تتكامل مباشرةً مع عملاء OpenAI وLangChain.
# pip install gptcache
# GPTCache integration example
# from gptcache import cache
# from gptcache.adapter import openai
# from gptcache.embedding import Onnx
# from gptcache.manager import CacheBase, VectorBase, get_data_manager
# from gptcache.similarity_evaluation.distance import SearchDistanceEvaluation
# Initialize GPTCache
# onnx = Onnx()
# data_manager = get_data_manager(
# CacheBase('sqlite'),
# VectorBase('faiss', dimension=onnx.dimension)
# )
# cache.init(
# embedding_func=onnx.to_embeddings,
# data_manager=data_manager,
# similarity_evaluation=SearchDistanceEvaluation(),
# )
# After init, use openai from gptcache.adapter instead of standard openai
# response = openai.ChatCompletion.create(
# model='gpt-4o-mini',
# messages=[{'role': 'user', 'content': 'What is Python?'}]
# )
# Same API, but cache is checked first
print('GPTCache: drop-in semantic cache for OpenAI API calls.')
print('Supports: FAISS, Redis, SQLite, Milvus as vector backends.')التخزين المؤقت الأصلي للمطالبات لدى Anthropic
توفر Anthropic تخزينًا مؤقتًا أصليًا للمطالبات، إذ تخزّن معالجة مطالبة النظام مؤقتًا على خوادمها. وعند العثور على النتيجة في ذاكرة التخزين المؤقت، تدفعون 10% فقط من السعر المعتاد لرموز الإدخال. وهذا منفصل عن التخزين المؤقت لنتائج الاستجابات على مستوى التطبيق.
import anthropic
client = anthropic.Anthropic(api_key='YOUR_API_KEY')
LONG_SYSTEM_PROMPT = '''You are an expert financial analyst with 20 years of experience.
''' + 'Domain knowledge: ' + 'analysis context...' * 500 # large system prompt
# Enable prompt caching with cache_control
response = client.messages.create(
model='claude-opus-4-5',
max_tokens=1024,
system=[
{
'type': 'text',
'text': LONG_SYSTEM_PROMPT,
'cache_control': {'type': 'ephemeral'} # cache this prefix
}
],
messages=[{'role': 'user', 'content': 'Analyze Q3 2024 earnings.'}]
)
print('Cache write tokens:', response.usage.cache_creation_input_tokens)
print('Cache read tokens: ', response.usage.cache_read_input_tokens)
print('Regular input tokens:', response.usage.input_tokens)
# On cache HIT: cache_read_input_tokens shows the cached tokens
# Cost: cached tokens charged at 10% of normal rateمدة صلاحية التخزين المؤقت واستراتيجيات الإخلاء
تصبح النتائج المخزنة مؤقتًا قديمة عند تغير المعرفة الأساسية أو تحديث النموذج. وتدير مدة الصلاحية TTL واستراتيجيات الإخلاء حداثة البيانات.
import time
from collections import OrderedDict
class TTLCache:
def __init__(self, max_size=1000, default_ttl=3600):
self.store = OrderedDict() # key: (value, expire_at)
self.max_size = max_size
self.default_ttl = default_ttl
def set(self, key, value, ttl=None):
ttl = ttl or self.default_ttl
expire_at = time.time() + ttl
if key in self.store:
del self.store[key]
self.store[key] = (value, expire_at)
# LRU eviction: remove oldest if over capacity
if len(self.store) > self.max_size:
self.store.popitem(last=False)
def get(self, key):
if key not in self.store:
return None
value, expire_at = self.store[key]
if time.time() > expire_at:
del self.store[key]
return None # expired
# Move to end (LRU update)
self.store.move_to_end(key)
return value
# TTL strategy guidelines
ttl_guidelines = {
'Static knowledge': 86400, # 24h (facts, definitions)
'Semi-static': 3600, # 1h (product info, FAQs)
'Dynamic content': 300, # 5min (news, prices)
'Personalized': 0 # no cache (user-specific)
}
for k, v in ttl_guidelines.items():
print(f'{k}: {v}s TTL')أنماط إبطال التخزين المؤقت
يُعد إبطال التخزين المؤقت — أي معرفة وقت مسح البيانات القديمة — من أصعب المشكلات في الحوسبة. وبالنسبة إلى ذاكرات التخزين المؤقت الخاصة بـ LLM، تعالج هذه الأنماط أكثر احتياجات الإبطال شيوعًا.
class InvalidationAwareCache(TTLCache):
def __init__(self, *args, **kwargs):
super().__init__(*args, **kwargs)
self.tags = {} # key: set of tags
self.tag_index = {} # tag: set of keys
def set_with_tags(self, key, value, tags, ttl=None):
self.set(key, value, ttl)
self.tags[key] = set(tags)
for tag in tags:
self.tag_index.setdefault(tag, set()).add(key)
def invalidate_by_tag(self, tag):
keys_to_delete = self.tag_index.pop(tag, set())
for key in keys_to_delete:
self.store.pop(key, None)
self.tags.pop(key, None)
print(f'Invalidated {len(keys_to_delete)} entries with tag={tag}')
# Usage: tag cache entries by data source
cache = InvalidationAwareCache()
cache.set_with_tags('product_faq_123', 'Product FAQs...', tags=['product:123', 'faqs'])
cache.set_with_tags('product_spec_123', 'Spec sheet...', tags=['product:123', 'specs'])
# When product 123 is updated, invalidate all its cache entries
cache.invalidate_by_tag('product:123') # Invalidated 2 entriesقياس أداء التخزين المؤقت
تتبعوا مقاييس أداء التخزين المؤقت لفهم تأثيره في التكلفة ووقت الاستجابة. وينبغي لذاكرة التخزين المؤقت المضبوطة جيدًا أن تحقق معدل إصابة >50% في معظم حالات استخدام الإنتاج.
class CacheMetrics:
def __init__(self):
self.hits = 0
self.misses = 0
self.total_latency_saved_ms = 0
self.total_cost_saved_usd = 0
self.avg_api_latency_ms = 1500 # typical LLM call latency
self.avg_api_cost_usd = 0.002 # typical cost per call
def record_hit(self):
self.hits += 1
self.total_latency_saved_ms += self.avg_api_latency_ms
self.total_cost_saved_usd += self.avg_api_cost_usd
def record_miss(self):
self.misses += 1
def report(self):
total = self.hits + self.misses
hit_rate = self.hits / total if total else 0
return {
'hit_rate': f'{hit_rate:.1%}',
'total_requests': total,
'cache_hits': self.hits,
'latency_saved_sec': round(self.total_latency_saved_ms / 1000, 1),
'cost_saved_usd': round(self.total_cost_saved_usd, 2)
}
metrics = CacheMetrics()
for i in range(100):
if i % 3 == 0: # simulate 33% hit rate
metrics.record_hit()
else:
metrics.record_miss()
print(metrics.report())التخزين المؤقت المدعوم بـ Redis للإنتاج
تُفقد ذاكرات التخزين المؤقت الموجودة في الذاكرة عند إعادة التشغيل، ولا يمكن مشاركتها بين مثيلات الخادم. يوفّر Redis ذاكرة تخزين مؤقت دائمة ومشتركة تعمل عبر خوادم API متعددة في بيئة إنتاج.
import redis
import json
import hashlib
class RedisLLMCache:
def __init__(self, host='localhost', port=6379, db=0, default_ttl=3600):
self.client = redis.Redis(host=host, port=port, db=db,
decode_responses=True)
self.default_ttl = default_ttl
def _key(self, messages, model):
content = json.dumps({'messages': messages, 'model': model},
sort_keys=True)
return 'llmcache:' + hashlib.sha256(content.encode()).hexdigest()
def get(self, messages, model):
key = self._key(messages, model)
value = self.client.get(key)
if value:
self.client.expire(key, self.default_ttl) # refresh TTL on hit
return json.loads(value)
return None
def set(self, messages, model, result, ttl=None):
key = self._key(messages, model)
self.client.setex(key, ttl or self.default_ttl, json.dumps(result))
def stats(self):
keys = self.client.keys('llmcache:*')
return {'cached_entries': len(keys),
'memory_bytes': self.client.memory_usage('llmcache:') or 0}
# Usage: drop-in replacement for in-memory cache
# cache = RedisLLMCache(host='redis.internal', port=6379)
print('RedisLLMCache: shared across all server instances, survives restarts.')متى لا ينبغي استخدام التخزين المؤقت
لا يناسب التخزين المؤقت جميع استدعاءات LLM. إن فهم الحالات التي ينبغي فيها تخطي التخزين المؤقت يمنع تقديم نتائج قديمة أو غير صحيحة.
DONT_CACHE_WHEN = {
'High temperature': (
'temperature > 0 produces different outputs for the same input. '
'Caching would always return the first generation, defeating the purpose.'
),
'Real-time data required': (
'Queries about current prices, live news, or real-time status '
'must always hit the API and live data source.'
),
'Personalized responses': (
'Responses that depend on user_id, session context, or personal data '
'should not be shared across users.'
),
'Safety-critical': (
'Medical, legal, or financial responses where staleness could cause harm '
'require fresh responses with the most current model version.'
),
'Non-deterministic tools': (
'If the prompt includes a current timestamp or random seed, '
'the response is by design non-repeatable.'
)
}
for condition, reason in DONT_CACHE_WHEN.items():
print(f'Skip cache: {condition}')
print(f' Reason: {reason[:60]}...')
print()تحقق سريع
ما الفرق الأساسي بين التخزين المؤقت بالمطابقة التامة والتخزين المؤقت الدلالي لاستجابات LLM؟
ملخص استراتيجيات التخزين المؤقت
يجمع التخزين المؤقت الفعّال للمطالبات عدة استراتيجيات:
- المطابقة التامة: تعتمد على التجزئة، ولا تتطلب أي تكلفة إضافية عند العثور على النتيجة في الذاكرة المؤقتة، لكن معدل العثور على النتائج منخفض عند تنوع الصياغات
- التخزين المؤقت الدلالي: يعثر تشابه التضمينات على الصياغات المعادَة، ما يحقق معدل العثور على النتائج أعلى
- GPTCache: مكتبة مفتوحة المصدر تجمع الاستراتيجيتين باستخدام واجهتي FAISS/Redis الخلفيتين
- التخزين المؤقت الأصلي في Anthropic: تخزين مؤقت من جهة الخادم لمطالبة النظام بتكلفة تبلغ 10% من تكلفة الرموز
- TTL + إخلاء LRU: حداثة تعتمد على الوقت + إدارة السعة
- الإبطال القائم على الوسوم: إبطال الإدخالات المرتبطة عند تغيّر بيانات المصدر
- متى لا ينبغي استخدام التخزين المؤقت: قيمة temperature غير صفرية، بيانات آنية، محتوى مخصص، حالات حرجة للسلامة
الأسئلة الشائعة
هل درس «استراتيجيات التخزين المؤقت للمطالبات» مجاني؟
نعم — نص درس «استراتيجيات التخزين المؤقت للمطالبات» كامل متاح مجاناً هنا على الويب. لتمرينه بشكل تفاعلي (محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7) وفتح باقي دورة AI Prompt Engineering، انتقل إلى CoddyKit PRO. تتضمن دورة AI Prompt Engineering 4 دروس في المجموع.
ماذا ستتعلم في «استراتيجيات التخزين المؤقت للمطالبات»؟
التخزين المؤقت الدلالي، والتخزين المؤقت للتطابق التام، والتخزين المؤقت لمطالبات Anthropic تتمرن على AI Prompt Engineering مع أكواد عملية تشغلها مباشرة في المتصفح، ومدرس ذكاء اصطناعي متاح 24/7 يجيب على أسئلتك أثناء عملك.
هل أحتاج إلى خبرة سابقة لأبدأ AI Prompt Engineering؟
لا تُشترط خبرة سابقة. AI Prompt Engineering على CoddyKit منظم للمبتدئين حتى المتقدمين، لذا يمكنك البدء من هنا أو من البداية والتقدم بسرعتك الخاصة. هذا هو الدرس 1 من أصل 4.
كم من الوقت يستغرق درس «استراتيجيات التخزين المؤقت للمطالبات»؟
معظم دروس CoddyKit تستغرق حوالي 5–10 دقائق. كل منها موجز وتفاعلي، لذا تحرز تقدماً مستمراً وتستأنف من حيث توقفت عبر الويب والتطبيق.
هل يمكنني كتابة وتشغيل أكواد في درس AI Prompt Engineering هذا؟
نعم. كل درس في AI Prompt Engineering يتضمن محرر أكواد مدمج، لذا تكتب وتشغل أكواداً حقيقية مباشرة في متصفحك وتحصل على تعليقات فورية من الذكاء الاصطناعي — بدون إعداد محلي.
جميع الدروس في هذه الدورة
- استراتيجيات التخزين المؤقت للمطالبات
- المعالجة الدفعية والتنفيذ غير المتزامن
- موازنة الحمل بين النماذج
- مراقبة مسارات المطالبات والتنبيه بشأنها