دمج التخزين المؤقت في مسار RAG
نفّذوا طبقات التخزين المؤقت داخل تطبيق RAG لتخزين الاستجابات أو السياقات المسترجعة سابقًا واستردادها.
دمج التخزين المؤقت في مسار RAG درس مجاني في LLM Apps in Production (RAG + Vector DB + Caching) على CoddyKit. هذا هو الدرس 3 من أصل 4. يمكنك قراءة الدرس كاملاً أدناه مجاناً — ثم تمرن عليه مباشرة في المتصفح باستخدام محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7. هذا الدرس جزء من مسار التعلم في LLM Apps in Production (RAG + Vector DB + Caching)، وتقدمك يتزامن عبر الويب وتطبيق CoddyKit. تتضمن دورة LLM Apps in Production (RAG + Vector DB + Caching) 4 دروس في المجموع.
بعض أجزاء هذا الدرس لم تُترجم بعد وتظهر باللغة الإنجليزية.
Intro to RAG Caching
Welcome to the final lesson on caching! We've learned why caching is vital and explored different strategies.
Now, let's get practical. This lesson focuses on integrating caching directly into your RAG pipeline to boost performance and cut costs.
Where to Cache in RAG
In a RAG pipeline, there are two primary points where caching offers significant benefits:
- Retrieval Step: Caching the documents retrieved by your vector database.
- Generation Step: Caching the final response generated by the Large Language Model (LLM).
Each point addresses different bottlenecks.
Caching Retrieved Context
When a user asks a question, your RAG system first queries a vector database to find relevant documents (the 'context').
If the same question (or a very similar one) is asked again, why re-query the vector database? Caching the retrieved documents can save significant time and resources.
- Key: User's query (or its embedding).
- Value: List of retrieved documents/chunks.
Caching LLM Responses
After retrieving context, your RAG system sends the user's query and the context to an LLM to generate a final answer.
LLM calls are often the most expensive and slowest part. Caching the final generated response for a given query and context pair is highly effective.
- Key: Tuple of (User Query, Retrieved Context).
- Value: LLM's generated answer.
Simple In-Memory Cache
For demonstration, we'll use a basic Python dict as an in-memory cache. In real-world scenarios, you'd use dedicated caching libraries or external services like Redis.
The core idea is to:
- Check if a result for the current input exists in the cache.
- If yes (cache hit), return the cached result immediately.
- If no (cache miss), compute the result, store it in the cache, then return it.
Python: Caching LLM Calls
Here's a simple Python example demonstrating how to cache results from a simulated LLM call. Notice how the 'actual LLM call' only happens once for the same input.
import time
# Simulate an expensive LLM call
def mock_llm_call(prompt, context):
print(f"DEBUG: Making actual LLM call for: '{prompt}'")
time.sleep(1) # Simulate network delay
return f"Response to '{prompt}' with context: {context}"
# Simple in-memory cache
llm_cache = {}
def get_llm_response_cached(prompt, context):
cache_key = (prompt, context) # Use a tuple as the key
if cache_key in llm_cache:
print("DEBUG: Cache hit!")
return llm_cache[cache_key]
else:
print("DEBUG: Cache miss. Calling LLM...")
response = mock_llm_call(prompt, context)
llm_cache[cache_key] = response
return response
# Main execution
if __name__ == "__main__":
print("--- First call ---")
response1 = get_llm_response_cached(
"What is RAG?",
"RAG combines retrieval with generation."
)
print(f"Result 1: {response1}\n")
print("--- Second call (same query/context) ---")
response2 = get_llm_response_cached(
"What is RAG?",
"RAG combines retrieval with generation."
)
print(f"Result 2: {response2}\n")
print("--- Third call (different query) ---")
response3 = get_llm_response_cached(
"How does RAG work?",
"RAG uses a retriever and a generator."
)
print(f"Result 3: {response3}\n")Understanding the Cache Output
Run the code and observe the output:
- For the first call, you'll see "DEBUG: Making actual LLM call...".
- For the second call with identical inputs, you'll see "DEBUG: Cache hit!" and no actual LLM call. This saves time and cost!
- For the third call with different inputs, it's a cache miss, so another LLM call is made.
This demonstrates the core mechanism of caching LLM responses.
Integrating Cache into Retrieval
You can apply a similar caching pattern to the retrieval step. Before querying your vector database, check if the user's query (or its embedding) has been seen before.
If a cached result (the list of relevant documents) exists, skip the vector database lookup and proceed directly to the LLM call with the cached context.
This reduces load on your vector database and speeds up retrieval.
Cache Invalidation & TTL
While caching is powerful, cached data can become stale. For dynamic information, you need a strategy to clear or update the cache.
- Time-To-Live (TTL): Automatically remove entries after a set period.
- Least Recently Used (LRU): Evict the oldest entries when the cache is full.
- Event-driven: Invalidate cache entries when source data changes.
Choosing the right strategy depends on your data's freshness requirements.
Cache Integration Check
You've learned how to integrate caching at different points in a RAG pipeline. Let's test your understanding!
Recap: Caching in RAG
Great job! In this lesson, we put theory into practice.
- We identified key integration points for caching in a RAG pipeline: retrieval and generation.
- We explored how caching retrieved contexts and LLM responses can significantly improve performance and reduce operational costs.
- You saw a practical Python example of how to implement a basic in-memory cache for LLM calls.
- We briefly touched on cache invalidation strategies like TTL.
You're now equipped to start integrating caching into your own RAG applications!
الأسئلة الشائعة
هل درس «دمج التخزين المؤقت في مسار RAG» مجاني؟
نعم — نص درس «دمج التخزين المؤقت في مسار RAG» كامل متاح مجاناً هنا على الويب. لتمرينه بشكل تفاعلي (محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7) وفتح باقي دورة LLM Apps in Production (RAG + Vector DB + Caching)، انتقل إلى CoddyKit PRO. تتضمن دورة LLM Apps in Production (RAG + Vector DB + Caching) 4 دروس في المجموع.
ماذا ستتعلم في «دمج التخزين المؤقت في مسار RAG»؟
نفّذوا طبقات التخزين المؤقت داخل تطبيق RAG لتخزين الاستجابات أو السياقات المسترجعة سابقًا واستردادها. تتمرن على LLM Apps in Production (RAG + Vector DB + Caching) مع أكواد عملية تشغلها مباشرة في المتصفح، ومدرس ذكاء اصطناعي متاح 24/7 يجيب على أسئلتك أثناء عملك.
هل أحتاج إلى خبرة سابقة لأبدأ LLM Apps in Production (RAG + Vector DB + Caching)؟
لا تُشترط خبرة سابقة. LLM Apps in Production (RAG + Vector DB + Caching) على CoddyKit منظم للمبتدئين حتى المتقدمين، لذا يمكنك البدء من هنا أو من البداية والتقدم بسرعتك الخاصة. هذا هو الدرس 3 من أصل 4.
كم من الوقت يستغرق درس «دمج التخزين المؤقت في مسار RAG»؟
معظم دروس CoddyKit تستغرق حوالي 5–10 دقائق. كل منها موجز وتفاعلي، لذا تحرز تقدماً مستمراً وتستأنف من حيث توقفت عبر الويب والتطبيق.
هل يمكنني كتابة وتشغيل أكواد في درس LLM Apps in Production (RAG + Vector DB + Caching) هذا؟
نعم. كل درس في LLM Apps in Production (RAG + Vector DB + Caching) يتضمن محرر أكواد مدمج، لذا تكتب وتشغل أكواداً حقيقية مباشرة في متصفحك وتحصل على تعليقات فورية من الذكاء الاصطناعي — بدون إعداد محلي.
جميع الدروس في هذه الدورة
- أهمية التخزين المؤقت لاستدعاءات نماذج اللغة الكبيرة
- استراتيجيات التخزين المؤقت داخل الذاكرة وخارجها
- دمج التخزين المؤقت في مسار RAG
- التخزين المؤقت الدلالي لتطبيقات LLM