0Pricing
LLM Apps in Production (RAG + Vector DB + Caching) · 课时

将缓存集成到 RAG 流水线

在 RAG 应用中实现缓存层,以存储和获取之前生成的响应或检索到的上下文。

将缓存集成到 RAG 流水线 是 CoddyKit 上的免费 LLM Apps in Production (RAG + Vector DB + Caching) 课时。 这是第 3 节课,共 4 节。 你可以在下方免费阅读本课时的完整内容 — 然后在浏览器中使用内置代码编辑器和全天候 AI 导师进行实践。 这是 LLM Apps in Production (RAG + Vector DB + Caching) 学习路径的一部分,你的进度在网页和 CoddyKit 应用中同步。 LLM Apps in Production (RAG + Vector DB + Caching) 课程共包含 4 节课。

本课时的部分内容尚未翻译,以英文显示。

Intro to RAG Caching

Welcome to the final lesson on caching! We've learned why caching is vital and explored different strategies.

Now, let's get practical. This lesson focuses on integrating caching directly into your RAG pipeline to boost performance and cut costs.

Where to Cache in RAG

In a RAG pipeline, there are two primary points where caching offers significant benefits:

  • Retrieval Step: Caching the documents retrieved by your vector database.
  • Generation Step: Caching the final response generated by the Large Language Model (LLM).

Each point addresses different bottlenecks.

Caching Retrieved Context

When a user asks a question, your RAG system first queries a vector database to find relevant documents (the 'context').

If the same question (or a very similar one) is asked again, why re-query the vector database? Caching the retrieved documents can save significant time and resources.

  • Key: User's query (or its embedding).
  • Value: List of retrieved documents/chunks.

Caching LLM Responses

After retrieving context, your RAG system sends the user's query and the context to an LLM to generate a final answer.

LLM calls are often the most expensive and slowest part. Caching the final generated response for a given query and context pair is highly effective.

  • Key: Tuple of (User Query, Retrieved Context).
  • Value: LLM's generated answer.

Simple In-Memory Cache

For demonstration, we'll use a basic Python dict as an in-memory cache. In real-world scenarios, you'd use dedicated caching libraries or external services like Redis.

The core idea is to:

  1. Check if a result for the current input exists in the cache.
  2. If yes (cache hit), return the cached result immediately.
  3. If no (cache miss), compute the result, store it in the cache, then return it.

Python: Caching LLM Calls

Here's a simple Python example demonstrating how to cache results from a simulated LLM call. Notice how the 'actual LLM call' only happens once for the same input.

import time

# Simulate an expensive LLM call
def mock_llm_call(prompt, context):
    print(f"DEBUG: Making actual LLM call for: '{prompt}'")
    time.sleep(1) # Simulate network delay
    return f"Response to '{prompt}' with context: {context}"

# Simple in-memory cache
llm_cache = {}

def get_llm_response_cached(prompt, context):
    cache_key = (prompt, context) # Use a tuple as the key
    if cache_key in llm_cache:
        print("DEBUG: Cache hit!")
        return llm_cache[cache_key]
    else:
        print("DEBUG: Cache miss. Calling LLM...")
        response = mock_llm_call(prompt, context)
        llm_cache[cache_key] = response
        return response

# Main execution
if __name__ == "__main__":
    print("--- First call ---")
    response1 = get_llm_response_cached(
        "What is RAG?", 
        "RAG combines retrieval with generation."
    )
    print(f"Result 1: {response1}\n")

    print("--- Second call (same query/context) ---")
    response2 = get_llm_response_cached(
        "What is RAG?", 
        "RAG combines retrieval with generation."
    )
    print(f"Result 2: {response2}\n")

    print("--- Third call (different query) ---")
    response3 = get_llm_response_cached(
        "How does RAG work?", 
        "RAG uses a retriever and a generator."
    )
    print(f"Result 3: {response3}\n")

Understanding the Cache Output

Run the code and observe the output:

  • For the first call, you'll see "DEBUG: Making actual LLM call...".
  • For the second call with identical inputs, you'll see "DEBUG: Cache hit!" and no actual LLM call. This saves time and cost!
  • For the third call with different inputs, it's a cache miss, so another LLM call is made.

This demonstrates the core mechanism of caching LLM responses.

Integrating Cache into Retrieval

You can apply a similar caching pattern to the retrieval step. Before querying your vector database, check if the user's query (or its embedding) has been seen before.

If a cached result (the list of relevant documents) exists, skip the vector database lookup and proceed directly to the LLM call with the cached context.

This reduces load on your vector database and speeds up retrieval.

Cache Invalidation & TTL

While caching is powerful, cached data can become stale. For dynamic information, you need a strategy to clear or update the cache.

  • Time-To-Live (TTL): Automatically remove entries after a set period.
  • Least Recently Used (LRU): Evict the oldest entries when the cache is full.
  • Event-driven: Invalidate cache entries when source data changes.

Choosing the right strategy depends on your data's freshness requirements.

Cache Integration Check

You've learned how to integrate caching at different points in a RAG pipeline. Let's test your understanding!

Recap: Caching in RAG

Great job! In this lesson, we put theory into practice.

  • We identified key integration points for caching in a RAG pipeline: retrieval and generation.
  • We explored how caching retrieved contexts and LLM responses can significantly improve performance and reduce operational costs.
  • You saw a practical Python example of how to implement a basic in-memory cache for LLM calls.
  • We briefly touched on cache invalidation strategies like TTL.

You're now equipped to start integrating caching into your own RAG applications!

常见问题解答

「将缓存集成到 RAG 流水线」课时是免费的吗?

是的 — 「将缓存集成到 RAG 流水线」的完整文本可在网页上免费阅读。要进行交互式练习(内置代码编辑器和全天候 AI 导师)并解锁 LLM Apps in Production (RAG + Vector DB + Caching) 课程的其余内容,请升级到 CoddyKit PRO。 LLM Apps in Production (RAG + Vector DB + Caching) 课程共包含 4 节课。

「将缓存集成到 RAG 流水线」这节课中我会学到什么?

在 RAG 应用中实现缓存层,以存储和获取之前生成的响应或检索到的上下文。 你通过在浏览器中直接运行的动手代码来练习 LLM Apps in Production (RAG + Vector DB + Caching),全天候 AI 导师会在你学习这节课的过程中回答你的问题。

学习 LLM Apps in Production (RAG + Vector DB + Caching) 需要有经验吗?

无需任何先前经验。CoddyKit 上的 LLM Apps in Production (RAG + Vector DB + Caching) 课程适合初学者到高级学习者,你可以从这里开始或从头开始,按照自己的节奏学习。 这是第 3 节课,共 4 节。

「将缓存集成到 RAG 流水线」课时需要多长时间?

大多数 CoddyKit 课程大约需要 5–10 分钟。每节课都很精短且互动,所以你能稳步进步,并在网页和应用中从离开的地方继续。

我能在这节 LLM Apps in Production (RAG + Vector DB + Caching) 课中编写并运行代码吗?

能。每节 LLM Apps in Production (RAG + Vector DB + Caching) 课都包含内置代码编辑器,你可以在浏览器中直接编写并运行真实代码,并获得即时 AI 反馈 — 无需本地设置。

此课程中的所有课时

  1. 缓存 LLM 调用的重要性
  2. 内存缓存与外部缓存策略
  3. 将缓存集成到 RAG 流水线
  4. LLM 应用的语义缓存
← 返回 LLM Apps in Production (RAG + Vector DB + Caching)