0Pricing
LLM Apps in Production (RAG + Vector DB + Caching) · 课时

内存缓存与外部缓存策略

比较不同的缓存方式,包括简单的内存缓存和 Redis 等健壮的外部解决方案。

内存缓存与外部缓存策略 是 CoddyKit 上的免费 LLM Apps in Production (RAG + Vector DB + Caching) 课时。 这是第 2 节课,共 4 节。 你可以在下方免费阅读本课时的完整内容 — 然后在浏览器中使用内置代码编辑器和全天候 AI 导师进行实践。 这是 LLM Apps in Production (RAG + Vector DB + Caching) 学习路径的一部分,你的进度在网页和 CoddyKit 应用中同步。 LLM Apps in Production (RAG + Vector DB + Caching) 课程共包含 4 节课。

本课时的部分内容尚未翻译,以英文显示。

Intro to Caching Strategies

Caching is vital for making LLM applications faster and more cost-effective. But not all caches are built the same!

In this lesson, we'll dive into two primary strategies: in-memory caching and external caching. Each has unique benefits and drawbacks depending on your application's needs.

In-Memory Caching: The Basics

In-memory caching means storing data directly within your application's Random Access Memory (RAM). Think of it like a temporary notepad your app keeps handy.

  • Speed: Accessing data from RAM is incredibly fast.
  • Simplicity: Often easy to set up, using built-in language features (like dictionaries or hash maps).
  • No External Dependencies: Your app doesn't need to connect to another service.

Simple In-Memory Cache (Python)

Here's a basic Python example using a dictionary to simulate an in-memory cache for LLM responses. Notice how subsequent requests for the same prompt hit the cache.

cache = {}

def get_llm_response(prompt):
    if prompt in cache:
        print("Cache hit!")
        return cache[prompt]
    else:
        print("Cache miss, calling LLM...")
        # Simulate a slow LLM call
        response = f"LLM response for: {prompt}"
        cache[prompt] = response
        return response

if __name__ == "__main__":
    print(get_llm_response("What is RAG?"))
    print(get_llm_response("What is RAG?"))
    print(get_llm_response("Tell me a joke."))
    print(get_llm_response("Tell me a joke."))

Limitations of In-Memory Caches

While fast and simple, in-memory caches have significant drawbacks for production LLM applications:

  • Ephemeral Data: All cached data is lost if your application restarts or crashes.
  • Limited Scale: Each instance of your application has its own separate cache. If you run multiple servers, they won't share data, leading to duplicated work.
  • Memory Usage: Large caches can consume a lot of RAM, potentially impacting your application's overall performance.

Introducing External Caching

To overcome the limitations of in-memory caches, we use external caching solutions. These store cached data outside your application, typically in a dedicated server or service.

This allows multiple instances of your application to access and share the same cached data, making it ideal for scalable, distributed systems.

Redis: A Popular External Cache

Redis (Remote Dictionary Server) is a popular open-source, in-memory data store. It's widely used as a cache, database, and message broker due to its high performance and versatile data structures.

It's an excellent choice for external caching in LLM applications because it's incredibly fast and designed for network-based access.

Advantages of External Caching

External caches like Redis provide several key advantages:

  • Distributed: Multiple application instances can share a single, consistent cache.
  • Persistent: Data can be configured to be saved to disk, so it survives application or cache server restarts.
  • Scalable: The cache can be scaled independently of your application, handling massive amounts of data and requests.
  • Rich Features: Redis offers advanced features like Time-To-Live (TTL) for automatic cache expiration and various data structures.

Using Redis for LLM Caching (Python)

Here's how you might interact with Redis from Python using the redis-py library. This code assumes a Redis server is running locally on localhost:6379.

It demonstrates setting a key with an expiration (TTL) and retrieving it.

import redis
import json

# Connect to Redis. Ensure a Redis server is running!
# e.g., on Docker: docker run --name my-redis -p 6379:6379 -d redis
try:
    r = redis.Redis(host='localhost', port=6379, db=0)
    r.ping() # Check connection
    print("Connected to Redis successfully!")
except redis.exceptions.ConnectionError as e:
    print(f"Could not connect to Redis: {e}")
    print("Please ensure a Redis server is running on localhost:6379")
    r = None # Set r to None if connection fails

def get_llm_response_from_redis(prompt):
    if r is None:
        return {"error": "Redis not connected, cannot cache."}

    cache_key = f"llm_response:{prompt}"
    cached_data = r.get(cache_key)

    if cached_data:
        print("Redis Cache hit!")
        return json.loads(cached_data.decode('utf-8'))
    else:
        print("Redis Cache miss, calling LLM...")
        # Simulate LLM call and create a response structure
        response_data = {"text": f"LLM response for: {prompt}", "source": "LLM"}
        # Cache the response for 3600 seconds (1 hour)
        r.setex(cache_key, 3600, json.dumps(response_data))
        return response_data

if __name__ == "__main__":
    print(get_llm_response_from_redis("What is the capital of France?"))
    print(get_llm_response_from_redis("What is the capital of France?"))
    print(get_llm_response_from_redis("Who invented the light bulb?"))
    print(get_llm_response_from_redis("Who invented the light bulb?"))

Choosing the Right Caching Strategy

Your choice of caching strategy depends on your application's requirements:

  • Use In-Memory Caches if:
    Your application runs as a single instance, data loss on restart is acceptable, or you're caching very small, temporary datasets.
  • Use External Caches (e.g., Redis) if:
    You need distributed caching across multiple application instances, data persistence is critical, your dataset is large, or you require advanced caching features and scalability.

For most production LLM applications, external caching is the robust choice.

Caching Strategy Quiz

Test your understanding of caching strategies!

Recap: Cache Your Knowledge

We've explored the two main caching strategies for LLM applications:

  • In-memory caches are fast and simple but are limited to a single application instance and lose data on restart.
  • External caches like Redis offer persistence, distributed sharing, and independent scalability, making them ideal for robust production LLM systems.

Choosing the right strategy depends on your application's scale, data persistence needs, and operational complexity.

常见问题解答

「内存缓存与外部缓存策略」课时是免费的吗?

是的 — 「内存缓存与外部缓存策略」的完整文本可在网页上免费阅读。要进行交互式练习(内置代码编辑器和全天候 AI 导师)并解锁 LLM Apps in Production (RAG + Vector DB + Caching) 课程的其余内容,请升级到 CoddyKit PRO。 LLM Apps in Production (RAG + Vector DB + Caching) 课程共包含 4 节课。

「内存缓存与外部缓存策略」这节课中我会学到什么?

比较不同的缓存方式,包括简单的内存缓存和 Redis 等健壮的外部解决方案。 你通过在浏览器中直接运行的动手代码来练习 LLM Apps in Production (RAG + Vector DB + Caching),全天候 AI 导师会在你学习这节课的过程中回答你的问题。

学习 LLM Apps in Production (RAG + Vector DB + Caching) 需要有经验吗?

无需任何先前经验。CoddyKit 上的 LLM Apps in Production (RAG + Vector DB + Caching) 课程适合初学者到高级学习者,你可以从这里开始或从头开始,按照自己的节奏学习。 这是第 2 节课,共 4 节。

「内存缓存与外部缓存策略」课时需要多长时间?

大多数 CoddyKit 课程大约需要 5–10 分钟。每节课都很精短且互动,所以你能稳步进步,并在网页和应用中从离开的地方继续。

我能在这节 LLM Apps in Production (RAG + Vector DB + Caching) 课中编写并运行代码吗?

能。每节 LLM Apps in Production (RAG + Vector DB + Caching) 课都包含内置代码编辑器,你可以在浏览器中直接编写并运行真实代码,并获得即时 AI 反馈 — 无需本地设置。

此课程中的所有课时

  1. 缓存 LLM 调用的重要性
  2. 内存缓存与外部缓存策略
  3. 将缓存集成到 RAG 流水线
  4. LLM 应用的语义缓存
← 返回 LLM Apps in Production (RAG + Vector DB + Caching)