การแคชเชิงความหมายสำหรับคำตอบจาก LLM
เรียนรู้ว่าการแคชเชิงความหมายใช้คำตอบซ้ำสำหรับคำถามที่คล้ายกันอย่างไร โดยจับคู่จากความหมายแทนข้อความที่ตรงกันทุกประการ ช่วยลดค่าใช้จ่ายและเวลาแฝงของ LLM ได้อย่างมาก
การแคชเชิงความหมายสำหรับคำตอบจาก LLM เป็นบทเรียน LLM Apps in Production (RAG + Vector DB + Caching) ฟรีบน CoddyKit นี่คือบทเรียนที่ 4 จากทั้งหมด 4 บทเรียน คุณสามารถอ่านบทเรียนทั้งหมดด้านล่างฟรี — จากนั้นลองปฏิบัติด้วยตัวคุณเองในเบราว์เซอร์พร้อมตัวแก้ไขโค้ดในตัวและติวเตอร์ AI ตลอด 24/7 บทเรียนนี้เป็นส่วนหนึ่งของเส้นทางการเรียน LLM Apps in Production (RAG + Vector DB + Caching) และความก้าวหน้าของคุณจะซิงค์ข้ามเว็บและแอป CoddyKit คอร์ส LLM Apps in Production (RAG + Vector DB + Caching) มีบทเรียนทั้งหมด 4 บทเรียน
บางส่วนของบทเรียนนี้ยังไม่ได้รับการแปล และแสดงเป็นภาษาอังกฤษ
Beyond Exact-Match Caching
A normal cache only hits when the key is byte-identical. But 'What is your refund policy?' and 'How do refunds work?' mean the same thing yet miss an exact cache.
Semantic caching matches on meaning, so paraphrases reuse the same answer.
How It Works
The flow:
- Embed the incoming query into a vector
- Search the cache for a near-by stored query
- If similarity exceeds a threshold, return the cached answer
- Otherwise call the LLM and store the new pair
Embedding the Query
Each query is converted to a vector by an embedding model. Similar meanings produce nearby vectors.
def embed(text):
return [len(text), text.count('refund'), text.count('?')]
print(embed('How do refunds work?'))Cosine Similarity
Similarity between query vectors is usually measured with cosine similarity.
import math
def cosine(a, b):
dot = sum(x*y for x, y in zip(a, b))
na = math.sqrt(sum(x*x for x in a))
nb = math.sqrt(sum(y*y for y in b))
return dot / (na * nb)
print(round(cosine([1,2,1],[1,2,0]), 3))Choosing the Threshold
The similarity threshold is the key tuning knob:
- Too low -> false hits, wrong answers served
- Too high -> few hits, little savings
Tune it on real traffic and err conservative for high-stakes domains.
A Minimal Semantic Cache
Putting embedding, similarity, and a threshold together.
cache = []
THRESH = 0.95
def get(query, qvec):
for stored_vec, ans in cache:
if cosine(qvec, stored_vec) >= THRESH:
return ans
return None
def cosine(a, b):
return 1.0 if a == b else 0.0
cache.append(([1,0], 'Refunds take 5 days'))
print(get('q', [1,0]))When NOT to Cache
Semantic caching is wrong for queries whose answer depends on changing or personal state:
- 'What is my account balance?'
- 'What is today's weather?'
- Anything user-specific or time-sensitive
Cache only stable, general knowledge.
Scoping the Cache
To avoid leaking one user's data to another, scope cache keys by tenant, language, and any relevant context. A global cache for personalized answers is a privacy bug.
Eviction and Freshness
Cached answers go stale when source data changes. Add TTLs and invalidate entries when underlying documents update, so the cache does not serve outdated answers.
Measuring Savings
Track hit rate, cost saved, and latency improvement. A 40 percent semantic hit rate can roughly translate into a 40 percent reduction in LLM spend for cacheable traffic.
Production Stack
In production, store query embeddings in a vector DB or Redis with vector search, set a tuned threshold, scope by tenant, apply TTLs, and monitor hit rate. Combine with exact caching for the best coverage.
Quick Check
Test your understanding of semantic caching.
Recap
You learned that semantic caching reuses answers for paraphrased queries by embedding them and matching via cosine similarity above a tuned threshold. Cache only stable knowledge, scope by tenant for privacy, apply TTLs for freshness, and monitor hit rate to quantify savings.
เรียนรู้ LLM Apps in Production (RAG + Vector DB + Caching) ด้วย AI tutor — ฟรี
เขียนและเรียกใช้โค้ดจริงในเบราว์เซอร์ของคุณ รับความช่วยเหลือทันทีจาก AI tutor 24/7 และเรียนรู้ต่อจากที่คุณหยุดบนเว็บหรือในแอป
- คอร์ส
- 12
- บทเรียน
- 48
คำถามที่พบบ่อย
บทเรียน “การแคชเชิงความหมายสำหรับคำตอบจาก LLM” ฟรีหรือไม่
ใช่ — ข้อความเต็มของ “การแคชเชิงความหมายสำหรับคำตอบจาก LLM” ฟรีให้อ่านที่นี่บนเว็บ เพื่อปฏิบัติแบบโต้ตอบ (ตัวแก้ไขโค้ดในตัวและติวเตอร์ AI ตลอด 24/7) และปลดล็อคส่วนที่เหลือของคอร์ส LLM Apps in Production (RAG + Vector DB + Caching) ให้อัปเกรดเป็น CoddyKit PRO คอร์ส LLM Apps in Production (RAG + Vector DB + Caching) มีบทเรียนทั้งหมด 4 บทเรียน
คุณจะเรียนรู้อะไรในบทเรียน “การแคชเชิงความหมายสำหรับคำตอบจาก LLM”
เรียนรู้ว่าการแคชเชิงความหมายใช้คำตอบซ้ำสำหรับคำถามที่คล้ายกันอย่างไร โดยจับคู่จากความหมายแทนข้อความที่ตรงกันทุกประการ ช่วยลดค่าใช้จ่ายและเวลาแฝงของ LLM ได้อย่างมาก คุณปฏิบัติ LLM Apps in Production (RAG + Vector DB + Caching) ด้วยโค้ดที่ใช้งานได้จริงที่คุณเรียกใช้โดยตรงในเบราว์เซอร์ และติวเตอร์ AI ตลอด 24/7 ตอบคำถามของคุณขณะที่คุณไปผ่านบทเรียน
คุณต้องมีประสบการณ์ก่อนที่จะเริ่มเรียน LLM Apps in Production (RAG + Vector DB + Caching) หรือไม่
ไม่จำเป็นต้องมีประสบการณ์มาก่อน LLM Apps in Production (RAG + Vector DB + Caching) บน CoddyKit ออกแบบมาสำหรับผู้เริ่มต้นไปจนถึงผู้เรียนขั้นสูง คุณสามารถเริ่มต้นที่นี่หรือเริ่มจากตัวแรกและเรียนด้วยความเร็วของคุณเอง นี่คือบทเรียนที่ 4 จากทั้งหมด 4 บทเรียน
บทเรียน “การแคชเชิงความหมายสำหรับคำตอบจาก LLM” ใช้เวลานานแค่ไหน
บทเรียน CoddyKit ส่วนใหญ่ใช้เวลาประมาณ 5–10 นาที แต่ละบทเรียนจึงสั้นและเป็นแบบโต้ตอบ คุณสามารถก้าวหน้าอย่างต่อเนื่องและกลับมาเรียนต่อจากตรงที่เพิ่งหยุดบนเว็บและแอปได้เลย
ฉันเขียนและรันโค้ดในบทเรียน LLM Apps in Production (RAG + Vector DB + Caching) นี้ได้ไหม
ได้ บทเรียน LLM Apps in Production (RAG + Vector DB + Caching) ทุกบทมีตัวแก้ไขโค้ดในตัว คุณจึงเขียนและรันโค้ดจริงได้เลยในเบราว์เซอร์ และได้รับข้อเสนอแนะจาก AI ในทันที — ไม่ต้องติดตั้งในเครื่องของคุณ
บทเรียนทั้งหมดในหลักสูตรนี้
- การแคชแบบกระจายด้วย Redis/Memcached
- การจัดการเซสชันและการคงอยู่ของบริบท
- กลยุทธ์ขั้นสูงสำหรับการทำให้แคชเป็นปัจจุบัน
- การแคชเชิงความหมายสำหรับคำตอบจาก LLM