Cache semântico para respostas de LLM
Aprenda como o cache semântico reutiliza respostas para consultas semelhantes comparando significados em vez de textos exatos, reduzindo drasticamente o custo e a latência do LLM.
Cache semântico para respostas de LLM é uma aula grátis de LLM Apps in Production (RAG + Vector DB + Caching) no CoddyKit. Esta é a aula 4 de 4. Você pode ler a aula completa abaixo gratuitamente — depois pratica ao vivo no navegador com um editor de código integrado e um tutor de IA 24/7. Faz parte do caminho de aprendizado de LLM Apps in Production (RAG + Vector DB + Caching), e seu progresso é sincronizado entre a web e o app CoddyKit. O curso de LLM Apps in Production (RAG + Vector DB + Caching) inclui 4 aulas no total.
Partes desta aula ainda não foram traduzidas e aparecem em inglês.
Beyond Exact-Match Caching
A normal cache only hits when the key is byte-identical. But 'What is your refund policy?' and 'How do refunds work?' mean the same thing yet miss an exact cache.
Semantic caching matches on meaning, so paraphrases reuse the same answer.
How It Works
The flow:
- Embed the incoming query into a vector
- Search the cache for a near-by stored query
- If similarity exceeds a threshold, return the cached answer
- Otherwise call the LLM and store the new pair
Embedding the Query
Each query is converted to a vector by an embedding model. Similar meanings produce nearby vectors.
def embed(text):
return [len(text), text.count('refund'), text.count('?')]
print(embed('How do refunds work?'))Cosine Similarity
Similarity between query vectors is usually measured with cosine similarity.
import math
def cosine(a, b):
dot = sum(x*y for x, y in zip(a, b))
na = math.sqrt(sum(x*x for x in a))
nb = math.sqrt(sum(y*y for y in b))
return dot / (na * nb)
print(round(cosine([1,2,1],[1,2,0]), 3))Choosing the Threshold
The similarity threshold is the key tuning knob:
- Too low -> false hits, wrong answers served
- Too high -> few hits, little savings
Tune it on real traffic and err conservative for high-stakes domains.
A Minimal Semantic Cache
Putting embedding, similarity, and a threshold together.
cache = []
THRESH = 0.95
def get(query, qvec):
for stored_vec, ans in cache:
if cosine(qvec, stored_vec) >= THRESH:
return ans
return None
def cosine(a, b):
return 1.0 if a == b else 0.0
cache.append(([1,0], 'Refunds take 5 days'))
print(get('q', [1,0]))When NOT to Cache
Semantic caching is wrong for queries whose answer depends on changing or personal state:
- 'What is my account balance?'
- 'What is today's weather?'
- Anything user-specific or time-sensitive
Cache only stable, general knowledge.
Scoping the Cache
To avoid leaking one user's data to another, scope cache keys by tenant, language, and any relevant context. A global cache for personalized answers is a privacy bug.
Eviction and Freshness
Cached answers go stale when source data changes. Add TTLs and invalidate entries when underlying documents update, so the cache does not serve outdated answers.
Measuring Savings
Track hit rate, cost saved, and latency improvement. A 40 percent semantic hit rate can roughly translate into a 40 percent reduction in LLM spend for cacheable traffic.
Production Stack
In production, store query embeddings in a vector DB or Redis with vector search, set a tuned threshold, scope by tenant, apply TTLs, and monitor hit rate. Combine with exact caching for the best coverage.
Quick Check
Test your understanding of semantic caching.
Recap
You learned that semantic caching reuses answers for paraphrased queries by embedding them and matching via cosine similarity above a tuned threshold. Cache only stable knowledge, scope by tenant for privacy, apply TTLs for freshness, and monitor hit rate to quantify savings.
Perguntas Frequentes
A aula “Cache semântico para respostas de LLM” é grátis?
Sim — o texto completo de “Cache semântico para respostas de LLM” é grátis para ler aqui na web. Para praticá-la interativamente (um editor de código integrado e um tutor de IA 24/7) e desbloquear o restante do curso de LLM Apps in Production (RAG + Vector DB + Caching), atualize para CoddyKit PRO. O curso de LLM Apps in Production (RAG + Vector DB + Caching) inclui 4 aulas no total.
O que vou aprender em “Cache semântico para respostas de LLM”?
Aprenda como o cache semântico reutiliza respostas para consultas semelhantes comparando significados em vez de textos exatos, reduzindo drasticamente o custo e a latência do LLM. Você pratica LLM Apps in Production (RAG + Vector DB + Caching) com código prático que executa diretamente no navegador, e um tutor de IA 24/7 responde suas dúvidas enquanto trabalha na aula.
Preciso ter experiência prévia para começar LLM Apps in Production (RAG + Vector DB + Caching)?
Nenhuma experiência prévia é necessária. LLM Apps in Production (RAG + Vector DB + Caching) no CoddyKit é estruturado para alunos iniciantes até avançados, então você pode começar aqui ou desde o início e aprender no seu ritmo. Esta é a aula 4 de 4.
Quanto tempo leva a aula “Cache semântico para respostas de LLM”?
A maioria das aulas CoddyKit leva cerca de 5–10 minutos. Cada uma é compacta e interativa, então você faz progresso constante e retoma exatamente de onde parou entre web e app.
Posso escrever e executar código nesta aula de LLM Apps in Production (RAG + Vector DB + Caching)?
Sim. Cada aula de LLM Apps in Production (RAG + Vector DB + Caching) inclui um editor de código integrado, então você escreve e executa código real direto no navegador e recebe feedback de IA instantaneamente — nenhuma configuração local necessária.
Todas as aulas deste curso
- Cache Distribuído com Redis/Memcached
- Gerenciamento de Sessões e Persistência de Contexto
- Estratégias Avançadas de Invalidação de Cache
- Cache semântico para respostas de LLM