Por que LLMs sem estado precisam de memória externa
Entenda por que cada chamada à API começa do zero, como inserir contexto ingenuamente causa explosões no limite de tokens e o espaço de estratégias de memória, das mais simples às mais complexas.
Por que LLMs sem estado precisam de memória externa é uma aula grátis de AI Engineering Academy no CoddyKit. Esta é a aula 1 de 4. Você pode ler a aula completa abaixo gratuitamente — depois pratica ao vivo no navegador com um editor de código integrado e um tutor de IA 24/7. Faz parte do caminho de aprendizado de AI Engineering Academy, e seu progresso é sincronizado entre a web e o app CoddyKit. O curso de AI Engineering Academy inclui 4 aulas no total.
Partes desta aula ainda não foram traduzidas e aparecem em inglês.
LLMs Have No Memory by Default
Every API call to an LLM is completely independent. The model receives the messages you send in that single request and processes them — nothing more. When you make the next call, the model has no recollection of the previous conversation. It is as if you are speaking to someone with no short-term memory. This statelessness is by design: it makes models easier to scale and deploy, but it shifts the memory burden onto your application.
The Amnesia Problem in Practice
Without memory, a chatbot will fail basic multi-turn tasks. If a user says 'My name is Alice' in turn 1, then asks 'What is my name?' in turn 3, the model will say it does not know. Every turn appears to be a fresh conversation. Users find this deeply frustrating. The solution is for your application to maintain conversation history and include it in every API call.
# Naive stateless approach — model forgets everything
from openai import OpenAI
client = OpenAI()
def chat_stateless(user_message: str) -> str:
response = client.chat.completions.create(
model='gpt-4o-mini',
messages=[ # Only the new message — no history!
{'role': 'user', 'content': user_message}
]
)
return response.choices[0].message.content
chat_stateless('My name is Alice.') # Model: 'Hello Alice!'
chat_stateless('What is my name?') # Model: 'I do not know your name.'Naive Fix: Stuffing All History
The simplest memory approach is to append every turn to a growing list and send the entire list with each request. This works, but it has a critical flaw: the context window has a finite size. A 128K token window sounds large, but a long customer support chat with code snippets can exhaust it in minutes. Sending all history also means paying for the same tokens over and over.
messages = [] # grows with every turn
def chat_with_full_history(user_message: str) -> str:
messages.append({'role': 'user', 'content': user_message})
response = client.chat.completions.create(
model='gpt-4o-mini',
messages=messages # all history every time
)
reply = response.choices[0].message.content
messages.append({'role': 'assistant', 'content': reply})
return reply
# Works at first, but messages list grows unboundedly
# After 100 turns: could be 50,000+ tokens per requestThe Memory Design Space
There is a spectrum of memory strategies, each making different trade-offs between context fidelity (how much history is remembered) and token cost (how many tokens are used per request). The strategies from simplest to most sophisticated are: full buffer memory, sliding window memory, summary memory, entity memory, and vector-based episodic memory. The right choice depends on your conversation length and budget.
Token Cost of Conversation History
Each API call costs tokens based on the combined length of all messages — both the input (your messages array) and the output (the model's response). If you include the full history in every request, token cost grows quadratically with conversation length: turn N sends N previous messages. A 50-turn conversation with 200 tokens per turn sends 200+400+600+...+10,000 = over 250,000 input tokens total.
import tiktoken
def estimate_conversation_cost(
turns: int,
tokens_per_turn: int,
price_per_1k_input: float = 0.00015 # gpt-4o-mini
) -> float:
total_input_tokens = sum(
(i + 1) * tokens_per_turn # each turn sends all previous turns
for i in range(turns)
)
total_output_tokens = turns * tokens_per_turn
cost = (total_input_tokens / 1000) * price_per_1k_input
print(f'{turns} turns: {total_input_tokens:,} input tokens, ${cost:.4f}')
return cost
estimate_conversation_cost(50, 200)Where to Store Conversation History
Conversation history needs to be stored outside the Python process to survive server restarts and scale across multiple instances. Common storage backends include: Redis for fast in-memory access with TTL expiry (great for active sessions), PostgreSQL for durable long-term storage and analytics, and DynamoDB for serverless auto-scaling. LangChain provides connectors for all of these out of the box.
# Three storage options for conversation history
# 1. In-process dict (development only — lost on restart)
from langchain_core.chat_history import InMemoryChatMessageHistory
# 2. Redis (production — fast, TTL-based expiry)
from langchain_community.chat_message_histories import RedisChatMessageHistory
history = RedisChatMessageHistory(session_id='user-123', url='redis://localhost:6379')
# 3. PostgreSQL (production — durable, queryable)
from langchain_community.chat_message_histories import PostgresChatMessageHistory
history = PostgresChatMessageHistory(
session_id='user-123',
connection_string='postgresql://user:pass@localhost/db'
)Session IDs: Separating User Conversations
When your app serves multiple users, you need to maintain separate histories per conversation. A session_id (typically a UUID or a combination of user ID and conversation ID) identifies which history to load for each request. LangChain's RunnableWithMessageHistory accepts a get_session_history function that takes a session ID and returns the appropriate history object.
import uuid
from langchain_core.runnables.history import RunnableWithMessageHistory
from langchain_core.chat_history import InMemoryChatMessageHistory
store = {} # session_id -> history (use Redis in production)
def get_session_history(session_id: str) -> InMemoryChatMessageHistory:
if session_id not in store:
store[session_id] = InMemoryChatMessageHistory()
return store[session_id]
# Create a new session for each user conversation
def new_session() -> str:
return str(uuid.uuid4())
alice_session = new_session()
bob_session = new_session()
# Alice and Bob's histories are completely independentThe RunnableWithMessageHistory Wrapper
RunnableWithMessageHistory is LangChain's LCEL-native way to add memory to any chain. It wraps your chain, automatically loads history before each invocation, appends the new user message and AI response, and saves everything back to storage. You specify which input key contains the user message and which prompt variable should receive the history.
from langchain_openai import ChatOpenAI
from langchain_core.prompts import ChatPromptTemplate, MessagesPlaceholder
from langchain_core.runnables.history import RunnableWithMessageHistory
from langchain_core.output_parsers import StrOutputParser
prompt = ChatPromptTemplate.from_messages([
('system', 'You are a helpful assistant.'),
MessagesPlaceholder(variable_name='history'), # history injected here
('human', '{input}'),
])
chain = prompt | ChatOpenAI(model='gpt-4o-mini') | StrOutputParser()
chain_with_memory = RunnableWithMessageHistory(
chain,
get_session_history,
input_messages_key='input',
history_messages_key='history'
)Invoking a Memory-Enabled Chain
When invoking a chain wrapped with RunnableWithMessageHistory, you pass a config dict with configurable: {session_id: ...}. This tells the wrapper which history store to load. The chain handles everything else: loading history before the call, injecting it into the prompt, and saving the new turn after the response arrives.
session_id = 'user-alice-session-1'
# First message
response1 = chain_with_memory.invoke(
{'input': 'My name is Alice and I love Python.'},
config={'configurable': {'session_id': session_id}}
)
print(response1) # 'Hello Alice! Nice to meet you.'
# Second message — model now knows Alice's name and interest
response2 = chain_with_memory.invoke(
{'input': 'What is my name and what do I love?'},
config={'configurable': {'session_id': session_id}}
)
print(response2) # 'Your name is Alice and you love Python!'Memory Failure Modes to Avoid
Three common memory implementation mistakes: Forgetting session isolation — reusing one history object for all users leaks private data between conversations. Ignoring TTL — storing histories indefinitely fills your database; set expiry for inactive sessions. Over-trusting history — users can inject false memories ('I told you I am an admin') so validate claims against a source of truth rather than the conversation history alone.
# Mistake 1: Shared history for all users
global_history = InMemoryChatMessageHistory() # BAD!
# Fix: per-session history
store = {} # keyed by session_id
# Mistake 2: No TTL on Redis history
# BAD: history = RedisChatMessageHistory(session_id=sid, url=url)
# Good: set TTL to 24 hours
history = RedisChatMessageHistory(
session_id=session_id,
url='redis://localhost',
ttl=86400 # 24 hours in seconds
)Choosing Your Memory Strategy
Use full buffer memory only for short conversations where you know the total context will stay within limits. Use sliding window for general chatbots (keep last N turns). Use summary memory when conversations are open-ended and may be very long. Use vector memory when users need to recall specific facts from much earlier in a long conversation. We explore each of these in upcoming lessons.
Quick Check
Test your understanding of why stateless LLMs need external memory.
Lesson Recap
In this lesson you learned: LLMs are stateless by design — each API call sees only what you send in that request, naive full-history stuffing grows token costs quadratically and eventually hits context limits, and RunnableWithMessageHistory is LangChain's clean way to add external memory with any backend (Redis, PostgreSQL) keyed by session ID. Next up we explore specific memory strategies: buffer and window memory.
Perguntas Frequentes
A aula “Por que LLMs sem estado precisam de memória externa” é grátis?
Sim — o texto completo de “Por que LLMs sem estado precisam de memória externa” é grátis para ler aqui na web. Para praticá-la interativamente (um editor de código integrado e um tutor de IA 24/7) e desbloquear o restante do curso de AI Engineering Academy, atualize para CoddyKit PRO. O curso de AI Engineering Academy inclui 4 aulas no total.
O que vou aprender em “Por que LLMs sem estado precisam de memória externa”?
Entenda por que cada chamada à API começa do zero, como inserir contexto ingenuamente causa explosões no limite de tokens e o espaço de estratégias de memória, das mais simples às mais complexas. Você pratica AI Engineering Academy com código prático que executa diretamente no navegador, e um tutor de IA 24/7 responde suas dúvidas enquanto trabalha na aula.
Preciso ter experiência prévia para começar AI Engineering Academy?
Nenhuma experiência prévia é necessária. AI Engineering Academy no CoddyKit é estruturado para alunos iniciantes até avançados, então você pode começar aqui ou desde o início e aprender no seu ritmo. Esta é a aula 1 de 4.
Quanto tempo leva a aula “Por que LLMs sem estado precisam de memória externa”?
A maioria das aulas CoddyKit leva cerca de 5–10 minutos. Cada uma é compacta e interativa, então você faz progresso constante e retoma exatamente de onde parou entre web e app.
Posso escrever e executar código nesta aula de AI Engineering Academy?
Sim. Cada aula de AI Engineering Academy inclui um editor de código integrado, então você escreve e executa código real direto no navegador e recebe feedback de IA instantaneamente — nenhuma configuração local necessária.
Todas as aulas deste curso
- Por que LLMs sem estado precisam de memória externa
- Memória de buffer e de janela
- Memória de resumo e truncamento consciente dos tokens
- Persistindo o histórico de conversas no Redis e no PostgreSQL