Designing the Production Architecture
Choose a capstone project, sketch the full architecture including RAG pipeline, agent layer, caching, observability, and API, then document the design decisions and trade-offs.
Designing the Production Architecture is a free AI Engineering Academy lesson on CoddyKit — lesson 1 of 4. You can read the complete lesson below for free — then practise it hands-on in the browser with a built-in code editor and a 24/7 AI tutor. It is part of the AI Engineering Academy learning path, one of 4 lessons in the course, and your progress syncs across the web and the CoddyKit app.
Choosing a Capstone Project
The capstone project ties together everything from the track: RAG, agents, streaming, caching, observability, and security. A good capstone project is meaningfully complex — requiring at least three distinct AI components — yet scoped small enough to ship in days rather than months. Classic examples: a production-grade document Q&A assistant, an autonomous research agent with human-in-the-loop oversight, or an enterprise data extraction pipeline.
Identifying System Components
Start by listing the distinct components your system needs. A production AI engineering project typically includes: an ingestion layer (document loading, chunking, embedding, vector store indexing), a retrieval layer (hybrid search, re-ranking), an agent layer (function calling, tool execution), an API layer (FastAPI backend, streaming endpoints), and an observability layer (tracing, metrics, alerts). Sketch the data flow between them before writing code.
# Component inventory for a Document QA Assistant:
COMPONENTS = [
'document_ingestion', # PDF/Word -> chunks -> embeddings -> pgvector
'hybrid_retriever', # BM25 + dense + RRF
'reranker', # Cohere rerank
'qa_agent', # GPT-4o with RAG + function calling
'semantic_cache', # Redis + embedding similarity
'streaming_api', # FastAPI StreamingResponse
'tracing', # LangSmith or Langfuse
'prompt_injection_filter', # Input sanitization
'eval_pipeline', # Automated quality scoring
]Choosing the Right Stack
Select technology based on your team's familiarity and the system's actual requirements rather than novelty. A reasonable default stack: FastAPI for the API layer, pgvector for vector storage (reuses existing PostgreSQL infrastructure), LangChain LCEL for pipeline composition, Redis for semantic caching and rate-limit state, LangSmith for tracing, and PostgreSQL for user data and evaluation results. Add components only when simpler options do not meet requirements.
# Technology decisions and their rationale
STACK = {
'api': ('FastAPI', 'Async support, OpenAPI docs, streaming easy'),
'vector_store': ('pgvector', 'Already on PostgreSQL, no extra infra'),
'llm_primary': ('gpt-4o', 'Best quality for the use case'),
'llm_fallback': ('claude-3.5', 'Different provider for resilience'),
'cache': ('Redis', 'Sub-ms lookup, TTL support built-in'),
'tracing': ('LangSmith', 'Native LangChain integration'),
'embedding': ('text-embedding-3-small', 'Good quality, low cost'),
'reranker': ('Cohere', 'Best-in-class rerank API'),
}Drawing the Architecture Diagram
Document the system architecture as a data-flow diagram showing each component, the data that flows between them, and the direction of flow. Include both the ingestion path (offline: documents → chunks → embeddings → vector store) and the query path (online: user query → cache check → retrieval → reranking → LLM → streaming response). This diagram is your north star for implementation and helps new team members understand the system instantly.
# Data flow (ASCII art):
#
# INGESTION PATH (offline):
# Documents --> Loader --> Chunker --> Embedder --> pgvector
# |
# BM25 index
#
# QUERY PATH (online):
# User Query
# |-> Semantic Cache (hit: return) -> miss:
# |-> Hybrid Retriever (BM25 + dense)
# |-> Cohere Reranker
# |-> LangChain LCEL Chain
# |-> GPT-4o (streaming) --> FastAPI StreamingResponse
# |-> LangSmith (trace every step)Defining API Contracts Early
Define your API endpoints and their request/response schemas in FastAPI before implementing backend logic. This creates a contract between the API and any frontend consumers and makes parallel development possible. Document every endpoint with OpenAPI descriptions. At minimum, define endpoints for: document ingestion, chat/query, conversation history, evaluation results, and system health.
from fastapi import FastAPI
from pydantic import BaseModel
app = FastAPI(title='Document QA Assistant', version='1.0')
class QueryRequest(BaseModel):
question: str
conversation_id: str | None = None
max_chunks: int = 5
stream: bool = True
class QueryResponse(BaseModel):
answer: str
sources: list
cached: bool
latency_ms: int
trace_id: str
@app.post('/query', response_model=QueryResponse)
async def query_endpoint(request: QueryRequest):
pass # implementation next lessonEstimating and Budgeting Costs
Estimate your system's monthly API cost before writing a single line of backend code. Count: expected daily active users × queries per user × average tokens per query. For a system with 100 DAU, 10 queries/day, and 3,000 tokens per query at GPT-4o pricing, that is 3 million tokens per day — roughly $45/day or $1,350/month. This estimate tells you whether the system is economically viable and which optimizations (caching, model routing) are worth implementing.
# Cost estimation model
DAU = 100 # daily active users
QPD = 10 # queries per user per day
TOKENS_PER_QUERY = {
'prompt_tokens': 2000, # system + context + question
'completion_tokens': 500
}
PRICE_GPT4O_INPUT = 2.50 / 1_000_000 # per token
PRICE_GPT4O_OUTPUT = 10.00 / 1_000_000 # per token
daily_cost = DAU * QPD * (
TOKENS_PER_QUERY['prompt_tokens'] * PRICE_GPT4O_INPUT +
TOKENS_PER_QUERY['completion_tokens'] * PRICE_GPT4O_OUTPUT
)
print(f'Daily: ${daily_cost:.2f}, Monthly: ${daily_cost * 30:.2f}')Planning the Ingestion Pipeline
Design the ingestion pipeline as an offline batch process that runs on demand or on a schedule. Define what document types you will support (PDF, DOCX, HTML, plain text), the chunking strategy, the embedding model, and the metadata fields to store alongside each vector. Metadata is critical for filtered retrieval — without it, you cannot restrict search to documents from a specific date range, author, or category.
from dataclasses import dataclass
from typing import Optional
@dataclass
class ChunkMetadata:
doc_id: str
source_file: str
page_number: Optional[int]
section_title: Optional[str]
created_at: str
author: Optional[str]
doc_type: str # 'pdf', 'docx', 'html'
# Ingestion config
INGESTION_CONFIG = {
'chunk_size': 800,
'chunk_overlap': 100,
'embedding_model': 'text-embedding-3-small',
'embedding_dimensions': 1536,
'batch_size': 100, # chunks per embedding API call
}Security Architecture Decisions
Make security decisions upfront rather than retrofitting them later. Define: how users authenticate (JWT, API keys), what data can be retrieved per user (row-level security in pgvector queries), how prompt injection is detected, what output scanning is applied, and which actions require multi-factor confirmation. Each decision has performance implications that affect architecture choices throughout the system.
SECURITY_DECISIONS = {
'auth': 'JWT with 24h expiry',
'data_isolation': 'tenant_id column in all vector metadata, filter on every query',
'injection_detection': 'rule-based pre-filter + LLM secondary check for complex inputs',
'output_scanning': 'check for PII, system prompt leakage patterns',
'destructive_actions': 'require confirmation token for delete operations',
'rate_limiting': '20 queries/minute per user, 429 with Retry-After header',
'key_storage': 'AWS Secrets Manager, rotated every 90 days'
}Defining Success Metrics
Define what success looks like for your system before you build it. A set of concrete, measurable success criteria keeps development focused and gives you clear go/no-go criteria for deployment. Include metrics for: quality (eval score on test set), latency (p95 TTFT and total), cost (per-query cost target), and reliability (uptime SLA). Post these on the project README so all contributors share the same target.
SUCCESS_METRICS = {
# Quality
'min_correctness_score': 4.0, # out of 5, LLM-as-judge
'min_retrieval_hit_rate': 0.85, # top-5 chunk contains answer
# Latency
'p95_ttft_ms': 800, # time to first token
'p95_total_latency_ms': 8000, # full response
# Cost
'max_cost_per_query_usd': 0.05, # $0.05 per Q&A
# Reliability
'target_uptime': 0.999, # 99.9%
'cache_hit_rate_target': 0.25, # 25% queries served from cache
}Documenting Architecture Trade-offs
Every architecture decision involves trade-offs. Document them explicitly in an Architecture Decision Record (ADR): what decision was made, what alternatives were considered, and why this option was chosen. For example: 'Chose pgvector over Pinecone because we already run PostgreSQL, reducing operational overhead. Trade-off: maximum scale is limited to ~10M vectors without sharding.' Future team members will thank you for this transparency.
# Architecture Decision Records (ADRs):
#
# ADR-001: Use pgvector for vector storage
# Decision: pgvector in existing PostgreSQL
# Alternatives: Pinecone, Weaviate, Qdrant
# Reason: No new infra, row-level security native, familiar operations
# Trade-offs: Limited to ~5M vectors before performance degrades
#
# ADR-002: GPT-4o as primary model
# Decision: gpt-4o for all user-facing queries
# Alternatives: gpt-4o-mini (cheaper), Claude (alternative)
# Reason: Highest quality for use case, Claude as fallback
# Trade-offs: $0.04/query vs $0.002 for gpt-4o-miniSketching the Deployment Topology
Define how your system will be deployed before writing any infrastructure code. Map each component to a deployment unit: the FastAPI backend as a Docker container, the ingestion pipeline as a separate worker service, pgvector as a managed PostgreSQL instance, Redis as a managed cache. Specify which components are stateless (can be scaled horizontally) versus stateful (require careful scaling strategies). A deployment diagram saves hours of rework later.
# Deployment topology:
#
# Internet -> Load Balancer (AWS ALB)
# |
# FastAPI API (stateless, 2-4 containers, auto-scale)
# |
# +------+------+
# | |
# pgvector Redis Cache
# (AWS RDS Postgres) (Elasticache)
# |
# Ingestion Worker (separate container, manual trigger)
# |
# LangSmith (external SaaS, traces only)
#
# All containers in same VPC, no public access to DB/cacheQuick Check
Test your understanding of production AI system architecture design.
Lesson Recap
In this lesson you learned: component inventory and data-flow diagrams create a shared architectural vision before coding begins, API contracts defined upfront enable parallel development and make requirements explicit, and success metrics defined in advance keep the team aligned on quality, latency, cost, and reliability targets. Next up we implement the core RAG and agent features.
Frequently asked questions
Is the “Designing the Production Architecture” lesson free?
Yes — the full text of “Designing the Production Architecture” is free to read here on the web, and the AI Engineering Academy course includes 4 lessons in total. To practise it interactively (a built-in code editor and a 24/7 AI tutor) and unlock the rest of the AI Engineering Academy course, upgrade to CoddyKit PRO.
What will I learn in “Designing the Production Architecture”?
Choose a capstone project, sketch the full architecture including RAG pipeline, agent layer, caching, observability, and API, then document the design decisions and trade-offs. You practise AI Engineering Academy with hands-on code you run directly in the browser, and a 24/7 AI tutor answers your questions as you work through the lesson.
Do I need any experience to start AI Engineering Academy?
No prior experience is required. AI Engineering Academy on CoddyKit is structured for beginners through advanced learners; this is — lesson 1 of 4, so you can start here or from the beginning and move at your own pace.
How long does the “Designing the Production Architecture” lesson take?
Most CoddyKit lessons take about 5–10 minutes. Each one is bite-sized and interactive, so you make steady progress and pick up exactly where you left off across the web and the app.
Can I write and run code in this AI Engineering Academy lesson?
Yes. Every AI Engineering Academy lesson includes a built-in code editor, so you write and run real code right in your browser and get instant AI feedback — no local setup required.
All lessons in this course
- Designing the Production Architecture
- Implementing Core RAG and Agent Features
- Hardening: Security, Caching, and Reliability
- Evaluation, Deployment, and Retrospective