Test Generation & Standards
Document fixtures and standards to improve generated tests.
Test Generation & Standards is a free Claude Architect lesson on CoddyKit — lesson 4 of 4. You can read the complete lesson below for free — then practise it hands-on in the browser with a built-in code editor and a 24/7 AI tutor. It is part of the Claude Architect learning path, one of 4 lessons in the course, and your progress syncs across the web and the CoddyKit app.
Why Generated Tests Drift
Ask Claude Code to "write tests for this module" with no guidance and you get tests that run but don't match your house style: wrong framework helpers, invented fixtures, assertions that mirror the implementation instead of the contract.
The fix is not a better one-off prompt. It is persistent, shared standards the model reads on every run. This lesson shows how to document fixtures and conventions so generated tests come out idiomatic, deterministic, and reviewable, in interactive sessions and in CI alike.
Standards Live in Project-Scope CLAUDE.md
Test conventions belong in project-level config so every contributor and every CI runner sees them. Put them in ./CLAUDE.md or .claude/CLAUDE.md, which is shared via VCS.
Do NOT rely on user-level ~/.claude/CLAUDE.md: it is personal, NOT shared through version control, so new teammates and your pipeline simply won't have it. Anything your test generation depends on must live in project scope.
# ./CLAUDE.md (committed -> every dev + CI sees it)
## Testing standards
- Framework: pytest; one test file per module as tests/test_<module>.py
- Name tests test_<behavior>_<condition>_<expected>
- Assert on the public contract, never on private internals
- No network or real time in unit tests; use the provided fixturesModularize with @path Imports
A monolithic CLAUDE.md grows unreadable and burns context. Pull the detailed testing playbook into its own file and import it with @path syntax. This keeps the root file lean while still loading the standard.
The imported file is just markdown, version-controlled like everything else, so the standard is reusable and easy to review in isolation.
# ./CLAUDE.md
@./standards/testing-style.md
@./standards/fixtures.md
# Each imported file documents one slice of the standard,
# keeping the root CLAUDE.md short and scannable.Load Test Rules Only When Needed
Even better than always-on imports: put test conventions in a .claude/rules/ file with YAML frontmatter paths. The rule loads only when editing matching files, so you save context and tokens versus a monolithic CLAUDE.md that ships everything on every turn.
Scope a rule to your test directory and it activates exactly when Claude is generating or editing tests, and stays out of the way otherwise.
# .claude/rules/testing.md
---
paths:
- "tests/**"
- "**/*.test.ts"
---
# Loaded only when a matching test file is in play
- Arrange-Act-Assert, one logical assertion per test
- Reuse fixtures from conftest.py; never hand-roll a DB
- Cover the happy path, one edge case, and one failure caseDocument Fixtures as the Source of Truth
The single biggest cause of bad generated tests is invented fixtures: the model fabricates a user object or a DB stub instead of using yours. Document the real fixtures so Claude reuses them.
Spell out what each fixture provides, its shape, and when to use it. Treat this like a tool description: purpose, return value, input format, and applicability boundaries are what drive correct selection.
# ./standards/fixtures.md (imported into CLAUDE.md)
## Available pytest fixtures (use these, do NOT invent)
- `db` -> in-memory SQLite session, auto-rolled-back per test
- `client` -> FastAPI TestClient with auth middleware disabled
- `user` -> a persisted User(id=1, role="member"); returns the ORM obj
- `frozen_now`-> pins datetime.utcnow() to 2026-01-01T00:00:00Z
# Need a different state? Parametrize an existing fixture; don't create a new DB.Few-Shot Examples Beat Vague Rules
Prose alone leaves ambiguity. Add 2 to 4 targeted examples of a canonical test and the model generalizes the pattern, it does not merely copy it. Few-shot examples are the strongest lever for consistency, edge cases, and output format.
Show one complete, idiomatic test that uses your real fixtures. New tests will mirror its structure, naming, and assertion style.
# ./standards/testing-style.md (a canonical example to generalize from)
def test_transfer_rejects_when_balance_too_low(db, user):
account = make_account(db, owner=user, balance=50)
with pytest.raises(InsufficientFunds):
transfer(db, account, amount=100)
assert account.balance == 50 # state unchanged on failure
# ^ Note: AAA layout, real `db`/`user` fixtures, asserts the contract.Write Explicit Criteria, Not Vague Wishes
"Write good tests" is a vague wish. Explicit criteria produce reliable output. Compare "be thorough" with "cover the happy path, one boundary value, and one error path; never test private methods directly."
Concrete, checkable rules remove the guesswork that makes generated tests inconsistent across files and contributors.
# In CLAUDE.md or the generation prompt -- explicit and checkable:
- Each public function gets: 1 happy-path, 1 edge/boundary, 1 failure test
- A test may fail for exactly ONE reason; split otherwise
- Mock ONLY at process boundaries (network, clock, filesystem)
- Forbidden: sleeping on real time, hitting a live service, asserting log textEncapsulate Generation as a Skill
Make the workflow repeatable with a .claude/skills/ skill (project scope is shared via VCS; user scope is personal). The skill bundles your standard and can restrict tools and isolate output.
Use context: fork to isolate verbose generation output, allowed-tools to restrict what it can touch, and argument-hint to guide the caller. Now "generate tests to standard" is one reusable command instead of a re-typed paragraph.
# .claude/skills/gen-tests/SKILL.md
---
name: gen-tests
description: Generate tests for a module using project fixtures + style
context: fork
allowed-tools: [Read, Glob, Grep, Write]
argument-hint: <path/to/module.py>
---
Follow @./standards/testing-style.md and @./standards/fixtures.md.
Find siblings with Glob **/*test*, reuse existing fixtures, then Write the test file.Find Patterns Before Generating
Don't generate in a vacuum. Have Claude follow the incremental investigation pattern first: Glob for existing test files, Read a couple, Grep for how a fixture is used, then write a new test that matches what's already there.
Grounding generation in the real codebase beats reciting a style guide, because the model copies living, working conventions instead of guessing.
# Glob to discover the established test layout
claude -p "Glob tests/**/*.py, Read two existing tests, \
Grep for usages of the `client` fixture, then write tests/test_orders.py \
following the same fixtures and naming. Do not invent new fixtures."Generate Headless, Review Fresh in CI
In a pipeline, generate tests headless with -p (required: no human is present) and --output-format json so a later step can parse the result. Then review the generated tests in a separate, isolated session.
Fresh-instance review beats same-session self-review: the author retains its own reasoning and won't challenge its own tests. A clean reviewer catches tautological assertions and missing failure cases the generator overlooked.
# 1) Generate (headless, parseable)
claude -p "$(cat .ci/gen-tests-prompt.md)" --output-format json > gen.json
# 2) Review in a FRESH session, not the generation context
claude -p "Review the new tests in gen.json against ./standards/testing-style.md. \
Flag tautological asserts and any missing failure-path test." \
--output-format json > review.jsonValidate Structure, Then Retry with Feedback
When you ask for tests as structured output (tool_use + JSON Schema), validate the result. If it is malformed, use retry-with-feedback: resend the original request, the wrong output, and the exact validation error. This reliably fixes format and structural mistakes.
Two cautions from the fact sheet: retry does NOT help when the needed info is simply absent from the source, and you should mark a schema field required only if it is always present, or the model will fabricate it.
# Schema for emitted test cases -- 'edge_case' is optional, so NOT required
{
"type": "object",
"properties": {
"test_name": {"type": "string"},
"fixtures": {"type": "array", "items": {"type": "string"}},
"assertion": {"type": "string"},
"edge_case": {"type": "string"}
},
"required": ["test_name", "fixtures", "assertion"]
}
# On a validation failure: resend original + bad output + the exact error.Quick Check
Apply the lesson to a realistic standards decision.
Recap: Test Generation & Standards
Key takeaways:
- Put testing standards in project scope (
./CLAUDE.md,.claude/rules/withpaths), shared via VCS; never depend on personal~/.claude/CLAUDE.mdin CI. - Modularize with
@pathimports; rules with frontmatterpathsload only when editing matching files, saving context. - Document fixtures as the source of truth (purpose, shape, when to use) so the model reuses them instead of inventing stubs.
- Few-shot (2-4 canonical examples) plus explicit criteria beat vague instructions; the model generalizes the pattern.
- Wrap it in a
.claude/skills/skill; have Claude investigate existing tests (Glob/Read/Grep) before writing. - In CI, generate headless with
-p --output-format json, then review in a fresh, isolated session. - Validate structured output; retry-with-feedback fixes format errors but not missing information, and require a schema field only if it is always present.
Frequently asked questions
Is the “Test Generation & Standards” lesson free?
Yes — the full text of “Test Generation & Standards” is free to read here on the web, and the Claude Architect course includes 4 lessons in total. To practise it interactively (a built-in code editor and a 24/7 AI tutor) and unlock the rest of the Claude Architect course, upgrade to CoddyKit PRO.
What will I learn in “Test Generation & Standards”?
Document fixtures and standards to improve generated tests. You practise Claude Architect with hands-on code you run directly in the browser, and a 24/7 AI tutor answers your questions as you work through the lesson.
Do I need any experience to start Claude Architect?
No prior experience is required. Claude Architect on CoddyKit is structured for beginners through advanced learners; this is — lesson 4 of 4, so you can start here or from the beginning and move at your own pace.
How long does the “Test Generation & Standards” lesson take?
Most CoddyKit lessons take about 5–10 minutes. Each one is bite-sized and interactive, so you make steady progress and pick up exactly where you left off across the web and the app.
Can I write and run code in this Claude Architect lesson?
Yes. Every Claude Architect lesson includes a built-in code editor, so you write and run real code right in your browser and get instant AI feedback — no local setup required.
All lessons in this course
- Non-Interactive Mode
- Structured Output
- Session Isolation for Reviews
- Test Generation & Standards