汇总指标会掩盖失败
整体 97% 的结果可能掩盖一种失败的文档类型。
汇总指标会掩盖失败 是 CoddyKit 上的免费 Claude Architect 课时。 这是第 3 节课,共 4 节。 你可以在下方免费阅读本课时的完整内容 — 然后在浏览器中使用内置代码编辑器和全天候 AI 导师进行实践。 这是 Claude Architect 学习路径的一部分,你的进度在网页和 CoddyKit 应用中同步。 Claude Architect 课程共包含 4 节课。
本课时的部分内容尚未翻译,以英文显示。
The Headline Number Lies
Your extraction pipeline reports 97% accuracy. The dashboard is green, the stakeholders are happy, and someone proposes turning off human review entirely.
Stop. A single aggregate number is one of the most dangerous artifacts in a production Claude system. That 97% is an average over a mixed population. Averages are excellent at smoothing away exactly the failures that hurt you most.
In this lesson you'll learn why aggregate-only accuracy is a documented anti-pattern, and what to measure instead before you automate human oversight away.
Anatomy of a Misleading Average
Imagine your pipeline processes three document types in equal volume. The blended score is 97%. Looks uniform, right?
But blended numbers are weighted by volume, not by risk. A small, high-stakes document type can be drowned out entirely. The aggregate tells you nothing about where the 3% of errors land — and in practice, errors are almost never spread evenly.
# Same 97% aggregate, two very different realities
docs = {
"invoices": {"n": 1000, "correct": 990}, # 99.0%
"receipts": {"n": 1000, "correct": 985}, # 98.5%
"contracts": {"n": 1000, "correct": 935}, # 93.5%
}
total = sum(d["n"] for d in docs.values())
hits = sum(d["correct"] for d in docs.values())
print(f"aggregate = {hits/total:.1%}") # 97.0% — hides contracts
for name, d in docs.items():
print(name, f"{d['correct']/d['n']:.1%}")One Failing Document Type
Here is the failure mode the exam wants you to recognize: aggregate accuracy can hide poor performance on a specific document type or field.
Contracts at 93.5% might be your highest-value, highest-liability documents. A handful of wrong clauses extracted from contracts can cost more than thousands of correct receipts ever saved. Yet the dashboard happily reports 97% and invites you to automate.
The number isn't wrong. It's just answering the wrong question. "How good are we on average?" is rarely the question that matters. "Where are we weakest, and how much does that weakness cost?" is.
Field-Level Failures Hide Even Deeper
It gets subtler. Even within a single document type, the failure may live in one field. A contract extractor can nail the parties, dates, and addresses — and quietly mangle the termination_clause or liability_cap 20% of the time.
Average those fields together and you still see a comfortable score. So stratify by both axes: by document type AND by field. The cell where a critical document type meets a critical field is where your real risk concentrates.
# Stratify a labeled validation set by (doc_type, field)
import collections
stats = collections.defaultdict(lambda: [0, 0]) # [correct, total]
for row in labeled_validation_set:
key = (row["doc_type"], row["field"])
stats[key][1] += 1
stats[key][0] += int(row["pred"] == row["gold"])
for (doc, field), (ok, n) in sorted(stats.items()):
acc = ok / n
flag = " <-- REVIEW" if acc < 0.95 else ""
print(f"{doc:10} {field:18} {acc:.1%} (n={n}){flag}")Stratified Random Sampling
How do you surface these hidden cells? Not by sampling 100 random documents — random sampling reproduces your volume distribution, so rare-but-critical types get almost no coverage.
Use stratified random sampling: partition the population into strata (document type, source, field criticality), then sample randomly within each stratum. Now your low-volume contract type gets a statistically meaningful sample instead of three lucky draws.
This is the exam's recommended technique for catching what aggregates hide.
from collections import defaultdict
import random
def stratified_sample(docs, key_fn, per_stratum=50):
strata = defaultdict(list)
for d in docs:
strata[key_fn(d)].append(d)
sample = []
for stratum, items in strata.items():
k = min(per_stratum, len(items))
sample += random.sample(items, k) # random WITHIN stratum
return sample
audit_set = stratified_sample(all_docs, key_fn=lambda d: d["doc_type"])Field-Level Confidence, Calibrated
Stratified sampling tells you where you stand offline. To decide per document, in production whether to auto-accept or route to a human, you need field-level confidence calibrated on a labeled validation set.
"Calibrated" is the load-bearing word. A raw confidence-looking score means nothing until you've checked, against ground truth, that documents marked 0.9 are actually correct ~90% of the time. Calibrate first; only then can a threshold like "auto-accept above 0.97" mean what you think it means.
# Calibrate, then gate per field
def route_extraction(field_name, value, confidence, thresholds):
# thresholds[field] derived from a LABELED validation set,
# tighter for high-stakes fields
if confidence >= thresholds[field_name]:
return "auto_accept"
return "human_review"
thresholds = {
"vendor_name": 0.92,
"liability_cap": 0.99, # critical -> stricter gate
"termination_clause": 0.99,
}Self-Correction Surfaces Discrepancies
Calibrated confidence isn't the only signal. For numeric extraction you can have the model expose its own work so you can catch errors deterministically.
Extract both calculated_total (summed from line items) and stated_total (the printed total). When they disagree, you've found a discrepancy no aggregate metric would ever reveal — a verifiable, document-level red flag that routes straight to review.
# Self-correction: extract both, compare deterministically
schema = {
"type": "object",
"properties": {
"line_items": {"type": "array", "items": {"type": "number"}},
"calculated_total": {"type": "number"}, # model sums line items
"stated_total": {"type": "number"} # printed on the doc
},
"required": ["line_items", "calculated_total", "stated_total"]
}
def needs_review(out):
return abs(out["calculated_total"] - out["stated_total"]) > 0.01Don't Retry What Calibration Reveals
When a low-confidence or discrepant document surfaces, be precise about the fix. Retry-with-feedback repairs format, structural, and arithmetic errors — send the original document, the wrong output, and the exact validation error back to the model.
But retry does not help when the information is simply absent from the source. If the contract never states a liability cap, no amount of re-prompting will conjure one — and a model that fabricates one is exactly the failure your metrics must catch. Absent data is an escalation case, not a retry case.
def handle_low_confidence(doc, output, error):
if error.kind in ("format", "arithmetic", "schema"):
# retry with original doc + wrong output + exact error
return retry_with_feedback(doc, output, error)
if error.kind == "absent":
# info not in source -> never retry; route to human
return escalate_to_human(doc, reason="field absent in source")Schemas Must Not Force Fabrication
This connects to a structured-output rule that directly affects your metrics. Mark a schema field required only if it is always present. Never require a field that may be absent — the model will fabricate a value to satisfy the schema, and a fabricated value still counts as a confident answer.
That fabrication can sail straight past an aggregate accuracy check while corrupting your highest-stakes field. Make optional fields optional, and use an enum with an "other" value plus a free-text detail field for extensibility.
{
"type": "object",
"properties": {
"vendor_name": {"type": "string"},
"liability_cap": {"type": ["number", "null"]},
"doc_category": {
"type": "string",
"enum": ["invoice", "receipt", "contract", "other"]
},
"category_detail": {"type": "string"}
},
"required": ["vendor_name", "doc_category"]
}Provenance Makes Failures Auditable
To investigate a hidden failure you must be able to trace any extracted claim back to its origin. Maintain claim-to-source mappings: source document name, the exact quote, location, and publication date.
When a stratified audit flags the contract stratum, provenance lets a reviewer jump straight to the quote that produced a bad liability_cap — instead of re-reading the whole document. Provenance turns "our metric looks off somewhere" into "this field, from this quote, on this page, is wrong."
extraction = {
"field": "liability_cap",
"value": 500000,
"provenance": {
"source_doc": "acme_msa_2026.pdf",
"quote": "liability shall not exceed five hundred thousand dollars",
"page": 7,
"published": "2026-01-15"
}
}Decide on Evidence, Not Vibes
Put it together into an automation decision. You earn the right to reduce human oversight on a stratum only when the evidence for that specific stratum supports it.
- Stratified audit shows the stratum meets target accuracy.
- Field-level confidence is calibrated on labeled data.
- Self-correction and provenance catch the residual errors.
And note what is not on this list: a model self-rated confidence (1-10) or sentiment score is a bad escalation trigger. Good triggers are threshold violations, policy gaps, no progress, and explicit human requests — not the model's own untrained self-assessment.
Quick Check: Reading the 97%
A scenario-style question on metrics and oversight.
Recap: Make Hidden Failures Visible
Key takeaways:
- Aggregate accuracy hides per-type and per-field failures — it's volume-weighted, not risk-weighted.
- Stratify before you trust: stratified random sampling by document type and field surfaces weak cells that random sampling misses.
- Calibrate field-level confidence on a labeled validation set before using any threshold to auto-accept.
- Self-correction (extract calculated_total and stated_total) and provenance (claim-to-source quote, page, date) catch and explain residual errors.
- Don't require possibly-absent fields (fabrication) and don't escalate on self-rated confidence or sentiment — use threshold violations, policy gaps, no progress, and explicit human requests.
Earn automation per stratum, on evidence. The green dashboard is the start of the investigation, not the end.
常见问题解答
「汇总指标会掩盖失败」课时是免费的吗?
是的 — 「汇总指标会掩盖失败」的完整文本可在网页上免费阅读。要进行交互式练习(内置代码编辑器和全天候 AI 导师)并解锁 Claude Architect 课程的其余内容,请升级到 CoddyKit PRO。 Claude Architect 课程共包含 4 节课。
「汇总指标会掩盖失败」这节课中我会学到什么?
整体 97% 的结果可能掩盖一种失败的文档类型。 你通过在浏览器中直接运行的动手代码来练习 Claude Architect,全天候 AI 导师会在你学习这节课的过程中回答你的问题。
学习 Claude Architect 需要有经验吗?
无需任何先前经验。CoddyKit 上的 Claude Architect 课程适合初学者到高级学习者,你可以从这里开始或从头开始,按照自己的节奏学习。 这是第 3 节课,共 4 节。
「汇总指标会掩盖失败」课时需要多长时间?
大多数 CoddyKit 课程大约需要 5–10 分钟。每节课都很精短且互动,所以你能稳步进步,并在网页和应用中从离开的地方继续。
我能在这节 Claude Architect 课中编写并运行代码吗?
能。每节 Claude Architect 课都包含内置代码编辑器,你可以在浏览器中直接编写并运行真实代码,并获得即时 AI 反馈 — 无需本地设置。
此课程中的所有课时
- 声明与来源的对应关系
- 冲突的数据与日期
- 汇总指标会掩盖失败
- 分层抽样与校准