0Pricing
PostgreSQL Performance & Query Optimization · 课时

将估算值与实际行数进行验证

在 EXPLAIN ANALYZE 中比较计划基数和实际基数,确认统计信息修复已生效。

将估算值与实际行数进行验证 是 CoddyKit 上的免费 PostgreSQL Performance & Query Optimization 课时。 这是第 4 节课,共 4 节。 你可以在下方免费阅读本课时的完整内容 — 然后在浏览器中使用内置代码编辑器和全天候 AI 导师进行实践。 这是 PostgreSQL Performance & Query Optimization 学习路径的一部分,你的进度在网页和 CoddyKit 应用中同步。 PostgreSQL Performance & Query Optimization 课程共包含 4 节课。

本课时的部分内容尚未翻译,以英文显示。

Why Validate Estimates?

When you fix bad statistics with CREATE STATISTICS or by raising default_statistics_target, you need proof that the planner now estimates cardinalities correctly.

The single best tool is EXPLAIN ANALYZE. It runs the query and reports, for every plan node, both:

  • the planner's estimated row count (rows=)
  • the actual row count observed at runtime (actual rows=)

If estimate and actual are close, your statistics fix landed. If they diverge by 10x or 100x, the planner is still flying blind.

Reading the Two Numbers

Plain EXPLAIN shows only estimates. To get actuals you must execute the query with ANALYZE.

Each node line looks like this:

  • rows=120 — the estimate
  • actual ... rows=11500 — what really happened

A ~100x gap on the orders scan is exactly the kind of misestimate that leads to a nested loop where a hash join would have been far cheaper.

EXPLAIN ANALYZE
SELECT *
FROM orders
WHERE status = 'shipped'
  AND ship_country = 'DE';

The estimate / actual Ratio

The metric to watch is the estimation ratio per node:

  • ratio = actual_rows / estimated_rows

Interpretation:

  • ~1.0 — healthy, the planner sees the data correctly
  • > 10 or < 0.1 — a real misestimate worth investigating
  • > 100 — almost always the root cause of a bad plan

Always compare at the node where the filter or join actually applies, not just the top-level row count.

Always Multiply by loops

The most common reading mistake: actual rows is reported per loop, not as a total.

If a node shows actual ... rows=5 loops=2000, the true number of rows produced is 5 × 2000 = 10000.

Compare the planner's estimate against actual_rows × loops, never against the raw per-loop figure. Forgetting this makes a healthy inner-loop node look like a wild misestimate.

EXPLAIN ANALYZE
SELECT o.*, c.name
FROM customers c
JOIN orders o ON o.customer_id = c.id
WHERE c.region = 'EU';

A Healthy Plan Looks Like This

After a good statistics fix, estimate and actual should line up on the driving nodes. Read this fragment carefully:

  • Seq Scan estimate rows=11200 vs actual rows=11500 → ratio 1.03, healthy
  • The aggregate on top also matches closely

This is the outcome you are validating for: the numbers in parentheses on the left agree with the numbers after actual on the right.

-- Reading EXPLAIN ANALYZE output:
-- Seq Scan on orders
--   (cost=0.00..2310.0 rows=11200 width=64)
--   (actual time=0.02..14.3 rows=11500 loops=1)
--   Filter: (status = 'shipped')

Correlated Columns: The Classic Trap

The planner assumes columns are independent and multiplies their selectivities. When columns are correlated, the combined estimate collapses far below reality.

Example: most rows where city = 'Berlin' also have country = 'DE'. The planner multiplies the two fractions and predicts a tiny result, but the actual count is large.

This is precisely where multivariate extended statistics repair the estimate.

EXPLAIN ANALYZE
SELECT *
FROM addresses
WHERE city = 'Berlin'
  AND country = 'DE';

Create and Refresh Extended Statistics

To teach the planner about correlation, create an extended statistics object with the dependencies kind, then analyze the table so the new statistics are populated.

Critically: CREATE STATISTICS alone does nothing until ANALYZE runs. Validation is meaningless if you skip the refresh.

CREATE STATISTICS addr_city_country (dependencies)
  ON city, country
  FROM addresses;

ANALYZE addresses;

The Before/After Discipline

Validation is a comparison, so capture two snapshots:

  • Before: run EXPLAIN ANALYZE and record the estimate vs actual on the filtered node (e.g. rows=40 vs actual rows=9000)
  • After: create + analyze the statistics, then re-run the identical query

The fix landed only if the estimate moved toward the actual (e.g. now rows=8700 vs actual rows=9000). A changed plan shape is a bonus, not the proof — the estimate convergence is the proof.

Inspect What the Planner Now Knows

You can confirm extended statistics were computed without re-running the query, by reading pg_stats_ext.

If the dependency degrees are populated (close to 1.0 for strongly correlated pairs), ANALYZE did its job and the planner has the data it needs.

SELECT statistics_name,
       attnames,
       dependencies
FROM pg_stats_ext
WHERE tablename = 'addresses';

Use BUFFERS and Format for Clarity

For serious validation, add options to the command:

  • BUFFERS — shows shared block hits/reads, exposing I/O caused by a misestimated scan
  • FORMAT JSON — gives machine-readable Plan Rows and Plan Actual Rows fields you can diff programmatically
  • SETTINGS — echoes non-default planner settings that may be skewing the test

JSON output is ideal when scripting regression checks across many queries.

EXPLAIN (ANALYZE, BUFFERS, SETTINGS, FORMAT JSON)
SELECT *
FROM addresses
WHERE city = 'Berlin'
  AND country = 'DE';

Don't Be Fooled by Rows Removed by Filter

When the estimate still looks off, check the Rows Removed by Filter line. A node can scan millions of rows yet return few, and the estimate you validate is about returned rows.

Also confirm you are validating a representative parameter value. A query that is healthy for country = 'DE' may misestimate badly for a rare value — per-value MCV skew is normal and may need a higher statistics target rather than extended statistics.

ALTER TABLE addresses
  ALTER COLUMN country SET STATISTICS 1000;

ANALYZE addresses;

Quick Check

Test your understanding of validating estimates against actual rows.

Recap

To confirm a statistics fix landed, validate estimates against actuals with discipline:

  • Run EXPLAIN ANALYZE and read rows= (estimate) against actual ... rows= per plan node.
  • Compute the ratio; treat >10x or <0.1x as a real misestimate, >100x as a likely root cause.
  • Always multiply actual rows by loops to get the true total.
  • For correlated columns, create extended statistics (dependencies/ndistinct/mcv) and then ANALYZE — creation alone does nothing.
  • Use a before/after snapshot: the proof is the estimate converging toward the actual, not merely a changed plan.
  • Reach for BUFFERS, FORMAT JSON, and pg_stats_ext to make validation rigorous and scriptable.

常见问题解答

「将估算值与实际行数进行验证」课时是免费的吗?

是的 — 「将估算值与实际行数进行验证」的完整文本可在网页上免费阅读。要进行交互式练习(内置代码编辑器和全天候 AI 导师)并解锁 PostgreSQL Performance & Query Optimization 课程的其余内容,请升级到 CoddyKit PRO。 PostgreSQL Performance & Query Optimization 课程共包含 4 节课。

「将估算值与实际行数进行验证」这节课中我会学到什么?

在 EXPLAIN ANALYZE 中比较计划基数和实际基数,确认统计信息修复已生效。 你通过在浏览器中直接运行的动手代码来练习 PostgreSQL Performance & Query Optimization,全天候 AI 导师会在你学习这节课的过程中回答你的问题。

学习 PostgreSQL Performance & Query Optimization 需要有经验吗?

无需任何先前经验。CoddyKit 上的 PostgreSQL Performance & Query Optimization 课程适合初学者到高级学习者,你可以从这里开始或从头开始,按照自己的节奏学习。 这是第 4 节课,共 4 节。

「将估算值与实际行数进行验证」课时需要多长时间?

大多数 CoddyKit 课程大约需要 5–10 分钟。每节课都很精短且互动,所以你能稳步进步,并在网页和应用中从离开的地方继续。

我能在这节 PostgreSQL Performance & Query Optimization 课中编写并运行代码吗?

能。每节 PostgreSQL Performance & Query Optimization 课都包含内置代码编辑器,你可以在浏览器中直接编写并运行真实代码,并获得即时 AI 反馈 — 无需本地设置。

此课程中的所有课时

  1. 查询规划器如何估算行数
  2. 相关列的多变量统计信息
  3. MCV 与不同值数量修正
  4. 将估算值与实际行数进行验证
← 返回 PostgreSQL Performance & Query Optimization