0Pricing
PostgreSQL Performance & Query Optimization · 课时

MCV 与不同值数量修正

使用 ndistinct 和最常见值统计信息,修正连接和分组估算。

MCV 与不同值数量修正 是 CoddyKit 上的免费 PostgreSQL Performance & Query Optimization 课时。 这是第 3 节课,共 4 节。 你可以在下方免费阅读本课时的完整内容 — 然后在浏览器中使用内置代码编辑器和全天候 AI 导师进行实践。 这是 PostgreSQL Performance & Query Optimization 学习路径的一部分,你的进度在网页和 CoddyKit 应用中同步。 PostgreSQL Performance & Query Optimization 课程共包含 4 节课。

本课时的部分内容尚未翻译,以英文显示。

Why Estimates Drift

The PostgreSQL planner chooses join orders, join methods, and grouping strategies from row-count estimates. When those estimates are wrong, you get nested loops over millions of rows or a hash table sized for the wrong cardinality.

Two column-level statistics drive most of these estimates:

  • n_distinct — how many distinct values the planner believes a column holds. It feeds grouping and join cardinality.
  • most_common_vals (MCV) — the list of frequent values and their frequencies, used for selectivity of equality predicates.

This lesson shows how to read, diagnose, and correct both when the default sampling gets them wrong.

Reading pg_stats

Everything the planner knows about a column lives in the pg_stats view, a human-readable wrapper over pg_statistic. Start every diagnosis here.

Key columns: n_distinct, most_common_vals, most_common_freqs, and null_frac.

SELECT attname,
       n_distinct,
       null_frac,
       most_common_vals,
       most_common_freqs
FROM pg_stats
WHERE schemaname = 'public'
  AND tablename = 'orders'
  AND attname IN ('customer_id', 'status');

How n_distinct Is Encoded

The n_distinct value is overloaded with two meanings:

  • A positive number is an absolute count of distinct values (e.g. 4200).
  • A negative number between -1 and 0 is a ratio of distinct values to total rows. -1 means every row is unique; -0.5 means distinct count is half the row count.

Negative form is chosen by ANALYZE when the distinct count appears to grow with the table, so it scales as the table grows. This distinction matters when you override it manually.

The Sampling Problem

ANALYZE estimates n_distinct from a random sample (default ~300 × default_statistics_target rows), not a full scan. Estimating the number of distinct values from a sample is notoriously hard.

The classic failure: a high-cardinality column where distinct values are spread thinly. The sample sees few repeats, so the estimator under-counts badly. A column with 5 million real distinct values might be recorded as 50,000.

The planner then thinks a GROUP BY produces 50,000 groups, picks a hash aggregate sized for that, and spills to disk when reality hits 5 million.

Spotting a Bad n_distinct

Compare what the planner believes against ground truth. Run an exact distinct count and hold it next to pg_stats:

If n_distinct is stored as a small positive number but the real count is orders of magnitude larger, you have an underestimate. Remember to convert the negative ratio form: real estimate = -n_distinct × reltuples.

-- ground truth
SELECT count(DISTINCT customer_id) AS real_distinct
FROM orders;

-- what the planner thinks
SELECT n_distinct
FROM pg_stats
WHERE tablename = 'orders' AND attname = 'customer_id';

Overriding n_distinct

When you know the true cardinality better than the sampler ever will, pin it with ALTER TABLE ... ALTER COLUMN ... SET (n_distinct = ...).

Use the negative ratio form for columns that scale with table size — it survives growth. Use a positive integer only for a stable, bounded domain.

The override is stored in pg_attribute and applied on the next ANALYZE, so always re-analyze afterward.

-- distinct count grows ~linearly with rows: use the ratio form
ALTER TABLE orders
  ALTER COLUMN customer_id SET (n_distinct = -0.8);

ANALYZE orders;

n_distinct_inherited for Partitions

Partitioned tables have a second knob: n_distinct_inherited. The plain n_distinct override applies to the table's own rows; n_distinct_inherited applies to statistics gathered across the whole inheritance/partition tree.

For a partitioned orders table, queries usually scan the parent, so the inherited form is what the planner reads. Set both to be safe when a column is badly estimated.

ALTER TABLE orders
  ALTER COLUMN customer_id SET (n_distinct_inherited = -0.8);

ANALYZE orders;

MCV: Selectivity of Equality

For an equality predicate like status = 'shipped', the planner looks for the value in most_common_vals. If found, it uses the paired frequency from most_common_freqs directly. If not found, it assumes the value is one of the non-MCV values and spreads the remaining selectivity evenly across them.

So MCV accuracy decides whether a skewed predicate gets a sensible row estimate or a flat average that's wildly wrong for a hot value.

SELECT unnest(most_common_vals::text::text[]) AS val,
       unnest(most_common_freqs)            AS freq
FROM pg_stats
WHERE tablename = 'orders' AND attname = 'status';

When the MCV List Is Too Short

The MCV list length is capped by the column's statistics target. If a skewed column has 200 meaningfully frequent values but the target only keeps 100, the planner mis-estimates the values that fell off the list.

The fix is to widen the histogram and MCV list by raising the per-column statistics target, then re-analyze. This is the most common, lowest-risk correction for skewed equality and grouping estimates.

-- keep up to 1000 MCV entries + histogram buckets for this column
ALTER TABLE orders
  ALTER COLUMN status SET STATISTICS 1000;

ANALYZE orders;

Verifying the Fix with EXPLAIN

Never trust an override blindly — confirm the estimate moved toward reality. Run EXPLAIN ANALYZE and compare the planner's estimated rows to the actual rows the executor saw.

For grouping, look at the row count emitted by the HashAggregate / GroupAggregate node. A healthy plan has estimated and actual within a small factor of each other.

EXPLAIN (ANALYZE, BUFFERS)
SELECT customer_id, count(*)
FROM orders
GROUP BY customer_id;

Correlated Columns Need Extended Stats

Per-column MCV and n_distinct assume columns are independent. When two columns are correlated (e.g. city and country), the product of single-column selectivities under-estimates the combined group count.

That is exactly what multivariate CREATE STATISTICS ... (ndistinct, mcv) repairs — it stores a joint n_distinct and a joint MCV list for the column group, fixing multi-column GROUP BY and AND-predicate estimates.

CREATE STATISTICS orders_geo (ndistinct, mcv)
  ON city, country
  FROM orders;

ANALYZE orders;

Quick Check

You have a high-cardinality column whose distinct count grows linearly as the table grows, and ANALYZE keeps under-estimating it, wrecking GROUP BY plans. Which correction is best?

Recap

You learned to repair the two statistics that drive most cardinality errors:

  • Diagnose in pg_stats: read n_distinct, most_common_vals, most_common_freqs; compare against an exact count(DISTINCT ...).
  • n_distinct is positive for absolute counts, negative for a row-ratio. Override with ALTER COLUMN ... SET (n_distinct = ...), using the ratio form for growing columns and n_distinct_inherited for partitioned parents.
  • MCV drives equality selectivity. Lengthen it with SET STATISTICS when skewed values fall off the list.
  • Always ANALYZE after any change and confirm with EXPLAIN ANALYZE that estimated rows now track actual rows.
  • For correlated columns, reach for multivariate CREATE STATISTICS (ndistinct, mcv).

常见问题解答

「MCV 与不同值数量修正」课时是免费的吗?

是的 — 「MCV 与不同值数量修正」的完整文本可在网页上免费阅读。要进行交互式练习(内置代码编辑器和全天候 AI 导师)并解锁 PostgreSQL Performance & Query Optimization 课程的其余内容,请升级到 CoddyKit PRO。 PostgreSQL Performance & Query Optimization 课程共包含 4 节课。

「MCV 与不同值数量修正」这节课中我会学到什么?

使用 ndistinct 和最常见值统计信息,修正连接和分组估算。 你通过在浏览器中直接运行的动手代码来练习 PostgreSQL Performance & Query Optimization,全天候 AI 导师会在你学习这节课的过程中回答你的问题。

学习 PostgreSQL Performance & Query Optimization 需要有经验吗?

无需任何先前经验。CoddyKit 上的 PostgreSQL Performance & Query Optimization 课程适合初学者到高级学习者,你可以从这里开始或从头开始,按照自己的节奏学习。 这是第 3 节课,共 4 节。

「MCV 与不同值数量修正」课时需要多长时间?

大多数 CoddyKit 课程大约需要 5–10 分钟。每节课都很精短且互动,所以你能稳步进步,并在网页和应用中从离开的地方继续。

我能在这节 PostgreSQL Performance & Query Optimization 课中编写并运行代码吗?

能。每节 PostgreSQL Performance & Query Optimization 课都包含内置代码编辑器,你可以在浏览器中直接编写并运行真实代码,并获得即时 AI 反馈 — 无需本地设置。

此课程中的所有课时

  1. 查询规划器如何估算行数
  2. 相关列的多变量统计信息
  3. MCV 与不同值数量修正
  4. 将估算值与实际行数进行验证
← 返回 PostgreSQL Performance & Query Optimization