使用 ON CONFLICT 实现大规模插入或更新
为大批量数据实现高效的合并逻辑,同时避免锁竞争和膨胀。
使用 ON CONFLICT 实现大规模插入或更新 是 CoddyKit 上的免费 PostgreSQL Performance & Query Optimization 课时。 这是第 4 节课,共 4 节。 你可以在下方免费阅读本课时的完整内容 — 然后在浏览器中使用内置代码编辑器和全天候 AI 导师进行实践。 这是 PostgreSQL Performance & Query Optimization 学习路径的一部分,你的进度在网页和 CoddyKit 应用中同步。 PostgreSQL Performance & Query Optimization 课程共包含 4 节课。
本课时的部分内容尚未翻译,以英文显示。
Why Upserts Need Care at Scale
An upsert inserts a row, but if it would collide with an existing key, it updates the existing row instead. PostgreSQL spells this INSERT ... ON CONFLICT.
For small workloads it is trivial. For ETL-sized batches (tens of thousands to millions of rows) the naive approach causes three problems:
- Lock contention — concurrent writers fighting over the same rows or index pages.
- Table bloat — every UPDATE writes a new row version (dead tuple) that VACUUM must later reclaim.
- WAL and round-trip overhead — row-by-row upserts multiply network and transaction cost.
This lesson builds a merge that is both correct and throughput-friendly.
The ON CONFLICT Shape
ON CONFLICT requires a conflict target: the column(s) or constraint that define a duplicate. PostgreSQL needs a unique or exclusion constraint on that target to arbitrate.
The two actions are DO NOTHING (skip the colliding row) and DO UPDATE (merge new values in).
Inside DO UPDATE, the incoming row is exposed through the special EXCLUDED pseudo-table.
INSERT INTO products (sku, name, price)
VALUES ('A-100', 'Widget', 9.99)
ON CONFLICT (sku) DO UPDATE
SET name = EXCLUDED.name,
price = EXCLUDED.price;One Statement, Many Rows
The single most important throughput rule: batch your rows into one statement. A multi-row VALUES list (or a feeding SELECT) is parsed, planned, and committed once instead of N times.
This collapses N network round-trips and N transaction commits into one, often a 10-100x speedup over row-by-row upserts.
INSERT INTO products (sku, name, price)
VALUES
('A-100', 'Widget', 9.99),
('A-101', 'Gadget', 14.50),
('A-102', 'Gizmo', 7.25)
ON CONFLICT (sku) DO UPDATE
SET name = EXCLUDED.name,
price = EXCLUDED.price;Staging Table + INSERT...SELECT
For real ETL, load raw data into an unlogged staging table first (often via COPY, the fastest bulk path), then merge from staging into the target with one INSERT ... SELECT ... ON CONFLICT.
Benefits:
COPYavoids per-row INSERT overhead.- An
UNLOGGEDstaging table skips WAL for the load phase. - You can dedupe and transform in the SELECT before merging.
CREATE UNLOGGED TABLE products_stage (LIKE products);
-- bulk load: COPY products_stage FROM '/data/products.csv' CSV;
INSERT INTO products (sku, name, price)
SELECT sku, name, price FROM products_stage
ON CONFLICT (sku) DO UPDATE
SET name = EXCLUDED.name,
price = EXCLUDED.price;Deduplicate the Batch First
A subtle but fatal error: a single INSERT statement cannot update the same target row twice. If your batch contains two rows with the same conflict key, PostgreSQL raises:
ERROR: ON CONFLICT DO UPDATE command cannot affect row a second time
Fix it by collapsing duplicates in the source before merging. DISTINCT ON keeps one row per key — typically the newest.
INSERT INTO products (sku, name, price)
SELECT DISTINCT ON (sku) sku, name, price
FROM products_stage
ORDER BY sku, updated_at DESC
ON CONFLICT (sku) DO UPDATE
SET name = EXCLUDED.name,
price = EXCLUDED.price;Skip No-Op Updates to Cut Bloat
Every DO UPDATE writes a new row version, even when the new values are identical to the old ones. Those dead tuples bloat the table and create extra work for VACUUM.
Add a WHERE clause to the DO UPDATE so it fires only when something actually changed. Use IS DISTINCT FROM so NULLs compare correctly.
INSERT INTO products (sku, name, price)
SELECT sku, name, price FROM products_stage
ON CONFLICT (sku) DO UPDATE
SET name = EXCLUDED.name,
price = EXCLUDED.price
WHERE products.name IS DISTINCT FROM EXCLUDED.name
OR products.price IS DISTINCT FROM EXCLUDED.price;Order Batches to Tame Lock Contention
When several ETL workers run concurrently, deadlocks appear if they touch the same keys in different orders. Worker 1 locks key X then Y; worker 2 locks Y then X — both block forever until PostgreSQL kills one.
Defenses:
- Sort each batch by the conflict key so all workers acquire locks in the same order.
- Partition work by key range so no two workers share keys.
- Keep transactions short — long-held row locks magnify contention.
INSERT INTO products (sku, name, price)
SELECT sku, name, price
FROM products_stage
ORDER BY sku -- consistent lock acquisition order
ON CONFLICT (sku) DO UPDATE
SET name = EXCLUDED.name,
price = EXCLUDED.price;Chunk Giant Merges
One enormous transaction merging millions of rows holds locks for a long time, balloons WAL, and blocks autovacuum from cleaning up. Break the work into chunks (for example 10k-50k rows) and commit between them.
Smaller transactions release locks sooner, let autovacuum keep pace, and make retries cheap after a failure. The trade-off is slightly more commit overhead — tune the chunk size against your hardware.
-- Merge one bounded slice; loop over key ranges from the app side.
INSERT INTO products (sku, name, price)
SELECT sku, name, price
FROM products_stage
WHERE sku >= 'A-0000' AND sku < 'A-5000'
ON CONFLICT (sku) DO UPDATE
SET name = EXCLUDED.name,
price = EXCLUDED.price;Choosing the Right Conflict Target
The conflict target must match an actual unique/primary-key or exclusion constraint. You can target by column list ON CONFLICT (sku) or by constraint name ON CONFLICT ON CONSTRAINT products_sku_key.
For partial unique indexes, repeat the index predicate so PostgreSQL can pick the right one:
-- Unique only among active rows
CREATE UNIQUE INDEX products_active_sku
ON products (sku) WHERE is_active;
INSERT INTO products (sku, name, is_active)
VALUES ('A-100', 'Widget', true)
ON CONFLICT (sku) WHERE is_active DO UPDATE
SET name = EXCLUDED.name;ON CONFLICT vs MERGE
PostgreSQL 15+ adds the SQL-standard MERGE, which can INSERT, UPDATE, and DELETE in one pass. For high-throughput upserts, INSERT ... ON CONFLICT is usually still preferred:
ON CONFLICTis atomic against concurrent inserts — it handles a race where another transaction inserts the same key, retrying internally.- Classic
MERGEcan raise a unique-violation under heavy concurrency because it does not have that built-in conflict arbitration.
Use MERGE when you need DELETE branches or complex conditional logic; use ON CONFLICT for plain, concurrency-safe upserts.
Maintenance: VACUUM and Indexes
Even an optimized merge produces dead tuples on the updated rows. Keep performance steady with maintenance discipline:
- Ensure autovacuum keeps up; for hot ETL tables lower
autovacuum_vacuum_scale_factorso it triggers more often. - After a massive one-off backfill, run
VACUUM (ANALYZE)to reclaim space and refresh planner statistics. - Every extra index on the target slows the merge (each insert/update maintains all of them) — keep only the indexes you truly need.
VACUUM (ANALYZE) products;Quick Check
Test your understanding of safe, high-throughput merges.
Recap
Efficient upserts at scale come from a few combined habits:
- Batch rows into one statement; for ETL, stage with
COPYthenINSERT ... SELECT ... ON CONFLICT. - Deduplicate the batch (
DISTINCT ON) so no key appears twice. - Skip no-op updates with a
WHERE ... IS DISTINCT FROMclause to curb dead tuples and bloat. - Order by the conflict key and partition work to avoid deadlocks; chunk huge merges into short transactions.
- Prefer
ON CONFLICTfor concurrency-safe upserts; keep autovacuum healthy and indexes lean.
Together these turn a fragile row-by-row merge into a fast, low-contention ETL load.
常见问题解答
「使用 ON CONFLICT 实现大规模插入或更新」课时是免费的吗?
是的 — 「使用 ON CONFLICT 实现大规模插入或更新」的完整文本可在网页上免费阅读。要进行交互式练习(内置代码编辑器和全天候 AI 导师)并解锁 PostgreSQL Performance & Query Optimization 课程的其余内容,请升级到 CoddyKit PRO。 PostgreSQL Performance & Query Optimization 课程共包含 4 节课。
「使用 ON CONFLICT 实现大规模插入或更新」这节课中我会学到什么?
为大批量数据实现高效的合并逻辑,同时避免锁竞争和膨胀。 你通过在浏览器中直接运行的动手代码来练习 PostgreSQL Performance & Query Optimization,全天候 AI 导师会在你学习这节课的过程中回答你的问题。
学习 PostgreSQL Performance & Query Optimization 需要有经验吗?
无需任何先前经验。CoddyKit 上的 PostgreSQL Performance & Query Optimization 课程适合初学者到高级学习者,你可以从这里开始或从头开始,按照自己的节奏学习。 这是第 4 节课,共 4 节。
「使用 ON CONFLICT 实现大规模插入或更新」课时需要多长时间?
大多数 CoddyKit 课程大约需要 5–10 分钟。每节课都很精短且互动,所以你能稳步进步,并在网页和应用中从离开的地方继续。
我能在这节 PostgreSQL Performance & Query Optimization 课中编写并运行代码吗?
能。每节 PostgreSQL Performance & Query Optimization 课都包含内置代码编辑器,你可以在浏览器中直接编写并运行真实代码,并获得即时 AI 反馈 — 无需本地设置。
此课程中的所有课时
- COPY 与多行 INSERT 的吞吐量
- 导入期间延后建立索引与约束
- 调整 WAL 与检查点以优化数据导入
- 使用 ON CONFLICT 实现大规模插入或更新