PostgreSQL Performance & Query Optimization · Lección

Upserts a escala con ON CONFLICT

Implemente una lógica de combinación eficiente para lotes grandes, evitando la contención de bloqueos y la fragmentación.

Lección 4 de 413 pasos

Upserts a escala con ON CONFLICT es una lección gratuita de PostgreSQL Performance & Query Optimization en CoddyKit. Esta es la lección 4 de 4. Puedes leer la lección completa abajo gratuitamente — luego la practicas en el navegador con un editor de código integrado y un tutor de IA 24/7. Forma parte de la ruta de aprendizaje de PostgreSQL Performance & Query Optimization, y tu progreso se sincroniza en la web y la app de CoddyKit. El curso de PostgreSQL Performance & Query Optimization incluye 4 lecciones en total.

Partes de esta lección aún no han sido traducidas y se muestran en inglés.

Why Upserts Need Care at Scale

An upsert inserts a row, but if it would collide with an existing key, it updates the existing row instead. PostgreSQL spells this INSERT ... ON CONFLICT.

For small workloads it is trivial. For ETL-sized batches (tens of thousands to millions of rows) the naive approach causes three problems:

  • Lock contention — concurrent writers fighting over the same rows or index pages.
  • Table bloat — every UPDATE writes a new row version (dead tuple) that VACUUM must later reclaim.
  • WAL and round-trip overhead — row-by-row upserts multiply network and transaction cost.

This lesson builds a merge that is both correct and throughput-friendly.

The ON CONFLICT Shape

ON CONFLICT requires a conflict target: the column(s) or constraint that define a duplicate. PostgreSQL needs a unique or exclusion constraint on that target to arbitrate.

The two actions are DO NOTHING (skip the colliding row) and DO UPDATE (merge new values in).

Inside DO UPDATE, the incoming row is exposed through the special EXCLUDED pseudo-table.

INSERT INTO products (sku, name, price)
VALUES ('A-100', 'Widget', 9.99)
ON CONFLICT (sku) DO UPDATE
  SET name  = EXCLUDED.name,
      price = EXCLUDED.price;

One Statement, Many Rows

The single most important throughput rule: batch your rows into one statement. A multi-row VALUES list (or a feeding SELECT) is parsed, planned, and committed once instead of N times.

This collapses N network round-trips and N transaction commits into one, often a 10-100x speedup over row-by-row upserts.

INSERT INTO products (sku, name, price)
VALUES
  ('A-100', 'Widget', 9.99),
  ('A-101', 'Gadget', 14.50),
  ('A-102', 'Gizmo', 7.25)
ON CONFLICT (sku) DO UPDATE
  SET name  = EXCLUDED.name,
      price = EXCLUDED.price;

Staging Table + INSERT...SELECT

For real ETL, load raw data into an unlogged staging table first (often via COPY, the fastest bulk path), then merge from staging into the target with one INSERT ... SELECT ... ON CONFLICT.

Benefits:

  • COPY avoids per-row INSERT overhead.
  • An UNLOGGED staging table skips WAL for the load phase.
  • You can dedupe and transform in the SELECT before merging.
CREATE UNLOGGED TABLE products_stage (LIKE products);

-- bulk load: COPY products_stage FROM '/data/products.csv' CSV;

INSERT INTO products (sku, name, price)
SELECT sku, name, price FROM products_stage
ON CONFLICT (sku) DO UPDATE
  SET name  = EXCLUDED.name,
      price = EXCLUDED.price;

Deduplicate the Batch First

A subtle but fatal error: a single INSERT statement cannot update the same target row twice. If your batch contains two rows with the same conflict key, PostgreSQL raises:

ERROR: ON CONFLICT DO UPDATE command cannot affect row a second time

Fix it by collapsing duplicates in the source before merging. DISTINCT ON keeps one row per key — typically the newest.

INSERT INTO products (sku, name, price)
SELECT DISTINCT ON (sku) sku, name, price
FROM products_stage
ORDER BY sku, updated_at DESC
ON CONFLICT (sku) DO UPDATE
  SET name  = EXCLUDED.name,
      price = EXCLUDED.price;

Skip No-Op Updates to Cut Bloat

Every DO UPDATE writes a new row version, even when the new values are identical to the old ones. Those dead tuples bloat the table and create extra work for VACUUM.

Add a WHERE clause to the DO UPDATE so it fires only when something actually changed. Use IS DISTINCT FROM so NULLs compare correctly.

INSERT INTO products (sku, name, price)
SELECT sku, name, price FROM products_stage
ON CONFLICT (sku) DO UPDATE
  SET name  = EXCLUDED.name,
      price = EXCLUDED.price
WHERE products.name  IS DISTINCT FROM EXCLUDED.name
   OR products.price IS DISTINCT FROM EXCLUDED.price;

Order Batches to Tame Lock Contention

When several ETL workers run concurrently, deadlocks appear if they touch the same keys in different orders. Worker 1 locks key X then Y; worker 2 locks Y then X — both block forever until PostgreSQL kills one.

Defenses:

  • Sort each batch by the conflict key so all workers acquire locks in the same order.
  • Partition work by key range so no two workers share keys.
  • Keep transactions short — long-held row locks magnify contention.
INSERT INTO products (sku, name, price)
SELECT sku, name, price
FROM products_stage
ORDER BY sku            -- consistent lock acquisition order
ON CONFLICT (sku) DO UPDATE
  SET name  = EXCLUDED.name,
      price = EXCLUDED.price;

Chunk Giant Merges

One enormous transaction merging millions of rows holds locks for a long time, balloons WAL, and blocks autovacuum from cleaning up. Break the work into chunks (for example 10k-50k rows) and commit between them.

Smaller transactions release locks sooner, let autovacuum keep pace, and make retries cheap after a failure. The trade-off is slightly more commit overhead — tune the chunk size against your hardware.

-- Merge one bounded slice; loop over key ranges from the app side.
INSERT INTO products (sku, name, price)
SELECT sku, name, price
FROM products_stage
WHERE sku >= 'A-0000' AND sku < 'A-5000'
ON CONFLICT (sku) DO UPDATE
  SET name  = EXCLUDED.name,
      price = EXCLUDED.price;

Choosing the Right Conflict Target

The conflict target must match an actual unique/primary-key or exclusion constraint. You can target by column list ON CONFLICT (sku) or by constraint name ON CONFLICT ON CONSTRAINT products_sku_key.

For partial unique indexes, repeat the index predicate so PostgreSQL can pick the right one:

-- Unique only among active rows
CREATE UNIQUE INDEX products_active_sku
  ON products (sku) WHERE is_active;

INSERT INTO products (sku, name, is_active)
VALUES ('A-100', 'Widget', true)
ON CONFLICT (sku) WHERE is_active DO UPDATE
  SET name = EXCLUDED.name;

ON CONFLICT vs MERGE

PostgreSQL 15+ adds the SQL-standard MERGE, which can INSERT, UPDATE, and DELETE in one pass. For high-throughput upserts, INSERT ... ON CONFLICT is usually still preferred:

  • ON CONFLICT is atomic against concurrent inserts — it handles a race where another transaction inserts the same key, retrying internally.
  • Classic MERGE can raise a unique-violation under heavy concurrency because it does not have that built-in conflict arbitration.

Use MERGE when you need DELETE branches or complex conditional logic; use ON CONFLICT for plain, concurrency-safe upserts.

Maintenance: VACUUM and Indexes

Even an optimized merge produces dead tuples on the updated rows. Keep performance steady with maintenance discipline:

  • Ensure autovacuum keeps up; for hot ETL tables lower autovacuum_vacuum_scale_factor so it triggers more often.
  • After a massive one-off backfill, run VACUUM (ANALYZE) to reclaim space and refresh planner statistics.
  • Every extra index on the target slows the merge (each insert/update maintains all of them) — keep only the indexes you truly need.
VACUUM (ANALYZE) products;

Quick Check

Test your understanding of safe, high-throughput merges.

Recap

Efficient upserts at scale come from a few combined habits:

  • Batch rows into one statement; for ETL, stage with COPY then INSERT ... SELECT ... ON CONFLICT.
  • Deduplicate the batch (DISTINCT ON) so no key appears twice.
  • Skip no-op updates with a WHERE ... IS DISTINCT FROM clause to curb dead tuples and bloat.
  • Order by the conflict key and partition work to avoid deadlocks; chunk huge merges into short transactions.
  • Prefer ON CONFLICT for concurrency-safe upserts; keep autovacuum healthy and indexes lean.

Together these turn a fragile row-by-row merge into a fast, low-contention ETL load.

Gratis para empezar

Aprende SQL con un tutor de IA — gratis

Escribe y ejecuta código real en tu navegador, obtén ayuda instantánea de un tutor de IA disponible 24/7 y continúa donde lo dejaste en la web o en la aplicación.

Cursos
22
Lecciones
88

Preguntas frecuentes

¿La lección «Upserts a escala con ON CONFLICT» es gratis?

Sí — el texto completo de «Upserts a escala con ON CONFLICT» es gratis para leer aquí en la web. Para practicarla de forma interactiva (editor de código integrado y tutor de IA 24/7) y desbloquear el resto del curso de PostgreSQL Performance & Query Optimization, actualiza a CoddyKit PRO. El curso de PostgreSQL Performance & Query Optimization incluye 4 lecciones en total.

¿Qué aprenderé en «Upserts a escala con ON CONFLICT»?

Implemente una lógica de combinación eficiente para lotes grandes, evitando la contención de bloqueos y la fragmentación. Practicas PostgreSQL Performance & Query Optimization con código real que ejecutas directamente en el navegador, y un tutor de IA 24/7 responde tus preguntas mientras trabajas en la lección.

¿Necesito experiencia previa para empezar PostgreSQL Performance & Query Optimization?

No se requiere experiencia previa. PostgreSQL Performance & Query Optimization en CoddyKit está estructurado para principiantes hasta estudiantes avanzados, así que puedes empezar aquí o desde el inicio y avanzar a tu ritmo. Esta es la lección 4 de 4.

¿Cuánto tiempo toma la lección «Upserts a escala con ON CONFLICT»?

La mayoría de las lecciones de CoddyKit toman alrededor de 5–10 minutos. Cada una es compacta e interactiva, así que avanzas constantemente y retomas exactamente por donde dejaste en la web y la app.

¿Puedo escribir y ejecutar código en esta lección de PostgreSQL Performance & Query Optimization?

Sí. Cada lección de PostgreSQL Performance & Query Optimization incluye un editor de código integrado, así que escribes y ejecutas código real directamente en tu navegador y obtienes retroalimentación instantánea de IA — sin configuración local necesaria.

Todas las lecciones de este curso

  1. Rendimiento de COPY frente a INSERT de varias filas
  2. Aplazamiento de índices y restricciones durante la carga
  3. Ajuste de WAL y checkpoints para la ingesta
  4. Upserts a escala con ON CONFLICT
← Volver a PostgreSQL Performance & Query Optimization