การแก้ปัญหาคำค้นหา N+1 ด้วย DataLoaders
รวมชุดและแคชการค้นหาฐานข้อมูลด้วยตัวโหลดข้อมูล เพื่อกำจัดคำค้นหา N+1 ที่เพิ่มขึ้นอย่างรวดเร็วในตัวแก้ไข
การแก้ปัญหาคำค้นหา N+1 ด้วย DataLoaders เป็นบทเรียน FastAPI Backend Development Bootcamp ฟรีบน CoddyKit นี่คือบทเรียนที่ 2 จากทั้งหมด 4 บทเรียน คุณสามารถอ่านบทเรียนทั้งหมดด้านล่างฟรี — จากนั้นลองปฏิบัติด้วยตัวคุณเองในเบราว์เซอร์พร้อมตัวแก้ไขโค้ดในตัวและติวเตอร์ AI ตลอด 24/7 บทเรียนนี้เป็นส่วนหนึ่งของเส้นทางการเรียน FastAPI Backend Development Bootcamp และความก้าวหน้าของคุณจะซิงค์ข้ามเว็บและแอป CoddyKit คอร์ส FastAPI Backend Development Bootcamp มีบทเรียนทั้งหมด 4 บทเรียน
บางส่วนของบทเรียนนี้ยังไม่ได้รับการแปล และแสดงเป็นภาษาอังกฤษ
The N+1 Problem in GraphQL
GraphQL lets clients ask for nested data in a single request, like a list of posts and each post's author. The danger is hidden in the resolvers.
Suppose you fetch 100 posts with 1 query, then resolve each post's author by running one query per post. That is 1 + 100 = 101 queries — the classic N+1 problem.
- 1 query to load the list (the 1)
- N queries, one per item, to load a related field (the N)
At scale this destroys latency and hammers the database. DataLoaders are the standard fix.
Seeing N+1 in a Strawberry Resolver
Here is a naive Strawberry resolver that triggers N+1. Each author resolver issues its own database call.
If a query returns 50 posts, this author resolver fires 50 separate SELECT statements. The list query plus those 50 lookups is the N+1 explosion.
import strawberry
@strawberry.type
class Author:
id: int
name: str
@strawberry.type
class Post:
id: int
title: str
author_id: int
@strawberry.field
async def author(self) -> Author:
# BAD: one DB round-trip per post -> N+1
row = await db.fetch_one(
"SELECT id, name FROM authors WHERE id = :id",
{"id": self.author_id},
)
return Author(id=row["id"], name=row["name"])The Core Idea: Batch and Cache
A DataLoader solves N+1 with two techniques:
- Batching: instead of resolving each
author_idimmediately, the loader collects all the keys requested during one tick of the event loop and resolves them together in a single batched query (e.g.WHERE id = ANY(...)). - Caching: within a single request, the same key is only fetched once. Asking for author 7 ten times yields one lookup.
The result: 1 query for the posts + 1 batched query for all authors = 2 queries instead of 101.
How Batching Works on the Event Loop
Strawberry's DataLoader relies on the asyncio event loop. When several resolvers call loader.load(key), the loader does not run immediately. It records each key and returns a pending awaitable.
On the next tick, the loader takes every queued key, calls your batch function once with the full list of keys, and then resolves each individual awaitable with its matching result.
This is why DataLoaders only work in async code: the deferral mechanism depends on the loop scheduling the batch dispatch after the current synchronous work finishes.
Writing the Batch Load Function
The heart of a DataLoader is the batch function. It receives a list of keys and must return a list of results in the exact same order as the keys.
Two non-negotiable rules:
- The returned list length must equal the keys length.
- Result at index
imust correspond tokeys[i]. Missing rows should map toNone(or an Exception), never be dropped.
Below we map rows by id, then re-emit them in key order.
from typing import List, Optional
async def load_authors(keys: List[int]) -> List[Optional[Author]]:
rows = await db.fetch_all(
"SELECT id, name FROM authors WHERE id = ANY(:ids)",
{"ids": keys},
)
by_id = {row["id"]: Author(id=row["id"], name=row["name"]) for row in rows}
# Preserve order; None for missing keys
return [by_id.get(key) for key in keys]Order Alignment Demonstrated
The order-preservation contract is the most common source of DataLoader bugs. Here is a standalone simulation: rows arrive in arbitrary order from the database, but we must return them aligned to the requested keys.
Run this to see how a lookup dict plus a key-ordered comprehension guarantees correct alignment even when the DB returns rows out of order or omits a missing key.
def batch_load(keys, rows):
by_id = {row["id"]: row["name"] for row in rows}
return [by_id.get(k) for k in keys]
keys = [3, 1, 7, 4]
# DB returns rows shuffled and is missing id=7
rows = [
{"id": 1, "name": "Ada"},
{"id": 4, "name": "Linus"},
{"id": 3, "name": "Grace"},
]
result = batch_load(keys, rows)
print(result) # ['Grace', 'Ada', None, 'Linus']
assert len(result) == len(keys)
for key, name in zip(keys, result):
print(f"key={key} -> {name}")Creating a DataLoader in Strawberry
Strawberry ships a DataLoader class. You construct it with your batch function. Calling .load(key) returns an awaitable that resolves after batching.
Critically, a DataLoader instance holds a per-instance cache. You must create a fresh loader per request so stale data and cross-user leakage never happen. We will wire that up next via context.
from strawberry.dataloader import DataLoader
# batch function from the previous scene
author_loader = DataLoader(load_fn=load_authors)
# Inside a resolver you would now write:
# author = await author_loader.load(self.author_id)
# Many concurrent .load() calls collapse into ONE call to load_authors.Per-Request Loaders via GraphQL Context
The clean place to store request-scoped loaders is the GraphQL context. With FastAPI + Strawberry you override get_context to build fresh loaders on every request.
This guarantees the batch window and the cache are isolated to one request — exactly the lifetime you want.
from strawberry.fastapi import GraphQLRouter
from strawberry.dataloader import DataLoader
async def get_context() -> dict:
return {
"author_loader": DataLoader(load_fn=load_authors),
# one loader per relation, all rebuilt per request
}
graphql_app = GraphQLRouter(schema, context_getter=get_context)
# app.include_router(graphql_app, prefix="/graphql")Using the Loader Inside a Resolver
Now the author resolver reads the loader from info.context and calls .load(). Strawberry injects info when you declare it as a parameter.
Even though this resolver runs once per post, all those .load() calls are batched into a single SELECT ... WHERE id = ANY(...) — N+1 is gone.
import strawberry
from strawberry.types import Info
@strawberry.type
class Post:
id: int
title: str
author_id: int
@strawberry.field
async def author(self, info: Info) -> Author:
loader = info.context["author_loader"]
return await loader.load(self.author_id)Caching Wins and Their Limits
Within one request the loader caches by key, so repeated load(7) calls hit the DB once. This is great for fan-out queries where the same author appears across many posts.
Watch the trade-offs:
- The cache is per request by design — never share a loader across requests or you serve stale data.
- If a record changes mid-request and you re-read it, you get the cached copy. Call
loader.clear(key)after a mutation to invalidate. - The cache key is the raw key value, so keep keys hashable and consistent (e.g. always
int, not sometimesstr).
Loading Collections and Tuple Keys
DataLoaders are not only for one-to-one lookups. For one-to-many (a post's comments), the batch function returns a list per key. Group the rows by foreign key, then emit one list per requested key (empty list if none).
For composite lookups, use a hashable tuple as the key, e.g. (post_id, locale). Just keep the type stable so caching stays correct.
from collections import defaultdict
async def load_comments(post_ids):
rows = await db.fetch_all(
"SELECT id, post_id, body FROM comments WHERE post_id = ANY(:ids)",
{"ids": post_ids},
)
grouped = defaultdict(list)
for row in rows:
grouped[row["post_id"]].append(row)
# one list per key, in key order
return [grouped.get(pid, []) for pid in post_ids]Quick Check: DataLoader Lifetime
A teammate creates a single module-level DataLoader and reuses it for the whole app to "save memory." Why is this the wrong choice for a multi-user FastAPI GraphQL service?
Recap: DataLoaders Defeat N+1
You learned how to eliminate N+1 query explosions in Strawberry + FastAPI resolvers:
- N+1 happens when a nested resolver issues one query per parent item.
- A DataLoader fixes it by batching all keys from one event-loop tick into a single query and caching repeated keys within the request.
- The batch function must return results aligned to the input keys, same length, same order, with
Noneor empty lists for misses. - Build loaders per request in
get_contextand read them frominfo.contextinside resolvers. - Use lists-per-key for one-to-many relations and hashable tuple keys for composite lookups; call
clear()after mutations.
With this pattern, deeply nested GraphQL queries stay fast and your database stays calm.
คำถามที่พบบ่อย
บทเรียน “การแก้ปัญหาคำค้นหา N+1 ด้วย DataLoaders” ฟรีหรือไม่
ใช่ — ข้อความเต็มของ “การแก้ปัญหาคำค้นหา N+1 ด้วย DataLoaders” ฟรีให้อ่านที่นี่บนเว็บ เพื่อปฏิบัติแบบโต้ตอบ (ตัวแก้ไขโค้ดในตัวและติวเตอร์ AI ตลอด 24/7) และปลดล็อคส่วนที่เหลือของคอร์ส FastAPI Backend Development Bootcamp ให้อัปเกรดเป็น CoddyKit PRO คอร์ส FastAPI Backend Development Bootcamp มีบทเรียนทั้งหมด 4 บทเรียน
คุณจะเรียนรู้อะไรในบทเรียน “การแก้ปัญหาคำค้นหา N+1 ด้วย DataLoaders”
รวมชุดและแคชการค้นหาฐานข้อมูลด้วยตัวโหลดข้อมูล เพื่อกำจัดคำค้นหา N+1 ที่เพิ่มขึ้นอย่างรวดเร็วในตัวแก้ไข คุณปฏิบัติ FastAPI Backend Development Bootcamp ด้วยโค้ดที่ใช้งานได้จริงที่คุณเรียกใช้โดยตรงในเบราว์เซอร์ และติวเตอร์ AI ตลอด 24/7 ตอบคำถามของคุณขณะที่คุณไปผ่านบทเรียน
คุณต้องมีประสบการณ์ก่อนที่จะเริ่มเรียน FastAPI Backend Development Bootcamp หรือไม่
ไม่จำเป็นต้องมีประสบการณ์มาก่อน FastAPI Backend Development Bootcamp บน CoddyKit ออกแบบมาสำหรับผู้เริ่มต้นไปจนถึงผู้เรียนขั้นสูง คุณสามารถเริ่มต้นที่นี่หรือเริ่มจากตัวแรกและเรียนด้วยความเร็วของคุณเอง นี่คือบทเรียนที่ 2 จากทั้งหมด 4 บทเรียน
บทเรียน “การแก้ปัญหาคำค้นหา N+1 ด้วย DataLoaders” ใช้เวลานานแค่ไหน
บทเรียน CoddyKit ส่วนใหญ่ใช้เวลาประมาณ 5–10 นาที แต่ละบทเรียนจึงสั้นและเป็นแบบโต้ตอบ คุณสามารถก้าวหน้าอย่างต่อเนื่องและกลับมาเรียนต่อจากตรงที่เพิ่งหยุดบนเว็บและแอปได้เลย
ฉันเขียนและรันโค้ดในบทเรียน FastAPI Backend Development Bootcamp นี้ได้ไหม
ได้ บทเรียน FastAPI Backend Development Bootcamp ทุกบทมีตัวแก้ไขโค้ดในตัว คุณจึงเขียนและรันโค้ดจริงได้เลยในเบราว์เซอร์ และได้รับข้อเสนอแนะจาก AI ในทันที — ไม่ต้องติดตั้งในเครื่องของคุณ
บทเรียนทั้งหมดในหลักสูตรนี้
- การกำหนดชนิด คำค้นหา และมิวเทชัน
- การแก้ปัญหาคำค้นหา N+1 ด้วย DataLoaders
- การสมัครรับ GraphQL แบบเรียลไทม์
- การวิเคราะห์ต้นทุนคำค้นหาและการจำกัดความลึก