设计 DynamoDB 表
学习设计高效 DynamoDB 表架构的最佳实践,重点掌握分区键和排序键,以实现最佳性能
设计 DynamoDB 表 是 CoddyKit 上的免费 Serverless Backend with AWS Lambda & API Gateway 课时。 这是第 2 节课,共 4 节。 你可以在下方免费阅读本课时的完整内容 — 然后在浏览器中使用内置代码编辑器和全天候 AI 导师进行实践。 这是 Serverless Backend with AWS Lambda & API Gateway 学习路径的一部分,你的进度在网页和 CoddyKit 应用中同步。 Serverless Backend with AWS Lambda & API Gateway 课程共包含 4 节课。
本课时的部分内容尚未翻译,以英文显示。
Why DynamoDB Design Matters
Welcome to designing DynamoDB tables! Unlike traditional relational databases, DynamoDB is a NoSQL database that requires a different approach to schema design.
- Its serverless nature and performance at scale depend heavily on how you design your tables and choose your keys.
- A well-designed table ensures fast, consistent performance and cost efficiency.
- A poor design can lead to slow queries, high costs, and operational headaches.
The Partition Key (PK)
Every item in a DynamoDB table is uniquely identified by its primary key. The first part of any primary key is the Partition Key (sometimes called a Hash Key).
- DynamoDB uses the Partition Key's value as input to an internal hash function.
- This function determines the physical partition (storage location) where your data is stored.
- Good Partition Keys distribute data evenly across partitions, preventing 'hot spots' and ensuring scalability.
The Sort Key (SK)
The second part of a primary key, if you choose to have one, is the Sort Key (sometimes called a Range Key).
- Items with the same Partition Key are grouped together and sorted by their Sort Key value.
- This allows for efficient range queries (e.g., 'all orders from a user within a date range').
- The combination of Partition Key and Sort Key must be unique for each item in the table.
Understanding Primary Keys
A table's primary key can be either a simple primary key (only a Partition Key) or a composite primary key (a Partition Key and a Sort Key).
- Simple Primary Key: Ideal when each item needs a unique identifier, like a
userIdfor a Users table. - Composite Primary Key: Best for one-to-many relationships or when you need to query items that share a common partition key but differ by a secondary identifier, like
userIdandorderIdfor an Orders table.
Simple PK in Action
Let's consider a Users table where each user has a unique userId. We can use userId as the simple Partition Key. Here's how an item might be added:
import boto3
# This client is conceptual for illustration.
# In a real app, it would connect to your DynamoDB table.
class MockDynamoDBTable:
def __init__(self, name):
self.name = name
self.items = {}
def put_item(self, Item):
pk = Item['userId']
if pk in self.items:
print(f"Warning: Item with userId '{pk}' already exists. Overwriting.")
self.items[pk] = Item
print(f"Item added/updated in '{self.name}': {Item}")
# Simulate a DynamoDB table named 'Users'
table = MockDynamoDBTable('Users')
def add_user_item():
table.put_item(
Item={
'userId': 'user123',
'username': 'Alice',
'email': 'alice@example.com'
}
)
if __name__ == "__main__":
add_user_item()Composite PK in Action
Now, imagine an Orders table where a user can have multiple orders. We'd use userId as the Partition Key and orderId as the Sort Key. This allows us to retrieve all orders for a specific user, sorted by orderId.
import boto3
# This client is conceptual for illustration.
# In a real app, it would connect to your DynamoDB table.
class MockDynamoDBTable:
def __init__(self, name):
self.name = name
self.items = {}
def put_item(self, Item):
pk = Item['userId']
sk = Item['orderId']
if pk not in self.items:
self.items[pk] = {}
self.items[pk][sk] = Item
print(f"Item added/updated in '{self.name}': {Item}")
# Simulate a DynamoDB table named 'Orders'
table = MockDynamoDBTable('Orders') # Assume PK: userId, SK: orderId
def add_order_item():
table.put_item(
Item={
'userId': 'user123',
'orderId': 'order456',
'itemCount': 2,
'totalAmount': 50.00,
'orderDate': '2023-10-26'
}
)
if __name__ == "__main__":
add_order_item()Designing for Access Patterns
The most crucial aspect of DynamoDB design is understanding your access patterns. You should design your primary keys around how your application will query the data, not just how the data looks.
- What queries will you make? (e.g., 'Get all products by category', 'Get a specific user's latest posts').
- What data will be returned? (e.g., single item, list of items).
- Your Partition Key should typically be the attribute you query on most frequently for specific items or groups of items.
- Your Sort Key allows for flexible queries within that group (e.g., range queries, reverse order).
Key Selection Best Practices
Choosing the right keys is vital for performance:
- High Cardinality: Keys should have many unique values to prevent 'hot spots' on a single partition.
- Even Distribution: Values should be accessed roughly equally. Avoid keys where a few values are queried much more often than others.
- Query Efficiency: Design keys so that most common queries can be satisfied using
GetItem(PK only) orQuery(PK + optional SK condition). - Avoid Scans: Operations that read every item in a table (
Scan) are inefficient and costly. Your design should minimize the need for them.
Modeling One-to-Many Relationships
Composite Primary Keys are excellent for modeling one-to-many relationships, which are very common. For example, a user has many posts:
- Partition Key:
USER#<userId>(a common pattern to prefix keys for clarity). - Sort Key:
POST#<postId>. - This design allows you to fetch all posts for a user by querying only on the Partition Key, and even specific posts by adding a Sort Key condition.
This approach keeps related data together, enabling efficient retrieval.
Design Challenge
You are designing a table to store user comments on various articles. Your primary access patterns are:
- Get all comments for a specific article.
- Get all comments made by a specific user.
Which primary key design would best support getting all comments for a specific article efficiently?
Key Design Recap
Great job! In this lesson, you learned the foundational principles of designing DynamoDB tables:
- The importance of Partition Keys for data distribution.
- The utility of Sort Keys for ordering and range queries.
- How Primary Keys (simple or composite) uniquely identify items.
- The critical role of access patterns in driving your design choices.
- Best practices for selecting keys to ensure performance and cost efficiency.
Mastering these concepts is key to building scalable and performant serverless applications with DynamoDB!
常见问题解答
「设计 DynamoDB 表」课时是免费的吗?
是的 — 「设计 DynamoDB 表」的完整文本可在网页上免费阅读。要进行交互式练习(内置代码编辑器和全天候 AI 导师)并解锁 Serverless Backend with AWS Lambda & API Gateway 课程的其余内容,请升级到 CoddyKit PRO。 Serverless Backend with AWS Lambda & API Gateway 课程共包含 4 节课。
「设计 DynamoDB 表」这节课中我会学到什么?
学习设计高效 DynamoDB 表架构的最佳实践,重点掌握分区键和排序键,以实现最佳性能 你通过在浏览器中直接运行的动手代码来练习 Serverless Backend with AWS Lambda & API Gateway,全天候 AI 导师会在你学习这节课的过程中回答你的问题。
学习 Serverless Backend with AWS Lambda & API Gateway 需要有经验吗?
无需任何先前经验。CoddyKit 上的 Serverless Backend with AWS Lambda & API Gateway 课程适合初学者到高级学习者,你可以从这里开始或从头开始,按照自己的节奏学习。 这是第 2 节课,共 4 节。
「设计 DynamoDB 表」课时需要多长时间?
大多数 CoddyKit 课程大约需要 5–10 分钟。每节课都很精短且互动,所以你能稳步进步,并在网页和应用中从离开的地方继续。
我能在这节 Serverless Backend with AWS Lambda & API Gateway 课中编写并运行代码吗?
能。每节 Serverless Backend with AWS Lambda & API Gateway 课都包含内置代码编辑器,你可以在浏览器中直接编写并运行真实代码,并获得即时 AI 反馈 — 无需本地设置。
此课程中的所有课时
- DynamoDB 简介
- 设计 DynamoDB 表
- Lambda 与 DynamoDB 集成
- 使用二级索引查询:GSI 与 LSI