0Pricing
Deep Learning Academy · 课时

编写自定义数据集类

实现 __len__ 与 __getitem__

编写自定义数据集类 是 CoddyKit 上的免费 Deep Learning Academy 课时。 这是第 1 节课,共 4 节。 你可以在下方免费阅读本课时的完整内容 — 然后在浏览器中使用内置代码编辑器和全天候 AI 导师进行实践。 这是 Deep Learning Academy 学习路径的一部分,你的进度在网页和 CoddyKit 应用中同步。 Deep Learning Academy 课程共包含 4 节课。

本课时的部分内容尚未翻译,以英文显示。

Your Data Needs a Front Door

Before a model can learn, PyTorch needs a clean way to reach your samples one at a time. That front door is a Dataset class. 🚪

Start by Subclassing

You build a custom dataset by subclassing torch.utils.data.Dataset. PyTorch then knows exactly how to ask your object for data.

from torch.utils.data import Dataset

class MyData(Dataset):
    pass

Stash Your Data in __init__

The __init__ method runs once when you create the dataset. Use it to load file paths, arrays, or labels into the object's fields.

def __init__(self, X, y):
    self.X = X
    self.y = y

Two Methods Make It Work

A working dataset only needs two methods: __len__ to report its size and __getitem__ to fetch one sample. That is the whole contract.

__len__ Counts Your Samples

The __len__ method returns how many samples you have. PyTorch reads this to know when an epoch ends and how far an index can go.

def __len__(self):
    return len(self.X)

__getitem__ Returns One Sample

Given an index, __getitem__ returns a single sample, usually a feature and its label. This is where one row of data is handed over.

def __getitem__(self, idx):
    return self.X[idx], self.y[idx]

Return Tensors, Not Lists

__getitem__ should hand back tensors so the model can use them directly. Convert NumPy arrays or Python lists right here if needed.

import torch
x = torch.tensor(self.X[idx], dtype=torch.float32)

Lazy Loading for Big Data

For huge datasets, do not load everything in __init__. Instead read each file inside __getitem__ so only one sample sits in memory at a time.

Apply Transforms Per Sample

__getitem__ is the natural place to apply a transform, like resizing an image. Store the transform in __init__, then call it before returning.

if self.transform:
    x = self.transform(x)

Index It Like a List

Once built, your dataset behaves like a list. Calling len(ds) or ds[0] just triggers the two methods you defined. Test it before training.

ds = MyData(X, y)
print(len(ds), ds[0])

Now It Plugs Into Everything

This tidy interface is why a custom Dataset drops straight into a DataLoader. You write two methods and the rest of PyTorch just works.

Quick Check

Which method does PyTorch call to fetch a single sample by index?

Recap

A custom dataset subclasses Dataset and defines two methods: __len__ for its size and __getitem__ to return one sample. Two methods, full power. 🎉

常见问题解答

「编写自定义数据集类」课时是免费的吗?

是的 — 「编写自定义数据集类」的完整文本可在网页上免费阅读。要进行交互式练习(内置代码编辑器和全天候 AI 导师)并解锁 Deep Learning Academy 课程的其余内容,请升级到 CoddyKit PRO。 Deep Learning Academy 课程共包含 4 节课。

「编写自定义数据集类」这节课中我会学到什么?

实现 __len__ 与 __getitem__ 你通过在浏览器中直接运行的动手代码来练习 Deep Learning Academy,全天候 AI 导师会在你学习这节课的过程中回答你的问题。

学习 Deep Learning Academy 需要有经验吗?

无需任何先前经验。CoddyKit 上的 Deep Learning Academy 课程适合初学者到高级学习者,你可以从这里开始或从头开始,按照自己的节奏学习。 这是第 1 节课,共 4 节。

「编写自定义数据集类」课时需要多长时间?

大多数 CoddyKit 课程大约需要 5–10 分钟。每节课都很精短且互动,所以你能稳步进步,并在网页和应用中从离开的地方继续。

我能在这节 Deep Learning Academy 课中编写并运行代码吗?

能。每节 Deep Learning Academy 课都包含内置代码编辑器,你可以在浏览器中直接编写并运行真实代码,并获得即时 AI 反馈 — 无需本地设置。

此课程中的所有课时

  1. 编写自定义数据集类
  2. 批处理、打乱与 num_workers
  3. 使用 collate_fn 处理可变长度输入
  4. 归一化与标准化输入
← 返回 Deep Learning Academy