0Pricing
Data Science Academy · 课时

修正 dtypes 并删除重复行

转换类型并删除重复项

修正 dtypes 并删除重复行 是 CoddyKit 上的免费 Data Science Academy 课时。 这是第 4 节课,共 4 节。 你可以在下方免费阅读本课时的完整内容 — 然后在浏览器中使用内置代码编辑器和全天候 AI 导师进行实践。 这是 Data Science Academy 学习路径的一部分,你的进度在网页和 CoddyKit 应用中同步。 Data Science Academy 课程共包含 4 节课。

本课时的部分内容尚未翻译,以英文显示。

Clean Goes Beyond Gaps

Filling missing values is only half the job. Clean data also needs correct dtypes and no accidental duplicate rows muddying your counts. 🧹

What a dtype Is

Every column has a dtype, its data type, like int, float, object, or datetime. The dtype controls which operations the column even allows.

df.dtypes

Numbers Stuck as Text

A frequent mess: numbers loaded as object strings because one stray symbol crept in. You cannot do math on them until the type is fixed.

Coerce With to_numeric

Convert a text column to numbers with to_numeric. Pass errors='coerce' to turn anything unparseable into NaN instead of crashing.

df['price'] = pd.to_numeric(df['price'], errors='coerce')

Cast With astype

When a column is already clean, astype changes its type directly, for example turning whole-number floats back into compact integers.

df['count'] = df['count'].astype(int)

Fix Date Columns

Dates often arrive as text. Convert them with to_datetime so you can sort, filter by range, and extract parts like month later.

df['date'] = pd.to_datetime(df['date'])

Category to Save Memory

Columns with few repeated text values shrink dramatically as the category dtype, which stores each label once and references it.

df['city'] = df['city'].astype('category')

Spotting Duplicate Rows

Repeated records inflate totals and averages. Find them with duplicated(), which flags each row that has appeared before as True.

df.duplicated().sum()

Dropping Duplicates

Remove repeats with drop_duplicates(). By default it keeps the first occurrence of each row and discards the rest.

df = df.drop_duplicates()

Duplicates by Key Columns

Sometimes only certain columns define a duplicate. Pass subset to detect repeats based on a key like an id, ignoring the other columns.

df.drop_duplicates(subset=['user_id'])

Verify Before You Move On

After fixing types and repeats, run info() once more. Confirming dtypes and row counts protects you from cleaning errors slipping downstream.

df.info()

Quick Check

A price column loaded as object text. Which call safely turns bad entries into NaN?

Recap: Types and Dupes Tamed

You now fix dtypes with to_numeric, astype, and to_datetime, and clear repeats with drop_duplicates. Your table is finally analysis-ready.

常见问题解答

「修正 dtypes 并删除重复行」课时是免费的吗?

是的 — 「修正 dtypes 并删除重复行」的完整文本可在网页上免费阅读。要进行交互式练习(内置代码编辑器和全天候 AI 导师)并解锁 Data Science Academy 课程的其余内容,请升级到 CoddyKit PRO。 Data Science Academy 课程共包含 4 节课。

「修正 dtypes 并删除重复行」这节课中我会学到什么?

转换类型并删除重复项 你通过在浏览器中直接运行的动手代码来练习 Data Science Academy,全天候 AI 导师会在你学习这节课的过程中回答你的问题。

学习 Data Science Academy 需要有经验吗?

无需任何先前经验。CoddyKit 上的 Data Science Academy 课程适合初学者到高级学习者,你可以从这里开始或从头开始,按照自己的节奏学习。 这是第 4 节课,共 4 节。

「修正 dtypes 并删除重复行」课时需要多长时间?

大多数 CoddyKit 课程大约需要 5–10 分钟。每节课都很精短且互动,所以你能稳步进步,并在网页和应用中从离开的地方继续。

我能在这节 Data Science Academy 课中编写并运行代码吗?

能。每节 Data Science Academy 课都包含内置代码编辑器,你可以在浏览器中直接编写并运行真实代码,并获得即时 AI 反馈 — 无需本地设置。

此课程中的所有课时

  1. 找出表格中隐藏的 NaNs
  2. 删除还是填充:明智选择
  3. 使用均值、中位数或众数填补
  4. 修正 dtypes 并删除重复行
← 返回 Data Science Academy