修正 dtypes 并删除重复行
转换类型并删除重复项
修正 dtypes 并删除重复行 是 CoddyKit 上的免费 Data Science Academy 课时。 这是第 4 节课,共 4 节。 你可以在下方免费阅读本课时的完整内容 — 然后在浏览器中使用内置代码编辑器和全天候 AI 导师进行实践。 这是 Data Science Academy 学习路径的一部分,你的进度在网页和 CoddyKit 应用中同步。 Data Science Academy 课程共包含 4 节课。
本课时的部分内容尚未翻译,以英文显示。
Clean Goes Beyond Gaps
Filling missing values is only half the job. Clean data also needs correct dtypes and no accidental duplicate rows muddying your counts. 🧹
What a dtype Is
Every column has a dtype, its data type, like int, float, object, or datetime. The dtype controls which operations the column even allows.
df.dtypesNumbers Stuck as Text
A frequent mess: numbers loaded as object strings because one stray symbol crept in. You cannot do math on them until the type is fixed.
Coerce With to_numeric
Convert a text column to numbers with to_numeric. Pass errors='coerce' to turn anything unparseable into NaN instead of crashing.
df['price'] = pd.to_numeric(df['price'], errors='coerce')Cast With astype
When a column is already clean, astype changes its type directly, for example turning whole-number floats back into compact integers.
df['count'] = df['count'].astype(int)Fix Date Columns
Dates often arrive as text. Convert them with to_datetime so you can sort, filter by range, and extract parts like month later.
df['date'] = pd.to_datetime(df['date'])Category to Save Memory
Columns with few repeated text values shrink dramatically as the category dtype, which stores each label once and references it.
df['city'] = df['city'].astype('category')Spotting Duplicate Rows
Repeated records inflate totals and averages. Find them with duplicated(), which flags each row that has appeared before as True.
df.duplicated().sum()Dropping Duplicates
Remove repeats with drop_duplicates(). By default it keeps the first occurrence of each row and discards the rest.
df = df.drop_duplicates()Duplicates by Key Columns
Sometimes only certain columns define a duplicate. Pass subset to detect repeats based on a key like an id, ignoring the other columns.
df.drop_duplicates(subset=['user_id'])Verify Before You Move On
After fixing types and repeats, run info() once more. Confirming dtypes and row counts protects you from cleaning errors slipping downstream.
df.info()Quick Check
A price column loaded as object text. Which call safely turns bad entries into NaN?
Recap: Types and Dupes Tamed
You now fix dtypes with to_numeric, astype, and to_datetime, and clear repeats with drop_duplicates. Your table is finally analysis-ready.
常见问题解答
「修正 dtypes 并删除重复行」课时是免费的吗?
是的 — 「修正 dtypes 并删除重复行」的完整文本可在网页上免费阅读。要进行交互式练习(内置代码编辑器和全天候 AI 导师)并解锁 Data Science Academy 课程的其余内容,请升级到 CoddyKit PRO。 Data Science Academy 课程共包含 4 节课。
「修正 dtypes 并删除重复行」这节课中我会学到什么?
转换类型并删除重复项 你通过在浏览器中直接运行的动手代码来练习 Data Science Academy,全天候 AI 导师会在你学习这节课的过程中回答你的问题。
学习 Data Science Academy 需要有经验吗?
无需任何先前经验。CoddyKit 上的 Data Science Academy 课程适合初学者到高级学习者,你可以从这里开始或从头开始,按照自己的节奏学习。 这是第 4 节课,共 4 节。
「修正 dtypes 并删除重复行」课时需要多长时间?
大多数 CoddyKit 课程大约需要 5–10 分钟。每节课都很精短且互动,所以你能稳步进步,并在网页和应用中从离开的地方继续。
我能在这节 Data Science Academy 课中编写并运行代码吗?
能。每节 Data Science Academy 课都包含内置代码编辑器,你可以在浏览器中直接编写并运行真实代码,并获得即时 AI 反馈 — 无需本地设置。
此课程中的所有课时
- 找出表格中隐藏的 NaNs
- 删除还是填充:明智选择
- 使用均值、中位数或众数填补
- 修正 dtypes 并删除重复行