แทนค่าด้วยค่าเฉลี่ย มัธยฐาน หรือฐานนิยม
เติมค่าอย่างเหมาะสมตามชนิดคอลัมน์
แทนค่าด้วยค่าเฉลี่ย มัธยฐาน หรือฐานนิยม เป็นบทเรียน Data Science Academy ฟรีบน CoddyKit นี่คือบทเรียนที่ 3 จากทั้งหมด 4 บทเรียน คุณสามารถอ่านบทเรียนทั้งหมดด้านล่างฟรี — จากนั้นลองปฏิบัติด้วยตัวคุณเองในเบราว์เซอร์พร้อมตัวแก้ไขโค้ดในตัวและติวเตอร์ AI ตลอด 24/7 บทเรียนนี้เป็นส่วนหนึ่งของเส้นทางการเรียน Data Science Academy และความก้าวหน้าของคุณจะซิงค์ข้ามเว็บและแอป CoddyKit คอร์ส Data Science Academy มีบทเรียนทั้งหมด 4 บทเรียน
บางส่วนของบทเรียนนี้ยังไม่ได้รับการแปล และแสดงเป็นภาษาอังกฤษ
Filling With a Statistic
Instead of inventing random numbers, you fill gaps with a summary value from the column itself. This careful approach is called imputation. 🧮
The Mean Fill
For numbers, the simplest choice is the mean, the column average. It keeps the overall total roughly intact when gaps are few.
df['age'].fillna(df['age'].mean())When the Mean Misleads
The mean is fragile: a few huge values drag it far from the center. On skewed data, mean imputation can pull every gap toward an unrealistic number.
The Median Fill
The median is the middle value, and it shrugs off extreme outliers. For skewed numeric columns it is usually the safer fill.
df['income'].fillna(df['income'].median())Mean or Median?
A quick rule: reach for the median when a column has outliers or a long tail, and the mean only when values are fairly symmetric.
The Mode for Categories
Text and category columns have no average. For them you fill with the mode, the most frequent value, since that is the most likely fit.
top = df['city'].mode()[0]
df['city'].fillna(top)Why mode Returns a Series
A column can tie for most common, so mode() returns a Series of all winners. Pick the first with index 0 when you need one value.
df['city'].mode()[0]Fill Per Column Type
The best practice is to impute each column by its type: median for skewed numbers, mean for symmetric ones, and mode for categories.
Forward and Back Fill
For ordered data like time series, carry the last known value forward with ffill, or pull the next one backward with bfill.
df['temp'].ffill()The Hidden Cost
Every imputation shrinks the column's natural spread, because filled values cluster at one point. Note this, since it nudges your variance downward.
Flag What You Filled
A pro habit: add a boolean column marking which rows were imputed. That flag lets later analysis know which values were real and which were guessed.
df['age_filled'] = df['age'].isna()Quick Check
An income column is heavily skewed by a few millionaires. Which imputation fits best?
Recap: Smart Fills
You now impute with mean, median, or mode by column type, use ffill for ordered data, and flag filled rows so nothing gets hidden.
คำถามที่พบบ่อย
บทเรียน “แทนค่าด้วยค่าเฉลี่ย มัธยฐาน หรือฐานนิยม” ฟรีหรือไม่
ใช่ — ข้อความเต็มของ “แทนค่าด้วยค่าเฉลี่ย มัธยฐาน หรือฐานนิยม” ฟรีให้อ่านที่นี่บนเว็บ เพื่อปฏิบัติแบบโต้ตอบ (ตัวแก้ไขโค้ดในตัวและติวเตอร์ AI ตลอด 24/7) และปลดล็อคส่วนที่เหลือของคอร์ส Data Science Academy ให้อัปเกรดเป็น CoddyKit PRO คอร์ส Data Science Academy มีบทเรียนทั้งหมด 4 บทเรียน
คุณจะเรียนรู้อะไรในบทเรียน “แทนค่าด้วยค่าเฉลี่ย มัธยฐาน หรือฐานนิยม”
เติมค่าอย่างเหมาะสมตามชนิดคอลัมน์ คุณปฏิบัติ Data Science Academy ด้วยโค้ดที่ใช้งานได้จริงที่คุณเรียกใช้โดยตรงในเบราว์เซอร์ และติวเตอร์ AI ตลอด 24/7 ตอบคำถามของคุณขณะที่คุณไปผ่านบทเรียน
คุณต้องมีประสบการณ์ก่อนที่จะเริ่มเรียน Data Science Academy หรือไม่
ไม่จำเป็นต้องมีประสบการณ์มาก่อน Data Science Academy บน CoddyKit ออกแบบมาสำหรับผู้เริ่มต้นไปจนถึงผู้เรียนขั้นสูง คุณสามารถเริ่มต้นที่นี่หรือเริ่มจากตัวแรกและเรียนด้วยความเร็วของคุณเอง นี่คือบทเรียนที่ 3 จากทั้งหมด 4 บทเรียน
บทเรียน “แทนค่าด้วยค่าเฉลี่ย มัธยฐาน หรือฐานนิยม” ใช้เวลานานแค่ไหน
บทเรียน CoddyKit ส่วนใหญ่ใช้เวลาประมาณ 5–10 นาที แต่ละบทเรียนจึงสั้นและเป็นแบบโต้ตอบ คุณสามารถก้าวหน้าอย่างต่อเนื่องและกลับมาเรียนต่อจากตรงที่เพิ่งหยุดบนเว็บและแอปได้เลย
ฉันเขียนและรันโค้ดในบทเรียน Data Science Academy นี้ได้ไหม
ได้ บทเรียน Data Science Academy ทุกบทมีตัวแก้ไขโค้ดในตัว คุณจึงเขียนและรันโค้ดจริงได้เลยในเบราว์เซอร์ และได้รับข้อเสนอแนะจาก AI ในทันที — ไม่ต้องติดตั้งในเครื่องของคุณ
บทเรียนทั้งหมดในหลักสูตรนี้
- ค้นหา NaNs ที่ซ่อนอยู่ในตาราง
- ลบหรือเติม: เลือกให้เหมาะสม
- แทนค่าด้วยค่าเฉลี่ย มัธยฐาน หรือฐานนิยม
- แก้ไข dtypes และแถวซ้ำ