Pandas untuk Manipulasi Data
Peserta didik akan memuat file CSV ke dalam DataFrames, memfilter baris, memilih kolom, menangani nilai yang hilang, dan menghitung statistik ringkasan.
Pandas untuk Manipulasi Data adalah pelajaran Machine Learning Academy gratis di CoddyKit. Ini adalah pelajaran 3 dari 4. Kamu bisa membaca pelajaran lengkapnya di bawah secara gratis — lalu praktikkan langsung di browser dengan editor kode bawaan dan tutor AI 24/7. Ini adalah bagian dari jalur belajar Machine Learning Academy, dan progresmu tersinkronisasi di web dan aplikasi CoddyKit. Kursus Machine Learning Academy mencakup 4 pelajaran total.
Bagian dari pelajaran ini belum diterjemahkan dan ditampilkan dalam bahasa Inggris.
What Is Pandas and Why Use It?
Pandas handles tabular data — rows and columns, like a spreadsheet. Its DataFrame is where you load and clean raw data before it ever reaches your model.
import pandas as pd
import numpy as np
# Create a DataFrame manually
df = pd.DataFrame({
'name': ['Alice', 'Bob', 'Carol', 'Dave'],
'age': [25, 30, 35, 28],
'salary': [50000, 70000, 90000, 60000],
'department': ['Engineering', 'Marketing', 'Engineering', 'Sales']
})
print(df)
print('\nShape:', df.shape) # (4, 4)Loading Data from CSV Files
Load a file in one line with pd.read_csv(). Then always inspect it: .head(), .info(), and .describe() reveal shape, types, and missing values in seconds.
import pandas as pd
# Load a CSV
df = pd.read_csv('titanic.csv')
# First inspection
print(df.head()) # first 5 rows
print(df.tail(3)) # last 3 rows
print(df.info()) # column types + non-null counts
print(df.describe()) # count, mean, std, min, max for numeric cols
print(df.columns.tolist()) # column names
print(df.shape) # (891, 12)Selecting Columns and Rows
Pick columns with df['col'], and filter rows with boolean conditions. Use .loc[] for labels and .iloc[] for positions. Wrap each condition in parentheses!
import pandas as pd
df = pd.read_csv('titanic.csv')
# Select columns
age = df['Age'] # Series
subset = df[['Age', 'Fare', 'Survived']] # DataFrame
# Filter rows with boolean conditions
survivors = df[df['Survived'] == 1]
first_class_women = df[(df['Pclass'] == 1) & (df['Sex'] == 'female')]
# Label-based selection (row label, column label)
print(df.loc[0, 'Name']) # first row, Name column
# Position-based selection
print(df.iloc[0, 0]) # first row, first columnHandling Missing Values
Real data has gaps, shown as NaN, and most models choke on them. Your two moves: drop the rows or columns, or fill them with a value like the median.
import pandas as pd
import numpy as np
df = pd.read_csv('titanic.csv')
# Detect missing values
print(df.isnull().sum()) # count NaN per column
print(df.isnull().mean() * 100) # % missing per column
# Drop columns with >50% missing
df_clean = df.dropna(thresh=len(df) * 0.5, axis=1)
# Fill numeric missing with median
df['Age'] = df['Age'].fillna(df['Age'].median())
# Fill categorical missing with most frequent
df['Embarked'] = df['Embarked'].fillna(df['Embarked'].mode()[0])Data Types and Type Conversion
Every column has a dtype like int, float, or object (text). A wrong one causes silent bugs — numbers read as text can't do math. Use .astype() to convert.
import pandas as pd
df = pd.read_csv('titanic.csv')
print(df.dtypes) # see all column dtypes
# Convert a column's dtype
df['Survived'] = df['Survived'].astype(bool)
df['Pclass'] = df['Pclass'].astype('category')
# Convert object column to numeric (coerce errors to NaN)
df['Fare'] = pd.to_numeric(df['Fare'], errors='coerce')
# Parse dates
# df['date'] = pd.to_datetime(df['date'])
print(df.dtypes)Adding and Transforming Columns
Feature engineering means building new columns from existing ones. Combining SibSp and Parch into one FamilySize often helps a model learn better. See the code.
import pandas as pd
df = pd.read_csv('titanic.csv')
# Create a new column from existing ones
df['FamilySize'] = df['SibSp'] + df['Parch'] + 1 # +1 for self
# Binary flag: is the passenger alone?
df['IsAlone'] = (df['FamilySize'] == 1).astype(int)
# Bin a continuous feature into categories
df['AgeGroup'] = pd.cut(df['Age'], bins=[0, 12, 18, 60, 100],
labels=['Child', 'Teen', 'Adult', 'Senior'])
print(df[['FamilySize', 'IsAlone', 'AgeGroup']].head())GroupBy: Aggregating by Category
groupby() splits your data by a category, applies a function like mean, and combines the results — just like SQL's GROUP BY. Perfect for spotting patterns fast.
import pandas as pd
df = pd.read_csv('titanic.csv')
# Mean survival rate by class
survival_by_class = df.groupby('Pclass')['Survived'].mean()
print(survival_by_class)
# Pclass 1: ~0.63, Pclass 2: ~0.47, Pclass 3: ~0.24
# Multiple aggregations at once
summary = df.groupby('Pclass').agg(
passengers=('Survived', 'count'),
survival_rate=('Survived', 'mean'),
avg_fare=('Fare', 'mean')
)
print(summary)Merging and Joining DataFrames
Need data from two sources? pd.merge() joins DataFrames on a shared key, just like a SQL join. Use pd.concat() to stack tables into more rows or columns.
import pandas as pd
customers = pd.DataFrame({'id': [1, 2, 3], 'name': ['Alice', 'Bob', 'Carol']})
orders = pd.DataFrame({'id': [1, 1, 2], 'amount': [50, 30, 70]})
# Inner join on 'id'
merged = pd.merge(customers, orders, on='id', how='inner')
print(merged)
# Stack DataFrames with the same columns
df1 = pd.DataFrame({'A': [1, 2]})
df2 = pd.DataFrame({'A': [3, 4]})
combined = pd.concat([df1, df2], ignore_index=True)
print(combined)Summary Statistics and Value Counts
Before modeling, study your data: .describe() summarizes numbers and .value_counts() counts categories. Always check for class imbalance — it can fool accuracy.
import pandas as pd
df = pd.read_csv('titanic.csv')
# Numeric statistics
print(df['Age'].describe())
# count, mean, std, min, 25%, 50%, 75%, max
# Categorical distribution
print(df['Sex'].value_counts())
print(df['Pclass'].value_counts(normalize=True)) # proportions
# Correlation with target variable
correlations = df.corr()['Survived'].sort_values(ascending=False)
print(correlations)Converting a DataFrame to NumPy for scikit-learn
The last step before training: split features (X) from the target (y), then convert to NumPy with .to_numpy(). Now scikit-learn has exactly what it expects.
import pandas as pd
from sklearn.model_selection import train_test_split
df = pd.read_csv('titanic.csv')
# Select numeric features only for simplicity
features = ['Pclass', 'Age', 'SibSp', 'Parch', 'Fare']
df_model = df[features + ['Survived']].dropna()
X = df_model[features].to_numpy() # shape (n, 5)
y = df_model['Survived'].to_numpy() # shape (n,)
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.2, random_state=42, stratify=y
)
print('Train shape:', X_train.shape)
print('Test shape:', X_test.shape)Saving and Loading Processed Data
After all that cleaning, save your work so you never redo it. df.to_csv() is simple; Parquet is far smaller and faster for big datasets.
import pandas as pd
df = pd.read_csv('titanic.csv')
# Save as CSV
df.to_csv('titanic_processed.csv', index=False)
# Save as Parquet (much faster for large files)
# df.to_parquet('titanic_processed.parquet', index=False)
# Load back
df_reloaded = pd.read_csv('titanic_processed.csv')
print('Reloaded shape:', df_reloaded.shape)Quick Check
Test your understanding of Machine Learning with Python concepts from this lesson.
Lesson Recap
You learned to wrangle data with Pandas: DataFrames load and clean tables, dropna and fillna handle missing values, and groupby reveals patterns. Next: charts. 📊
Pertanyaan yang Sering Diajukan
Apakah pelajaran “Pandas untuk Manipulasi Data” gratis?
Ya — teks lengkap “Pandas untuk Manipulasi Data” gratis dibaca di sini di web. Untuk praktiknya secara interaktif (editor kode bawaan dan tutor AI 24/7) dan buka sisa kursus Machine Learning Academy, upgrade ke CoddyKit PRO. Kursus Machine Learning Academy mencakup 4 pelajaran total.
Apa yang akan aku pelajari di “Pandas untuk Manipulasi Data”?
Peserta didik akan memuat file CSV ke dalam DataFrames, memfilter baris, memilih kolom, menangani nilai yang hilang, dan menghitung statistik ringkasan. Kamu berlatih Machine Learning Academy dengan kode praktik yang langsung kamu jalankan di browser, dan tutor AI 24/7 menjawab pertanyaanmu saat kamu mengerjakan pelajaran ini.
Apakah aku perlu pengalaman untuk memulai Machine Learning Academy?
Tidak diperlukan pengalaman sebelumnya. Machine Learning Academy di CoddyKit dirancang untuk pemula hingga pelajar tingkat lanjut, jadi kamu bisa memulai di sini atau dari awal dan belajar sesuai kecepatan kamu sendiri. Ini adalah pelajaran 3 dari 4.
Berapa lama pelajaran “Pandas untuk Manipulasi Data” memakan waktu?
Sebagian besar pelajaran CoddyKit memakan waktu sekitar 5–10 menit. Setiap pelajaran ringkas dan interaktif, jadi kamu membuat kemajuan stabil dan melanjutkan dari tempat kamu tinggalkan di web dan aplikasi.
Bisakah aku menulis dan menjalankan kode dalam pelajaran Machine Learning Academy ini?
Ya. Setiap pelajaran Machine Learning Academy menyertakan editor kode bawaan, jadi kamu menulis dan menjalankan kode nyata langsung di browser dan mendapatkan umpan balik AI instan — tidak diperlukan penyiapan lokal.
Semua pelajaran dalam kursus ini
- Memasang Anaconda dan Jupyter Notebook
- Dasar-Dasar NumPy: Larik dan Operasi Matematika
- Pandas untuk Manipulasi Data
- Memvisualisasikan Data dengan Matplotlib dan Seaborn