0Pricing
Machine Learning Academy · Pelajaran

Pandas untuk Manipulasi Data

Peserta didik akan memuat file CSV ke dalam DataFrames, memfilter baris, memilih kolom, menangani nilai yang hilang, dan menghitung statistik ringkasan.

Pandas untuk Manipulasi Data adalah pelajaran Machine Learning Academy gratis di CoddyKit. Ini adalah pelajaran 3 dari 4. Kamu bisa membaca pelajaran lengkapnya di bawah secara gratis — lalu praktikkan langsung di browser dengan editor kode bawaan dan tutor AI 24/7. Ini adalah bagian dari jalur belajar Machine Learning Academy, dan progresmu tersinkronisasi di web dan aplikasi CoddyKit. Kursus Machine Learning Academy mencakup 4 pelajaran total.

Bagian dari pelajaran ini belum diterjemahkan dan ditampilkan dalam bahasa Inggris.

What Is Pandas and Why Use It?

Pandas handles tabular data — rows and columns, like a spreadsheet. Its DataFrame is where you load and clean raw data before it ever reaches your model.

import pandas as pd
import numpy as np

# Create a DataFrame manually
df = pd.DataFrame({
    'name': ['Alice', 'Bob', 'Carol', 'Dave'],
    'age': [25, 30, 35, 28],
    'salary': [50000, 70000, 90000, 60000],
    'department': ['Engineering', 'Marketing', 'Engineering', 'Sales']
})
print(df)
print('\nShape:', df.shape)  # (4, 4)

Loading Data from CSV Files

Load a file in one line with pd.read_csv(). Then always inspect it: .head(), .info(), and .describe() reveal shape, types, and missing values in seconds.

import pandas as pd

# Load a CSV
df = pd.read_csv('titanic.csv')

# First inspection
print(df.head())       # first 5 rows
print(df.tail(3))      # last 3 rows
print(df.info())       # column types + non-null counts
print(df.describe())   # count, mean, std, min, max for numeric cols
print(df.columns.tolist())  # column names
print(df.shape)        # (891, 12)

Selecting Columns and Rows

Pick columns with df['col'], and filter rows with boolean conditions. Use .loc[] for labels and .iloc[] for positions. Wrap each condition in parentheses!

import pandas as pd

df = pd.read_csv('titanic.csv')

# Select columns
age = df['Age']                          # Series
subset = df[['Age', 'Fare', 'Survived']] # DataFrame

# Filter rows with boolean conditions
survivors = df[df['Survived'] == 1]
first_class_women = df[(df['Pclass'] == 1) & (df['Sex'] == 'female')]

# Label-based selection (row label, column label)
print(df.loc[0, 'Name'])  # first row, Name column

# Position-based selection
print(df.iloc[0, 0])      # first row, first column

Handling Missing Values

Real data has gaps, shown as NaN, and most models choke on them. Your two moves: drop the rows or columns, or fill them with a value like the median.

import pandas as pd
import numpy as np

df = pd.read_csv('titanic.csv')

# Detect missing values
print(df.isnull().sum())  # count NaN per column
print(df.isnull().mean() * 100)  # % missing per column

# Drop columns with >50% missing
df_clean = df.dropna(thresh=len(df) * 0.5, axis=1)

# Fill numeric missing with median
df['Age'] = df['Age'].fillna(df['Age'].median())

# Fill categorical missing with most frequent
df['Embarked'] = df['Embarked'].fillna(df['Embarked'].mode()[0])

Data Types and Type Conversion

Every column has a dtype like int, float, or object (text). A wrong one causes silent bugs — numbers read as text can't do math. Use .astype() to convert.

import pandas as pd

df = pd.read_csv('titanic.csv')
print(df.dtypes)  # see all column dtypes

# Convert a column's dtype
df['Survived'] = df['Survived'].astype(bool)
df['Pclass'] = df['Pclass'].astype('category')

# Convert object column to numeric (coerce errors to NaN)
df['Fare'] = pd.to_numeric(df['Fare'], errors='coerce')

# Parse dates
# df['date'] = pd.to_datetime(df['date'])

print(df.dtypes)

Adding and Transforming Columns

Feature engineering means building new columns from existing ones. Combining SibSp and Parch into one FamilySize often helps a model learn better. See the code.

import pandas as pd

df = pd.read_csv('titanic.csv')

# Create a new column from existing ones
df['FamilySize'] = df['SibSp'] + df['Parch'] + 1  # +1 for self

# Binary flag: is the passenger alone?
df['IsAlone'] = (df['FamilySize'] == 1).astype(int)

# Bin a continuous feature into categories
df['AgeGroup'] = pd.cut(df['Age'], bins=[0, 12, 18, 60, 100],
                         labels=['Child', 'Teen', 'Adult', 'Senior'])

print(df[['FamilySize', 'IsAlone', 'AgeGroup']].head())

GroupBy: Aggregating by Category

groupby() splits your data by a category, applies a function like mean, and combines the results — just like SQL's GROUP BY. Perfect for spotting patterns fast.

import pandas as pd

df = pd.read_csv('titanic.csv')

# Mean survival rate by class
survival_by_class = df.groupby('Pclass')['Survived'].mean()
print(survival_by_class)
# Pclass 1: ~0.63, Pclass 2: ~0.47, Pclass 3: ~0.24

# Multiple aggregations at once
summary = df.groupby('Pclass').agg(
    passengers=('Survived', 'count'),
    survival_rate=('Survived', 'mean'),
    avg_fare=('Fare', 'mean')
)
print(summary)

Merging and Joining DataFrames

Need data from two sources? pd.merge() joins DataFrames on a shared key, just like a SQL join. Use pd.concat() to stack tables into more rows or columns.

import pandas as pd

customers = pd.DataFrame({'id': [1, 2, 3], 'name': ['Alice', 'Bob', 'Carol']})
orders = pd.DataFrame({'id': [1, 1, 2], 'amount': [50, 30, 70]})

# Inner join on 'id'
merged = pd.merge(customers, orders, on='id', how='inner')
print(merged)

# Stack DataFrames with the same columns
df1 = pd.DataFrame({'A': [1, 2]})
df2 = pd.DataFrame({'A': [3, 4]})
combined = pd.concat([df1, df2], ignore_index=True)
print(combined)

Summary Statistics and Value Counts

Before modeling, study your data: .describe() summarizes numbers and .value_counts() counts categories. Always check for class imbalance — it can fool accuracy.

import pandas as pd

df = pd.read_csv('titanic.csv')

# Numeric statistics
print(df['Age'].describe())
# count, mean, std, min, 25%, 50%, 75%, max

# Categorical distribution
print(df['Sex'].value_counts())
print(df['Pclass'].value_counts(normalize=True))  # proportions

# Correlation with target variable
correlations = df.corr()['Survived'].sort_values(ascending=False)
print(correlations)

Converting a DataFrame to NumPy for scikit-learn

The last step before training: split features (X) from the target (y), then convert to NumPy with .to_numpy(). Now scikit-learn has exactly what it expects.

import pandas as pd
from sklearn.model_selection import train_test_split

df = pd.read_csv('titanic.csv')

# Select numeric features only for simplicity
features = ['Pclass', 'Age', 'SibSp', 'Parch', 'Fare']
df_model = df[features + ['Survived']].dropna()

X = df_model[features].to_numpy()       # shape (n, 5)
y = df_model['Survived'].to_numpy()     # shape (n,)

X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.2, random_state=42, stratify=y
)
print('Train shape:', X_train.shape)
print('Test shape:', X_test.shape)

Saving and Loading Processed Data

After all that cleaning, save your work so you never redo it. df.to_csv() is simple; Parquet is far smaller and faster for big datasets.

import pandas as pd

df = pd.read_csv('titanic.csv')

# Save as CSV
df.to_csv('titanic_processed.csv', index=False)

# Save as Parquet (much faster for large files)
# df.to_parquet('titanic_processed.parquet', index=False)

# Load back
df_reloaded = pd.read_csv('titanic_processed.csv')
print('Reloaded shape:', df_reloaded.shape)

Quick Check

Test your understanding of Machine Learning with Python concepts from this lesson.

Lesson Recap

You learned to wrangle data with Pandas: DataFrames load and clean tables, dropna and fillna handle missing values, and groupby reveals patterns. Next: charts. 📊

Pertanyaan yang Sering Diajukan

Apakah pelajaran “Pandas untuk Manipulasi Data” gratis?

Ya — teks lengkap “Pandas untuk Manipulasi Data” gratis dibaca di sini di web. Untuk praktiknya secara interaktif (editor kode bawaan dan tutor AI 24/7) dan buka sisa kursus Machine Learning Academy, upgrade ke CoddyKit PRO. Kursus Machine Learning Academy mencakup 4 pelajaran total.

Apa yang akan aku pelajari di “Pandas untuk Manipulasi Data”?

Peserta didik akan memuat file CSV ke dalam DataFrames, memfilter baris, memilih kolom, menangani nilai yang hilang, dan menghitung statistik ringkasan. Kamu berlatih Machine Learning Academy dengan kode praktik yang langsung kamu jalankan di browser, dan tutor AI 24/7 menjawab pertanyaanmu saat kamu mengerjakan pelajaran ini.

Apakah aku perlu pengalaman untuk memulai Machine Learning Academy?

Tidak diperlukan pengalaman sebelumnya. Machine Learning Academy di CoddyKit dirancang untuk pemula hingga pelajar tingkat lanjut, jadi kamu bisa memulai di sini atau dari awal dan belajar sesuai kecepatan kamu sendiri. Ini adalah pelajaran 3 dari 4.

Berapa lama pelajaran “Pandas untuk Manipulasi Data” memakan waktu?

Sebagian besar pelajaran CoddyKit memakan waktu sekitar 5–10 menit. Setiap pelajaran ringkas dan interaktif, jadi kamu membuat kemajuan stabil dan melanjutkan dari tempat kamu tinggalkan di web dan aplikasi.

Bisakah aku menulis dan menjalankan kode dalam pelajaran Machine Learning Academy ini?

Ya. Setiap pelajaran Machine Learning Academy menyertakan editor kode bawaan, jadi kamu menulis dan menjalankan kode nyata langsung di browser dan mendapatkan umpan balik AI instan — tidak diperlukan penyiapan lokal.

Semua pelajaran dalam kursus ini

  1. Memasang Anaconda dan Jupyter Notebook
  2. Dasar-Dasar NumPy: Larik dan Operasi Matematika
  3. Pandas untuk Manipulasi Data
  4. Memvisualisasikan Data dengan Matplotlib dan Seaborn
← Kembali ke Machine Learning Academy