カテゴリ型
カーディナリティの低い文字列列をpandasのCategoricalに変換してメモリ使用量を減らし、groupby操作を高速化します。
「カテゴリ型」はCoddyKit上の無料Pandas & NumPy Academyレッスンです。 これはレッスン3/4です。 下記で完全なレッスンを無料で読むことができます。その後、ブラウザ内の組み込みコードエディタと24時間対応のAIチューターでハンズオン演習できます。 これはPandas & NumPy Academy学習パスの一部であり、ウェブとCoddyKitアプリ全体で進捗が同期されます。 Pandas & NumPy Academyコースには全4レッスンが含まれています。
このレッスンの一部はまだ翻訳されておらず、英語で表示されています。
What Is the Categorical Dtype?
The Categorical dtype in Pandas is designed for columns that contain a limited set of discrete values (low cardinality), such as gender, status, region, or product category. Instead of storing the full string for every row, Pandas stores the unique values (called categories) once and uses an integer code per row to reference them. This is similar to how databases use lookup tables or enum types.
import pandas as pd
df = pd.DataFrame({
'region': ['North', 'South', 'East', 'North', 'East', 'South'] * 1000
})
# Without Categorical: object dtype stores every string
print('object dtype memory:', df['region'].memory_usage(deep=True))
# With Categorical: only stores 3 unique values + integer codes
df['region_cat'] = df['region'].astype('category')
print('category dtype memory:', df['region_cat'].memory_usage(deep=True))Creating a Categorical Series
Convert a column to Categorical by calling .astype('category') on any Series. The resulting Categorical Series stores the unique values as .cat.categories and the integer position codes as .cat.codes. You can also create a Categorical directly with pd.Categorical() to specify the categories and their order up front.
import pandas as pd
s = pd.Series(['low', 'high', 'med', 'high', 'low'])
s_cat = s.astype('category')
print(s_cat.cat.categories) # Index(['high', 'low', 'med'], dtype='object')
print(s_cat.cat.codes) # integer codes per row
# 0 1
# 1 0
# 2 2
# 3 0
# 4 1Memory Savings with Categorical
The memory savings from Categorical dtype depend on the cardinality ratio (unique values ÷ total rows). With 5 unique strings in a 1,000,000-row column, the object dtype stores 1 million string pointers (~50+ bytes each) while Categorical stores just 5 strings plus 1 million 8-bit integers. The savings can be 5-20x, turning a 50 MB column into 2-5 MB.
import pandas as pd
import numpy as np
n = 1_000_000
statuses = np.random.choice(['active', 'inactive', 'pending'], size=n)
df = pd.DataFrame({'status': statuses})
object_mem = df['status'].memory_usage(deep=True) / 1e6
df['status'] = df['status'].astype('category')
cat_mem = df['status'].memory_usage(deep=True) / 1e6
print(f'Object dtype: {object_mem:.1f} MB')
print(f'Category dtype: {cat_mem:.1f} MB')
print(f'Reduction: {object_mem/cat_mem:.1f}x')Ordered Categorical for Ranking
Sometimes categories have a natural order — e.g., 'low' < 'medium' < 'high', or T-shirt sizes XS < S < M < L < XL. An ordered Categorical captures this ranking, enabling meaningful comparison operators (<, >=, etc.) and ensuring groupby results are sorted in the correct semantic order rather than alphabetically.
import pandas as pd
df = pd.DataFrame({
'priority': ['high', 'low', 'medium', 'high', 'low']
})
priority_type = pd.CategoricalDtype(
categories=['low', 'medium', 'high'],
ordered=True
)
df['priority'] = df['priority'].astype(priority_type)
# Comparison now works correctly
print(df['priority'] >= 'medium')
# 0 True
# 1 False
# 2 True
# 3 True
# 4 FalseSorting Ordered Categoricals
When you sort an ordered Categorical column, Pandas uses the defined category order rather than alphabetical order. This means 'low' comes before 'medium' before 'high' regardless of how they would sort as strings. This is critical for producing correctly ordered summary tables and charts.
import pandas as pd
df = pd.DataFrame({
'severity': pd.Categorical(
['high', 'low', 'medium', 'low', 'high'],
categories=['low', 'medium', 'high'],
ordered=True
),
'count': [5, 20, 8, 15, 3]
})
# sort_values respects the categorical order
print(df.sort_values('severity')[['severity', 'count']])
# severity count
# 1 low 20
# 3 low 15
# 2 medium 8
# 0 high 5
# 4 high 3GroupBy Speed with Categorical
Pandas can optimise groupby() operations on Categorical columns because it knows all possible groups in advance. This can make groupby 2-5x faster on large DataFrames. Additionally, groupby on Categorical returns all categories — including empty ones — which ensures your summary tables always have a row for every expected group, even if some have zero observations.
import pandas as pd
import numpy as np
np.random.seed(0)
df = pd.DataFrame({
'region': pd.Categorical(
np.random.choice(['North', 'South', 'East', 'West'], 10),
categories=['North', 'South', 'East', 'West']
),
'sales': np.random.randint(100, 1000, 10)
})
# observed=False shows ALL categories even empty ones
print(df.groupby('region', observed=False)['sales'].sum())Adding and Removing Categories
Use the .cat accessor to manage the category list. .cat.add_categories() adds new allowed values (useful before inserting new data), and .cat.remove_unused_categories() drops category labels that have no corresponding rows — handy after filtering. You cannot assign a value that is not in the categories without adding it first.
import pandas as pd
s = pd.Categorical(['a', 'b', 'a'], categories=['a', 'b', 'c'])
s = pd.Series(s)
print(s.cat.categories) # Index(['a', 'b', 'c'], dtype='object')
print(s.value_counts()) # a:2, b:1, c:0
# Remove unused category 'c'
s = s.cat.remove_unused_categories()
print(s.cat.categories) # Index(['a', 'b'], dtype='object')Renaming Category Labels
.cat.rename_categories() lets you relabel categories without changing the underlying codes. This is useful when you want to display friendlier names (e.g., 'M' → 'Male') or fix inconsistent label capitalisation without converting back to object and remapping. The method accepts a list (positional) or a dictionary (targeted).
import pandas as pd
s = pd.Series(pd.Categorical(['M', 'F', 'M', 'F'], categories=['M', 'F']))
# Rename using a dict
s = s.cat.rename_categories({'M': 'Male', 'F': 'Female'})
print(s)
# 0 Male
# 1 Female
# 2 Male
# 3 Female
print(s.cat.categories) # Index(['Male', 'Female'], dtype='object')When NOT to Use Categorical
Categorical dtype is a poor choice when cardinality is high — if almost every row has a unique value (like user IDs, free-text comments, or UUIDs), the category index is nearly as large as the original data, giving no memory saving. A rule of thumb: use Categorical only when the number of unique values is less than roughly 50% of the total rows, and especially when unique values number in the tens or hundreds.
import pandas as pd
import numpy as np
df = pd.DataFrame({'user_id': range(1_000_000)})
# High-cardinality: every value is unique — don't use Categorical
object_mem = df['user_id'].astype(str).memory_usage(deep=True) / 1e6
cat_mem = df['user_id'].astype(str).astype('category').memory_usage(deep=True) / 1e6
print(f'String: {object_mem:.1f} MB')
print(f'Category: {cat_mem:.1f} MB')
# Category is WORSE for high-cardinality data!Categorical in Pivot Tables and Groupby
Using Categorical ensures consistent output shape in groupby and pivot table results. Without Categorical, groups that happen to have zero observations are silently omitted from the result — which can cause misaligned output when comparing across different data slices. With Categorical and observed=False, all groups always appear in the output.
import pandas as pd
df = pd.DataFrame({
'quarter': pd.Categorical(
['Q1', 'Q1', 'Q3'],
categories=['Q1', 'Q2', 'Q3', 'Q4']
),
'revenue': [100, 200, 300]
})
# observed=False includes Q2 and Q4 even though they have no rows
result = df.groupby('quarter', observed=False)['revenue'].sum()
print(result)
# quarter
# Q1 300
# Q2 0
# Q3 300
# Q4 0Categorical and Machine Learning
Many scikit-learn estimators expect numeric input. To convert a Categorical column to numbers for a model, use .cat.codes (label encoding) or pd.get_dummies() (one-hot encoding). One-hot encoding avoids implying an ordinal relationship between categories. For ordered categoricals like priority levels, label encoding (the codes directly) is appropriate and carries the ordering information.
import pandas as pd
df = pd.DataFrame({
'colour': pd.Categorical(['red', 'blue', 'green', 'red'])
})
# One-hot encode for unordered categories
one_hot = pd.get_dummies(df['colour'], prefix='colour')
print(one_hot)
# colour_blue colour_green colour_red
# 0 0 0 1
# 1 1 0 0
# 2 0 1 0
# 3 0 0 1Quick Check
Test your understanding of the Categorical dtype.
Lesson Recap
In this lesson you learned: Categorical dtype stores unique values once and uses integer codes per row, providing large memory savings for low-cardinality columns, ordered Categorical supports meaningful comparisons and correct sorting, and groupby with observed=False includes all categories even empty ones. Avoid Categorical for high-cardinality columns. Next up we parse date strings correctly with pd.to_datetime().
よくある質問
「カテゴリ型」レッスンは無料ですか?
はい。「カテゴリ型」の完全なテキストはこのウェブで無料で読めます。インタラクティブに演習し(組み込みコードエディタと24時間対応のAIチューター)、Pandas & NumPy Academyコースの残りをアンロックするには、CoddyKit PROにアップグレードしてください。 Pandas & NumPy Academyコースには全4レッスンが含まれています。
「カテゴリ型」で何を学びますか?
カーディナリティの低い文字列列をpandasのCategoricalに変換してメモリ使用量を減らし、groupby操作を高速化します。 ブラウザで直接実行するハンズオンコードでPandas & NumPy Academyを演習し、24時間対応のAIチューターがレッスンを進める中での質問に答えます。
Pandas & NumPy Academyを始めるのに経験は必要ですか?
事前経験は必要ありません。CoddyKitのPandas & NumPy Academyは初級者から上級者向けに構成されているため、ここから始めるか最初から始めて、自分のペースで進むことができます。 これはレッスン3/4です。
「カテゴリ型」レッスンにはどのくらい時間がかかりますか?
ほとんどのCoddyKitレッスンは約5~10分かかります。各レッスンはコンパクトでインタラクティブなので、着実に進歩し、ウェブとアプリ全体で正確に前回の場所から再開できます。
このPandas & NumPy Academyレッスンでコードを書いて実行できますか?
はい。すべてのPandas & NumPy Academyレッスンに組み込みコードエディタが含まれているため、ブラウザでリアルコードを書いて実行し、即座のAIフィードバックを取得できます。ローカル設定は不要です。