0Pricing
Pandas & NumPy Academy · Lesson

Categorical Data Type

Convert low-cardinality string columns to pandas Categorical to reduce memory usage and speed up groupby operations.

Categorical Data Type is a free Pandas & NumPy Academy lesson on CoddyKit — lesson 3 of 4. You can read the complete lesson below for free — then practise it hands-on in the browser with a built-in code editor and a 24/7 AI tutor. It is part of the Pandas & NumPy Academy learning path, one of 4 lessons in the course, and your progress syncs across the web and the CoddyKit app.

What Is the Categorical Dtype?

The Categorical dtype in Pandas is designed for columns that contain a limited set of discrete values (low cardinality), such as gender, status, region, or product category. Instead of storing the full string for every row, Pandas stores the unique values (called categories) once and uses an integer code per row to reference them. This is similar to how databases use lookup tables or enum types.

import pandas as pd

df = pd.DataFrame({
    'region': ['North', 'South', 'East', 'North', 'East', 'South'] * 1000
})

# Without Categorical: object dtype stores every string
print('object dtype memory:', df['region'].memory_usage(deep=True))

# With Categorical: only stores 3 unique values + integer codes
df['region_cat'] = df['region'].astype('category')
print('category dtype memory:', df['region_cat'].memory_usage(deep=True))

Creating a Categorical Series

Convert a column to Categorical by calling .astype('category') on any Series. The resulting Categorical Series stores the unique values as .cat.categories and the integer position codes as .cat.codes. You can also create a Categorical directly with pd.Categorical() to specify the categories and their order up front.

import pandas as pd

s = pd.Series(['low', 'high', 'med', 'high', 'low'])
s_cat = s.astype('category')

print(s_cat.cat.categories)  # Index(['high', 'low', 'med'], dtype='object')
print(s_cat.cat.codes)       # integer codes per row
# 0    1
# 1    0
# 2    2
# 3    0
# 4    1

Memory Savings with Categorical

The memory savings from Categorical dtype depend on the cardinality ratio (unique values ÷ total rows). With 5 unique strings in a 1,000,000-row column, the object dtype stores 1 million string pointers (~50+ bytes each) while Categorical stores just 5 strings plus 1 million 8-bit integers. The savings can be 5-20x, turning a 50 MB column into 2-5 MB.

import pandas as pd
import numpy as np

n = 1_000_000
statuses = np.random.choice(['active', 'inactive', 'pending'], size=n)
df = pd.DataFrame({'status': statuses})

object_mem = df['status'].memory_usage(deep=True) / 1e6
df['status'] = df['status'].astype('category')
cat_mem = df['status'].memory_usage(deep=True) / 1e6

print(f'Object dtype:    {object_mem:.1f} MB')
print(f'Category dtype:  {cat_mem:.1f} MB')
print(f'Reduction:       {object_mem/cat_mem:.1f}x')

Ordered Categorical for Ranking

Sometimes categories have a natural order — e.g., 'low' < 'medium' < 'high', or T-shirt sizes XS < S < M < L < XL. An ordered Categorical captures this ranking, enabling meaningful comparison operators (<, >=, etc.) and ensuring groupby results are sorted in the correct semantic order rather than alphabetically.

import pandas as pd

df = pd.DataFrame({
    'priority': ['high', 'low', 'medium', 'high', 'low']
})

priority_type = pd.CategoricalDtype(
    categories=['low', 'medium', 'high'],
    ordered=True
)
df['priority'] = df['priority'].astype(priority_type)

# Comparison now works correctly
print(df['priority'] >= 'medium')
# 0     True
# 1    False
# 2     True
# 3     True
# 4    False

Sorting Ordered Categoricals

When you sort an ordered Categorical column, Pandas uses the defined category order rather than alphabetical order. This means 'low' comes before 'medium' before 'high' regardless of how they would sort as strings. This is critical for producing correctly ordered summary tables and charts.

import pandas as pd

df = pd.DataFrame({
    'severity': pd.Categorical(
        ['high', 'low', 'medium', 'low', 'high'],
        categories=['low', 'medium', 'high'],
        ordered=True
    ),
    'count': [5, 20, 8, 15, 3]
})

# sort_values respects the categorical order
print(df.sort_values('severity')[['severity', 'count']])
#   severity  count
# 1      low     20
# 3      low     15
# 2   medium      8
# 0     high      5
# 4     high      3

GroupBy Speed with Categorical

Pandas can optimise groupby() operations on Categorical columns because it knows all possible groups in advance. This can make groupby 2-5x faster on large DataFrames. Additionally, groupby on Categorical returns all categories — including empty ones — which ensures your summary tables always have a row for every expected group, even if some have zero observations.

import pandas as pd
import numpy as np

np.random.seed(0)
df = pd.DataFrame({
    'region': pd.Categorical(
        np.random.choice(['North', 'South', 'East', 'West'], 10),
        categories=['North', 'South', 'East', 'West']
    ),
    'sales': np.random.randint(100, 1000, 10)
})

# observed=False shows ALL categories even empty ones
print(df.groupby('region', observed=False)['sales'].sum())

Adding and Removing Categories

Use the .cat accessor to manage the category list. .cat.add_categories() adds new allowed values (useful before inserting new data), and .cat.remove_unused_categories() drops category labels that have no corresponding rows — handy after filtering. You cannot assign a value that is not in the categories without adding it first.

import pandas as pd

s = pd.Categorical(['a', 'b', 'a'], categories=['a', 'b', 'c'])
s = pd.Series(s)

print(s.cat.categories)  # Index(['a', 'b', 'c'], dtype='object')
print(s.value_counts())  # a:2, b:1, c:0

# Remove unused category 'c'
s = s.cat.remove_unused_categories()
print(s.cat.categories)  # Index(['a', 'b'], dtype='object')

Renaming Category Labels

.cat.rename_categories() lets you relabel categories without changing the underlying codes. This is useful when you want to display friendlier names (e.g., 'M''Male') or fix inconsistent label capitalisation without converting back to object and remapping. The method accepts a list (positional) or a dictionary (targeted).

import pandas as pd

s = pd.Series(pd.Categorical(['M', 'F', 'M', 'F'], categories=['M', 'F']))

# Rename using a dict
s = s.cat.rename_categories({'M': 'Male', 'F': 'Female'})
print(s)
# 0      Male
# 1    Female
# 2      Male
# 3    Female
print(s.cat.categories)  # Index(['Male', 'Female'], dtype='object')

When NOT to Use Categorical

Categorical dtype is a poor choice when cardinality is high — if almost every row has a unique value (like user IDs, free-text comments, or UUIDs), the category index is nearly as large as the original data, giving no memory saving. A rule of thumb: use Categorical only when the number of unique values is less than roughly 50% of the total rows, and especially when unique values number in the tens or hundreds.

import pandas as pd
import numpy as np

df = pd.DataFrame({'user_id': range(1_000_000)})

# High-cardinality: every value is unique — don't use Categorical
object_mem = df['user_id'].astype(str).memory_usage(deep=True) / 1e6
cat_mem = df['user_id'].astype(str).astype('category').memory_usage(deep=True) / 1e6

print(f'String: {object_mem:.1f} MB')
print(f'Category: {cat_mem:.1f} MB')
# Category is WORSE for high-cardinality data!

Categorical in Pivot Tables and Groupby

Using Categorical ensures consistent output shape in groupby and pivot table results. Without Categorical, groups that happen to have zero observations are silently omitted from the result — which can cause misaligned output when comparing across different data slices. With Categorical and observed=False, all groups always appear in the output.

import pandas as pd

df = pd.DataFrame({
    'quarter': pd.Categorical(
        ['Q1', 'Q1', 'Q3'],
        categories=['Q1', 'Q2', 'Q3', 'Q4']
    ),
    'revenue': [100, 200, 300]
})

# observed=False includes Q2 and Q4 even though they have no rows
result = df.groupby('quarter', observed=False)['revenue'].sum()
print(result)
# quarter
# Q1    300
# Q2      0
# Q3    300
# Q4      0

Categorical and Machine Learning

Many scikit-learn estimators expect numeric input. To convert a Categorical column to numbers for a model, use .cat.codes (label encoding) or pd.get_dummies() (one-hot encoding). One-hot encoding avoids implying an ordinal relationship between categories. For ordered categoricals like priority levels, label encoding (the codes directly) is appropriate and carries the ordering information.

import pandas as pd

df = pd.DataFrame({
    'colour': pd.Categorical(['red', 'blue', 'green', 'red'])
})

# One-hot encode for unordered categories
one_hot = pd.get_dummies(df['colour'], prefix='colour')
print(one_hot)
#    colour_blue  colour_green  colour_red
# 0            0             0           1
# 1            1             0           0
# 2            0             1           0
# 3            0             0           1

Quick Check

Test your understanding of the Categorical dtype.

Lesson Recap

In this lesson you learned: Categorical dtype stores unique values once and uses integer codes per row, providing large memory savings for low-cardinality columns, ordered Categorical supports meaningful comparisons and correct sorting, and groupby with observed=False includes all categories even empty ones. Avoid Categorical for high-cardinality columns. Next up we parse date strings correctly with pd.to_datetime().

Frequently asked questions

Is the “Categorical Data Type” lesson free?

Yes — the full text of “Categorical Data Type” is free to read here on the web, and the Pandas & NumPy Academy course includes 4 lessons in total. To practise it interactively (a built-in code editor and a 24/7 AI tutor) and unlock the rest of the Pandas & NumPy Academy course, upgrade to CoddyKit PRO.

What will I learn in “Categorical Data Type”?

Convert low-cardinality string columns to pandas Categorical to reduce memory usage and speed up groupby operations. You practise Pandas & NumPy Academy with hands-on code you run directly in the browser, and a 24/7 AI tutor answers your questions as you work through the lesson.

Do I need any experience to start Pandas & NumPy Academy?

No prior experience is required. Pandas & NumPy Academy on CoddyKit is structured for beginners through advanced learners; this is — lesson 3 of 4, so you can start here or from the beginning and move at your own pace.

How long does the “Categorical Data Type” lesson take?

Most CoddyKit lessons take about 5–10 minutes. Each one is bite-sized and interactive, so you make steady progress and pick up exactly where you left off across the web and the app.

Can I write and run code in this Pandas & NumPy Academy lesson?

Yes. Every Pandas & NumPy Academy lesson includes a built-in code editor, so you write and run real code right in your browser and get instant AI feedback — no local setup required.

All lessons in this course

  1. Inspecting Column Data Types
  2. Casting with astype()
  3. Categorical Data Type
  4. Parsing Dates Correctly
← Back to Pandas & NumPy Academy