0Pricing
Pandas & NumPy Academy · Aula

Tipo de dados categórico

Converta colunas de texto com baixa cardinalidade para Categorical do pandas a fim de reduzir o uso de memória e acelerar operações de agrupamento.

Tipo de dados categórico é uma aula grátis de Pandas & NumPy Academy no CoddyKit. Esta é a aula 3 de 4. Você pode ler a aula completa abaixo gratuitamente — depois pratica ao vivo no navegador com um editor de código integrado e um tutor de IA 24/7. Faz parte do caminho de aprendizado de Pandas & NumPy Academy, e seu progresso é sincronizado entre a web e o app CoddyKit. O curso de Pandas & NumPy Academy inclui 4 aulas no total.

Partes desta aula ainda não foram traduzidas e aparecem em inglês.

What Is the Categorical Dtype?

The Categorical dtype in Pandas is designed for columns that contain a limited set of discrete values (low cardinality), such as gender, status, region, or product category. Instead of storing the full string for every row, Pandas stores the unique values (called categories) once and uses an integer code per row to reference them. This is similar to how databases use lookup tables or enum types.

import pandas as pd

df = pd.DataFrame({
    'region': ['North', 'South', 'East', 'North', 'East', 'South'] * 1000
})

# Without Categorical: object dtype stores every string
print('object dtype memory:', df['region'].memory_usage(deep=True))

# With Categorical: only stores 3 unique values + integer codes
df['region_cat'] = df['region'].astype('category')
print('category dtype memory:', df['region_cat'].memory_usage(deep=True))

Creating a Categorical Series

Convert a column to Categorical by calling .astype('category') on any Series. The resulting Categorical Series stores the unique values as .cat.categories and the integer position codes as .cat.codes. You can also create a Categorical directly with pd.Categorical() to specify the categories and their order up front.

import pandas as pd

s = pd.Series(['low', 'high', 'med', 'high', 'low'])
s_cat = s.astype('category')

print(s_cat.cat.categories)  # Index(['high', 'low', 'med'], dtype='object')
print(s_cat.cat.codes)       # integer codes per row
# 0    1
# 1    0
# 2    2
# 3    0
# 4    1

Memory Savings with Categorical

The memory savings from Categorical dtype depend on the cardinality ratio (unique values ÷ total rows). With 5 unique strings in a 1,000,000-row column, the object dtype stores 1 million string pointers (~50+ bytes each) while Categorical stores just 5 strings plus 1 million 8-bit integers. The savings can be 5-20x, turning a 50 MB column into 2-5 MB.

import pandas as pd
import numpy as np

n = 1_000_000
statuses = np.random.choice(['active', 'inactive', 'pending'], size=n)
df = pd.DataFrame({'status': statuses})

object_mem = df['status'].memory_usage(deep=True) / 1e6
df['status'] = df['status'].astype('category')
cat_mem = df['status'].memory_usage(deep=True) / 1e6

print(f'Object dtype:    {object_mem:.1f} MB')
print(f'Category dtype:  {cat_mem:.1f} MB')
print(f'Reduction:       {object_mem/cat_mem:.1f}x')

Ordered Categorical for Ranking

Sometimes categories have a natural order — e.g., 'low' < 'medium' < 'high', or T-shirt sizes XS < S < M < L < XL. An ordered Categorical captures this ranking, enabling meaningful comparison operators (<, >=, etc.) and ensuring groupby results are sorted in the correct semantic order rather than alphabetically.

import pandas as pd

df = pd.DataFrame({
    'priority': ['high', 'low', 'medium', 'high', 'low']
})

priority_type = pd.CategoricalDtype(
    categories=['low', 'medium', 'high'],
    ordered=True
)
df['priority'] = df['priority'].astype(priority_type)

# Comparison now works correctly
print(df['priority'] >= 'medium')
# 0     True
# 1    False
# 2     True
# 3     True
# 4    False

Sorting Ordered Categoricals

When you sort an ordered Categorical column, Pandas uses the defined category order rather than alphabetical order. This means 'low' comes before 'medium' before 'high' regardless of how they would sort as strings. This is critical for producing correctly ordered summary tables and charts.

import pandas as pd

df = pd.DataFrame({
    'severity': pd.Categorical(
        ['high', 'low', 'medium', 'low', 'high'],
        categories=['low', 'medium', 'high'],
        ordered=True
    ),
    'count': [5, 20, 8, 15, 3]
})

# sort_values respects the categorical order
print(df.sort_values('severity')[['severity', 'count']])
#   severity  count
# 1      low     20
# 3      low     15
# 2   medium      8
# 0     high      5
# 4     high      3

GroupBy Speed with Categorical

Pandas can optimise groupby() operations on Categorical columns because it knows all possible groups in advance. This can make groupby 2-5x faster on large DataFrames. Additionally, groupby on Categorical returns all categories — including empty ones — which ensures your summary tables always have a row for every expected group, even if some have zero observations.

import pandas as pd
import numpy as np

np.random.seed(0)
df = pd.DataFrame({
    'region': pd.Categorical(
        np.random.choice(['North', 'South', 'East', 'West'], 10),
        categories=['North', 'South', 'East', 'West']
    ),
    'sales': np.random.randint(100, 1000, 10)
})

# observed=False shows ALL categories even empty ones
print(df.groupby('region', observed=False)['sales'].sum())

Adding and Removing Categories

Use the .cat accessor to manage the category list. .cat.add_categories() adds new allowed values (useful before inserting new data), and .cat.remove_unused_categories() drops category labels that have no corresponding rows — handy after filtering. You cannot assign a value that is not in the categories without adding it first.

import pandas as pd

s = pd.Categorical(['a', 'b', 'a'], categories=['a', 'b', 'c'])
s = pd.Series(s)

print(s.cat.categories)  # Index(['a', 'b', 'c'], dtype='object')
print(s.value_counts())  # a:2, b:1, c:0

# Remove unused category 'c'
s = s.cat.remove_unused_categories()
print(s.cat.categories)  # Index(['a', 'b'], dtype='object')

Renaming Category Labels

.cat.rename_categories() lets you relabel categories without changing the underlying codes. This is useful when you want to display friendlier names (e.g., 'M' → 'Male') or fix inconsistent label capitalisation without converting back to object and remapping. The method accepts a list (positional) or a dictionary (targeted).

import pandas as pd

s = pd.Series(pd.Categorical(['M', 'F', 'M', 'F'], categories=['M', 'F']))

# Rename using a dict
s = s.cat.rename_categories({'M': 'Male', 'F': 'Female'})
print(s)
# 0      Male
# 1    Female
# 2      Male
# 3    Female
print(s.cat.categories)  # Index(['Male', 'Female'], dtype='object')

When NOT to Use Categorical

Categorical dtype is a poor choice when cardinality is high — if almost every row has a unique value (like user IDs, free-text comments, or UUIDs), the category index is nearly as large as the original data, giving no memory saving. A rule of thumb: use Categorical only when the number of unique values is less than roughly 50% of the total rows, and especially when unique values number in the tens or hundreds.

import pandas as pd
import numpy as np

df = pd.DataFrame({'user_id': range(1_000_000)})

# High-cardinality: every value is unique — don't use Categorical
object_mem = df['user_id'].astype(str).memory_usage(deep=True) / 1e6
cat_mem = df['user_id'].astype(str).astype('category').memory_usage(deep=True) / 1e6

print(f'String: {object_mem:.1f} MB')
print(f'Category: {cat_mem:.1f} MB')
# Category is WORSE for high-cardinality data!

Categorical in Pivot Tables and Groupby

Using Categorical ensures consistent output shape in groupby and pivot table results. Without Categorical, groups that happen to have zero observations are silently omitted from the result — which can cause misaligned output when comparing across different data slices. With Categorical and observed=False, all groups always appear in the output.

import pandas as pd

df = pd.DataFrame({
    'quarter': pd.Categorical(
        ['Q1', 'Q1', 'Q3'],
        categories=['Q1', 'Q2', 'Q3', 'Q4']
    ),
    'revenue': [100, 200, 300]
})

# observed=False includes Q2 and Q4 even though they have no rows
result = df.groupby('quarter', observed=False)['revenue'].sum()
print(result)
# quarter
# Q1    300
# Q2      0
# Q3    300
# Q4      0

Categorical and Machine Learning

Many scikit-learn estimators expect numeric input. To convert a Categorical column to numbers for a model, use .cat.codes (label encoding) or pd.get_dummies() (one-hot encoding). One-hot encoding avoids implying an ordinal relationship between categories. For ordered categoricals like priority levels, label encoding (the codes directly) is appropriate and carries the ordering information.

import pandas as pd

df = pd.DataFrame({
    'colour': pd.Categorical(['red', 'blue', 'green', 'red'])
})

# One-hot encode for unordered categories
one_hot = pd.get_dummies(df['colour'], prefix='colour')
print(one_hot)
#    colour_blue  colour_green  colour_red
# 0            0             0           1
# 1            1             0           0
# 2            0             1           0
# 3            0             0           1

Quick Check

Test your understanding of the Categorical dtype.

Lesson Recap

In this lesson you learned: Categorical dtype stores unique values once and uses integer codes per row, providing large memory savings for low-cardinality columns, ordered Categorical supports meaningful comparisons and correct sorting, and groupby with observed=False includes all categories even empty ones. Avoid Categorical for high-cardinality columns. Next up we parse date strings correctly with pd.to_datetime().

Perguntas Frequentes

A aula “Tipo de dados categórico” é grátis?

Sim — o texto completo de “Tipo de dados categórico” é grátis para ler aqui na web. Para praticá-la interativamente (um editor de código integrado e um tutor de IA 24/7) e desbloquear o restante do curso de Pandas & NumPy Academy, atualize para CoddyKit PRO. O curso de Pandas & NumPy Academy inclui 4 aulas no total.

O que vou aprender em “Tipo de dados categórico”?

Converta colunas de texto com baixa cardinalidade para Categorical do pandas a fim de reduzir o uso de memória e acelerar operações de agrupamento. Você pratica Pandas & NumPy Academy com código prático que executa diretamente no navegador, e um tutor de IA 24/7 responde suas dúvidas enquanto trabalha na aula.

Preciso ter experiência prévia para começar Pandas & NumPy Academy?

Nenhuma experiência prévia é necessária. Pandas & NumPy Academy no CoddyKit é estruturado para alunos iniciantes até avançados, então você pode começar aqui ou desde o início e aprender no seu ritmo. Esta é a aula 3 de 4.

Quanto tempo leva a aula “Tipo de dados categórico”?

A maioria das aulas CoddyKit leva cerca de 5–10 minutos. Cada uma é compacta e interativa, então você faz progresso constante e retoma exatamente de onde parou entre web e app.

Posso escrever e executar código nesta aula de Pandas & NumPy Academy?

Sim. Cada aula de Pandas & NumPy Academy inclui um editor de código integrado, então você escreve e executa código real direto no navegador e recebe feedback de IA instantaneamente — nenhuma configuração local necessária.

Todas as aulas deste curso

  1. Inspecionando os tipos de dados das colunas
  2. Conversão de tipos com astype()
  3. Tipo de dados categórico
  4. Analisando datas corretamente
← Voltar para Pandas & NumPy Academy