Ordenando pelo índice
Reordene as linhas pelo rótulo do índice com sort_index() e entenda quando um índice ordenado melhora o desempenho.
Ordenando pelo índice é uma aula grátis de Pandas & NumPy Academy no CoddyKit. Esta é a aula 2 de 4. Você pode ler a aula completa abaixo gratuitamente — depois pratica ao vivo no navegador com um editor de código integrado e um tutor de IA 24/7. Faz parte do caminho de aprendizado de Pandas & NumPy Academy, e seu progresso é sincronizado entre a web e o app CoddyKit. O curso de Pandas & NumPy Academy inclui 4 aulas no total.
Partes desta aula ainda não foram traduzidas e aparecem em inglês.
Understanding the DataFrame Index
Every Pandas DataFrame has a row index — a set of labels used to identify and access rows. By default, this is a RangeIndex (0, 1, 2, …), but you can set it to any column (dates, names, IDs) using set_index(). When the index is meaningful (e.g., a DatetimeIndex or a customer ID), sorting by it rather than by a column value produces a logically organised output.
import pandas as pd
df = pd.DataFrame(
{'value': [10, 20, 30]},
index=['C', 'A', 'B'] # out-of-order alphabetic index
)
print(df)
# value
# C 10
# A 20
# B 30sort_index() — Ascending
DataFrame.sort_index() reorders rows by their index label rather than by column values. By default, sorting is ascending — alphabetically for string indices, numerically for integer indices, and chronologically for DatetimeIndex. This is the standard way to restore a dataset to a natural order after shuffling or appending records out of sequence.
import pandas as pd
df = pd.DataFrame(
{'temp': [22.5, 19.0, 25.1, 18.3]},
index=pd.to_datetime(['2024-03-01', '2024-01-15', '2024-06-10', '2024-01-01'])
)
sorted_df = df.sort_index()
print(sorted_df)
# temp
# 2024-01-01 18.3
# 2024-01-15 19.0
# 2024-03-01 22.5
# 2024-06-10 25.1sort_index() Descending
Pass ascending=False to sort the index from largest (or latest) to smallest (or earliest). For a DatetimeIndex this puts the most recent observations at the top, which is the typical layout for financial data, log files, and event streams where the latest event is most relevant.
import pandas as pd
df = pd.DataFrame(
{'price': [100, 110, 105, 115]},
index=pd.to_datetime(['2024-01', '2024-02', '2024-03', '2024-04'])
)
# Most recent first
print(df.sort_index(ascending=False))
# price
# 2024-04-30 115
# 2024-03-31 105
# 2024-02-29 110
# 2024-01-31 100Sorting Column Labels with axis=1
By default, sort_index() sorts the row index (axis=0). Pass axis=1 to sort the column labels alphabetically instead. This is useful for standardising wide DataFrames with many columns so columns appear in a predictable alphabetical order, making it easier to visually find a column or compare DataFrames.
import pandas as pd
df = pd.DataFrame({
'zebra': [1], 'apple': [2], 'mango': [3], 'banana': [4]
})
print('Before:', df.columns.tolist())
# ['zebra', 'apple', 'mango', 'banana']
sorted_cols = df.sort_index(axis=1)
print('After:', sorted_cols.columns.tolist())
# ['apple', 'banana', 'mango', 'zebra']When is a Sorted Index Faster?
Pandas can use binary search for label look-ups when the index is sorted (monotonic). A sorted index makes .loc['2024-01':'2024-06'] slices O(log n) instead of O(n). The method is_monotonic_increasing (or is_monotonic_decreasing) returns a boolean indicating whether the index is already sorted. Sorting a large index before slicing repeatedly is a worthwhile one-time cost.
import pandas as pd
import numpy as np
idx = pd.to_datetime(['2024-03-01', '2024-01-15', '2024-06-10'])
df = pd.DataFrame({'v': [1, 2, 3]}, index=idx)
print('Sorted?', df.index.is_monotonic_increasing) # False
df = df.sort_index()
print('Sorted?', df.index.is_monotonic_increasing) # True
# Now slicing is efficient
print(df.loc['2024-01':'2024-03'])sort_index() with MultiIndex
When a DataFrame has a MultiIndex (hierarchical row index), sort_index() sorts all levels in the hierarchy by default. You can restrict sorting to specific levels with the level parameter. A sorted MultiIndex is required for efficient hierarchical slicing with .loc[(outer, inner), :].
import pandas as pd
arrays = [
['B', 'B', 'A', 'A'],
['two', 'one', 'two', 'one']
]
idx = pd.MultiIndex.from_arrays(arrays, names=['first', 'second'])
df = pd.DataFrame({'value': [10, 20, 30, 40]}, index=idx)
print(df.sort_index())
# value
# first second
# A one 40
# two 30
# B one 20
# two 10Sorting Only One Level of MultiIndex
With a MultiIndex, you may want to sort on only the inner or outer level while keeping the other level's order intact. Pass level= to sort_index() — it accepts an integer (level position), a string (level name), or a list. This is useful when the outer level order is already correct and you only need to sort within each group.
import pandas as pd
df = pd.DataFrame({
'sales': [300, 100, 200, 400, 150, 250]
}, index=pd.MultiIndex.from_tuples([
('Eng', 'Dave'), ('Eng', 'Alice'), ('Eng', 'Bob'),
('HR', 'Zoe'), ('HR', 'Carol'), ('HR', 'Eve')
], names=['dept', 'name']))
# Sort only the 'name' level alphabetically within each dept
print(df.sort_index(level='name'))
# sales
# dept name
# Eng Alice 100
# Bob 200
# Dave 300
# HR Carol 150
# Eve 250
# Zoe 400Difference Between sort_index() and sort_values()
It is important to distinguish the two sort methods. sort_values(by='col') reorders rows based on the data values in a column. sort_index() reorders rows based on the row label (the index), which may or may not correspond to any column. When the index is the primary identifier (e.g., a DatetimeIndex or a meaningful string key), sort_index() is the right choice.
import pandas as pd
df = pd.DataFrame(
{'value': [30, 10, 20]},
index=['C', 'A', 'B']
)
# sort_index: sorted by row label A, B, C
print(df.sort_index())
# value
# A 10
# B 20
# C 30
# sort_values: sorted by data value 10, 20, 30
print(df.sort_values('value'))
# value
# A 10
# B 20
# C 30
# (same here because values happen to match alphabetical label order!)Restoring Original Order After Ops
Some operations (shuffling, random sampling with df.sample(frac=1)) scramble the row order. sort_index() is the clean way to restore the original sequential order. If the original order was a RangeIndex (0, 1, 2, …), sort_index() restores it; if it was a meaningful label, it restores that label's natural ordering.
import pandas as pd
df = pd.DataFrame({'x': [10, 20, 30, 40, 50]})
# Shuffle (random sample)
shuffled = df.sample(frac=1, random_state=42)
print('Shuffled index:', shuffled.index.tolist())
# e.g. [2, 4, 0, 1, 3]
# Restore original order by sorting the index
restored = shuffled.sort_index()
print('Restored index:', restored.index.tolist())
# [0, 1, 2, 3, 4]sort_index() with na_position
Like sort_values(), sort_index() also supports na_position for controlling where NaN index labels appear. This matters when a DataFrame has a string or date index that contains some NaN labels (possible after operations that introduce missing index values). The default is 'last'.
import pandas as pd
import numpy as np
df = pd.DataFrame(
{'v': [1, 2, 3, 4]},
index=['B', None, 'A', 'C']
)
print(df.sort_index(na_position='last'))
# v
# A 3
# B 1
# C 4
# NaN 2Performance Gain Measurement
You can empirically measure the performance benefit of a sorted index by using Python's timeit to compare a slice on an unsorted vs. sorted DatetimeIndex. The sorted case uses binary search and is typically 5-20x faster for large DataFrames. This demonstrates why sort_index() is not just a cosmetic operation — it has real runtime implications.
import pandas as pd
import numpy as np
np.random.seed(0)
random_dates = pd.to_datetime(
pd.Timestamp('2020-01-01').value + np.random.randint(0, 1_000_000_000_000_000, size=100_000),
unit='ns'
)
df = pd.DataFrame({'val': np.random.randn(100_000)}, index=random_dates)
# Without sort: O(n) scan
import timeit
t1 = timeit.timeit(lambda: df.loc['2022-01':'2022-06'], number=100)
df_sorted = df.sort_index()
t2 = timeit.timeit(lambda: df_sorted.loc['2022-01':'2022-06'], number=100)
print(f'Unsorted: {t1:.3f}s, Sorted: {t2:.3f}s, Speedup: {t1/t2:.1f}x')Quick Check
Test your understanding of sorting by index in Pandas.
Lesson Recap
In this lesson you learned: sort_index() reorders rows by their label (not column values), axis=1 sorts column labels alphabetically, and a sorted index enables binary search making time-based slicing much faster. For MultiIndex DataFrames, level= restricts sorting to one hierarchy level. Always check is_monotonic_increasing before relying on efficient slice lookups. Next up we rank values within a column.
Perguntas Frequentes
A aula “Ordenando pelo índice” é grátis?
Sim — o texto completo de “Ordenando pelo índice” é grátis para ler aqui na web. Para praticá-la interativamente (um editor de código integrado e um tutor de IA 24/7) e desbloquear o restante do curso de Pandas & NumPy Academy, atualize para CoddyKit PRO. O curso de Pandas & NumPy Academy inclui 4 aulas no total.
O que vou aprender em “Ordenando pelo índice”?
Reordene as linhas pelo rótulo do índice com sort_index() e entenda quando um índice ordenado melhora o desempenho. Você pratica Pandas & NumPy Academy com código prático que executa diretamente no navegador, e um tutor de IA 24/7 responde suas dúvidas enquanto trabalha na aula.
Preciso ter experiência prévia para começar Pandas & NumPy Academy?
Nenhuma experiência prévia é necessária. Pandas & NumPy Academy no CoddyKit é estruturado para alunos iniciantes até avançados, então você pode começar aqui ou desde o início e aprender no seu ritmo. Esta é a aula 2 de 4.
Quanto tempo leva a aula “Ordenando pelo índice”?
A maioria das aulas CoddyKit leva cerca de 5–10 minutos. Cada uma é compacta e interativa, então você faz progresso constante e retoma exatamente de onde parou entre web e app.
Posso escrever e executar código nesta aula de Pandas & NumPy Academy?
Sim. Cada aula de Pandas & NumPy Academy inclui um editor de código integrado, então você escreve e executa código real direto no navegador e recebe feedback de IA instantaneamente — nenhuma configuração local necessária.
Todas as aulas deste curso
- Ordenando por valores das colunas
- Ordenando pelo índice
- Classificando valores
- Definindo e redefinindo o índice