Сортировка по индексу
Переупорядочивайте строки по меткам индекса с помощью sort_index() и узнайте, когда отсортированный индекс повышает производительность.
«Сортировка по индексу» — бесплатный урок Pandas & NumPy Academy на CoddyKit. Это урок 2 из 4. Ты можешь прочитать весь урок бесплатно ниже — а потом практиковать его прямо в браузере с встроенным редактором кода и ИИ-репетитором 24/7. Это часть пути обучения Pandas & NumPy Academy, и твой прогресс синхронизируется между веб-версией и приложением CoddyKit. Курс Pandas & NumPy Academy содержит 4 уроков всего.
Части этого урока еще не переведены и отображаются на английском.
Understanding the DataFrame Index
Every Pandas DataFrame has a row index — a set of labels used to identify and access rows. By default, this is a RangeIndex (0, 1, 2, …), but you can set it to any column (dates, names, IDs) using set_index(). When the index is meaningful (e.g., a DatetimeIndex or a customer ID), sorting by it rather than by a column value produces a logically organised output.
import pandas as pd
df = pd.DataFrame(
{'value': [10, 20, 30]},
index=['C', 'A', 'B'] # out-of-order alphabetic index
)
print(df)
# value
# C 10
# A 20
# B 30sort_index() — Ascending
DataFrame.sort_index() reorders rows by their index label rather than by column values. By default, sorting is ascending — alphabetically for string indices, numerically for integer indices, and chronologically for DatetimeIndex. This is the standard way to restore a dataset to a natural order after shuffling or appending records out of sequence.
import pandas as pd
df = pd.DataFrame(
{'temp': [22.5, 19.0, 25.1, 18.3]},
index=pd.to_datetime(['2024-03-01', '2024-01-15', '2024-06-10', '2024-01-01'])
)
sorted_df = df.sort_index()
print(sorted_df)
# temp
# 2024-01-01 18.3
# 2024-01-15 19.0
# 2024-03-01 22.5
# 2024-06-10 25.1sort_index() Descending
Pass ascending=False to sort the index from largest (or latest) to smallest (or earliest). For a DatetimeIndex this puts the most recent observations at the top, which is the typical layout for financial data, log files, and event streams where the latest event is most relevant.
import pandas as pd
df = pd.DataFrame(
{'price': [100, 110, 105, 115]},
index=pd.to_datetime(['2024-01', '2024-02', '2024-03', '2024-04'])
)
# Most recent first
print(df.sort_index(ascending=False))
# price
# 2024-04-30 115
# 2024-03-31 105
# 2024-02-29 110
# 2024-01-31 100Sorting Column Labels with axis=1
By default, sort_index() sorts the row index (axis=0). Pass axis=1 to sort the column labels alphabetically instead. This is useful for standardising wide DataFrames with many columns so columns appear in a predictable alphabetical order, making it easier to visually find a column or compare DataFrames.
import pandas as pd
df = pd.DataFrame({
'zebra': [1], 'apple': [2], 'mango': [3], 'banana': [4]
})
print('Before:', df.columns.tolist())
# ['zebra', 'apple', 'mango', 'banana']
sorted_cols = df.sort_index(axis=1)
print('After:', sorted_cols.columns.tolist())
# ['apple', 'banana', 'mango', 'zebra']When is a Sorted Index Faster?
Pandas can use binary search for label look-ups when the index is sorted (monotonic). A sorted index makes .loc['2024-01':'2024-06'] slices O(log n) instead of O(n). The method is_monotonic_increasing (or is_monotonic_decreasing) returns a boolean indicating whether the index is already sorted. Sorting a large index before slicing repeatedly is a worthwhile one-time cost.
import pandas as pd
import numpy as np
idx = pd.to_datetime(['2024-03-01', '2024-01-15', '2024-06-10'])
df = pd.DataFrame({'v': [1, 2, 3]}, index=idx)
print('Sorted?', df.index.is_monotonic_increasing) # False
df = df.sort_index()
print('Sorted?', df.index.is_monotonic_increasing) # True
# Now slicing is efficient
print(df.loc['2024-01':'2024-03'])sort_index() with MultiIndex
When a DataFrame has a MultiIndex (hierarchical row index), sort_index() sorts all levels in the hierarchy by default. You can restrict sorting to specific levels with the level parameter. A sorted MultiIndex is required for efficient hierarchical slicing with .loc[(outer, inner), :].
import pandas as pd
arrays = [
['B', 'B', 'A', 'A'],
['two', 'one', 'two', 'one']
]
idx = pd.MultiIndex.from_arrays(arrays, names=['first', 'second'])
df = pd.DataFrame({'value': [10, 20, 30, 40]}, index=idx)
print(df.sort_index())
# value
# first second
# A one 40
# two 30
# B one 20
# two 10Sorting Only One Level of MultiIndex
With a MultiIndex, you may want to sort on only the inner or outer level while keeping the other level's order intact. Pass level= to sort_index() — it accepts an integer (level position), a string (level name), or a list. This is useful when the outer level order is already correct and you only need to sort within each group.
import pandas as pd
df = pd.DataFrame({
'sales': [300, 100, 200, 400, 150, 250]
}, index=pd.MultiIndex.from_tuples([
('Eng', 'Dave'), ('Eng', 'Alice'), ('Eng', 'Bob'),
('HR', 'Zoe'), ('HR', 'Carol'), ('HR', 'Eve')
], names=['dept', 'name']))
# Sort only the 'name' level alphabetically within each dept
print(df.sort_index(level='name'))
# sales
# dept name
# Eng Alice 100
# Bob 200
# Dave 300
# HR Carol 150
# Eve 250
# Zoe 400Difference Between sort_index() and sort_values()
It is important to distinguish the two sort methods. sort_values(by='col') reorders rows based on the data values in a column. sort_index() reorders rows based on the row label (the index), which may or may not correspond to any column. When the index is the primary identifier (e.g., a DatetimeIndex or a meaningful string key), sort_index() is the right choice.
import pandas as pd
df = pd.DataFrame(
{'value': [30, 10, 20]},
index=['C', 'A', 'B']
)
# sort_index: sorted by row label A, B, C
print(df.sort_index())
# value
# A 10
# B 20
# C 30
# sort_values: sorted by data value 10, 20, 30
print(df.sort_values('value'))
# value
# A 10
# B 20
# C 30
# (same here because values happen to match alphabetical label order!)Restoring Original Order After Ops
Some operations (shuffling, random sampling with df.sample(frac=1)) scramble the row order. sort_index() is the clean way to restore the original sequential order. If the original order was a RangeIndex (0, 1, 2, …), sort_index() restores it; if it was a meaningful label, it restores that label's natural ordering.
import pandas as pd
df = pd.DataFrame({'x': [10, 20, 30, 40, 50]})
# Shuffle (random sample)
shuffled = df.sample(frac=1, random_state=42)
print('Shuffled index:', shuffled.index.tolist())
# e.g. [2, 4, 0, 1, 3]
# Restore original order by sorting the index
restored = shuffled.sort_index()
print('Restored index:', restored.index.tolist())
# [0, 1, 2, 3, 4]sort_index() with na_position
Like sort_values(), sort_index() also supports na_position for controlling where NaN index labels appear. This matters when a DataFrame has a string or date index that contains some NaN labels (possible after operations that introduce missing index values). The default is 'last'.
import pandas as pd
import numpy as np
df = pd.DataFrame(
{'v': [1, 2, 3, 4]},
index=['B', None, 'A', 'C']
)
print(df.sort_index(na_position='last'))
# v
# A 3
# B 1
# C 4
# NaN 2Performance Gain Measurement
You can empirically measure the performance benefit of a sorted index by using Python's timeit to compare a slice on an unsorted vs. sorted DatetimeIndex. The sorted case uses binary search and is typically 5-20x faster for large DataFrames. This demonstrates why sort_index() is not just a cosmetic operation — it has real runtime implications.
import pandas as pd
import numpy as np
np.random.seed(0)
random_dates = pd.to_datetime(
pd.Timestamp('2020-01-01').value + np.random.randint(0, 1_000_000_000_000_000, size=100_000),
unit='ns'
)
df = pd.DataFrame({'val': np.random.randn(100_000)}, index=random_dates)
# Without sort: O(n) scan
import timeit
t1 = timeit.timeit(lambda: df.loc['2022-01':'2022-06'], number=100)
df_sorted = df.sort_index()
t2 = timeit.timeit(lambda: df_sorted.loc['2022-01':'2022-06'], number=100)
print(f'Unsorted: {t1:.3f}s, Sorted: {t2:.3f}s, Speedup: {t1/t2:.1f}x')Quick Check
Test your understanding of sorting by index in Pandas.
Lesson Recap
In this lesson you learned: sort_index() reorders rows by their label (not column values), axis=1 sorts column labels alphabetically, and a sorted index enables binary search making time-based slicing much faster. For MultiIndex DataFrames, level= restricts sorting to one hierarchy level. Always check is_monotonic_increasing before relying on efficient slice lookups. Next up we rank values within a column.
Часто задаваемые вопросы
Урок «Сортировка по индексу» бесплатный?
Да — полный текст урока «Сортировка по индексу» бесплатно доступен здесь в веб-версии. Чтобы практиковать его интерактивно (встроенный редактор кода и ИИ-репетитор 24/7) и разблокировать остальной курс Pandas & NumPy Academy, подпишись на CoddyKit PRO. Курс Pandas & NumPy Academy содержит 4 уроков всего.
Чему я научусь в уроке «Сортировка по индексу»?
Переупорядочивайте строки по меткам индекса с помощью sort_index() и узнайте, когда отсортированный индекс повышает производительность. Ты практикуешь Pandas & NumPy Academy с помощью реального кода, который запускаешь прямо в браузере, и ИИ-репетитор 24/7 отвечает на твои вопросы во время урока.
Нужен ли мне опыт, чтобы начать Pandas & NumPy Academy?
Предыдущий опыт не требуется. Pandas & NumPy Academy на CoddyKit структурирован для всех уровней — от новичков до продвинутых, поэтому ты можешь начать отсюда или с самого начала и учиться в своем темпе. Это урок 2 из 4.
Сколько времени занимает урок «Сортировка по индексу»?
Большинство уроков CoddyKit занимают около 5–10 минут. Каждый из них компактный и интерактивный, поэтому ты постоянно делаешь прогресс и продолжаешь с того же места в веб-версии и приложении.
Можно ли писать и запускать код в этом уроке Pandas & NumPy Academy?
Да. Каждый урок Pandas & NumPy Academy включает встроенный редактор кода, поэтому ты пишешь и запускаешь реальный код прямо в браузере и получаешь моментальную обратную связь от AI — локальная установка не требуется.
Все уроки этого курса
- Сортировка по значениям столбцов
- Сортировка по индексу
- Ранжирование значений
- Установка и сброс индекса