0Pricing
Pandas & NumPy Academy · Урок

pd.concat для объединения DataFrames

Объединяйте DataFrames вертикально или горизонтально с помощью pd.concat, выравнивая их по индексу или столбцам и обрабатывая дублирующиеся индексы.

«pd.concat для объединения DataFrames» — бесплатный урок Pandas & NumPy Academy на CoddyKit. Это урок 1 из 4. Ты можешь прочитать весь урок бесплатно ниже — а потом практиковать его прямо в браузере с встроенным редактором кода и ИИ-репетитором 24/7. Это часть пути обучения Pandas & NumPy Academy, и твой прогресс синхронизируется между веб-версией и приложением CoddyKit. Курс Pandas & NumPy Academy содержит 4 уроков всего.

Части этого урока еще не переведены и отображаются на английском.

Why Concatenate DataFrames?

In real projects, data rarely arrives in a single file. You might have monthly sales files, regional survey exports, or database query results split across multiple queries. Concatenation stacks these separate DataFrames into one combined table. Pandas provides pd.concat() for this purpose, handling both vertical (row-wise) and horizontal (column-wise) stacking.

Basic Vertical Concatenation

The most common use of pd.concat() is stacking DataFrames vertically (adding more rows). Pass a list of DataFrames and the function appends them one below the other. Both DataFrames must have compatible columns for a clean result. Pandas aligns on column names automatically, filling missing columns with NaN.

import pandas as pd

jan = pd.DataFrame({'product': ['A', 'B'], 'sales': [100, 200]})
feb = pd.DataFrame({'product': ['A', 'C'], 'sales': [150, 120]})

combined = pd.concat([jan, feb])
print(combined)
#   product  sales
# 0       A    100
# 1       B    200
# 0       A    150
# 1       C    120

Resetting the Index After Concat

Notice that pd.concat() preserves the original indices from each DataFrame, which can result in duplicate index values (as seen above with two rows at index 0 and index 1). Pass ignore_index=True to create a fresh sequential integer index in the combined result, which is almost always what you want.

combined = pd.concat([jan, feb], ignore_index=True)
print(combined)
#   product  sales
# 0       A    100
# 1       B    200
# 2       A    150
# 3       C    120

print(combined.index)
# RangeIndex(start=0, stop=4, step=1)

Tracking Source with keys

When concatenating data from multiple sources, you may want to know which original DataFrame each row came from. Pass the keys parameter with a list of labels. Pandas creates a MultiIndex where the outer level identifies the source. You can then use .loc['label'] to access rows from a specific source.

combined = pd.concat([jan, feb], keys=['January', 'February'])
print(combined)
#                product  sales
# January   0         A    100
#           1         B    200
# February  0         A    150
#           1         C    120

# Access only February rows
print(combined.loc['February'])

Horizontal Concatenation with axis=1

Pass axis=1 to concatenate DataFrames side by side (adding more columns). Pandas aligns on the row index, so both DataFrames should have the same index for a clean result. If indices differ, non-matching rows will be filled with NaN. This is useful for combining features computed from the same dataset in separate steps.

names = pd.DataFrame({'name': ['Alice', 'Bob', 'Carol']}, index=[1, 2, 3])
scores = pd.DataFrame({'score': [95, 88, 72]}, index=[1, 2, 3])

combined = pd.concat([names, scores], axis=1)
print(combined)
#     name  score
# 1  Alice     95
# 2    Bob     88
# 3  Carol     72

Handling Mismatched Columns

If the DataFrames being concatenated have different columns, Pandas keeps all columns and fills missing values with NaN. This is the default join='outer' behaviour. If you want to keep only the columns that appear in all DataFrames, pass join='inner', which drops any column not present in every source.

df1 = pd.DataFrame({'a': [1, 2], 'b': [3, 4]})
df2 = pd.DataFrame({'b': [5, 6], 'c': [7, 8]})

# Outer join (default): keeps all columns
print(pd.concat([df1, df2], ignore_index=True))
#      a  b    c
# 0  1.0  3  NaN
# 1  2.0  4  NaN
# 2  NaN  5  7.0
# 3  NaN  6  8.0

# Inner join: keeps only shared columns
print(pd.concat([df1, df2], join='inner', ignore_index=True))
#    b
# 0  3
# 1  4
# 2  5
# 3  6

Concatenating Many Files in a Loop

A typical pattern when loading multiple files is to collect all DataFrames in a list and call pd.concat() once at the end. Avoid concatenating inside a loop (e.g., df = pd.concat([df, new_chunk])) because this creates a new DataFrame copy on every iteration, leading to quadratic time complexity for large collections.

import glob

# Efficient: collect first, concat once
files = glob.glob('data/sales_*.csv')
frames = [pd.read_csv(f) for f in files]
combined = pd.concat(frames, ignore_index=True)

# Inefficient (avoid):
# result = pd.DataFrame()
# for f in files:
#     result = pd.concat([result, pd.read_csv(f)])  # slow!

Concatenating Series

pd.concat() works with Series as well as DataFrames. Concatenating a list of Series vertically gives a longer Series. Concatenating with axis=1 produces a DataFrame where each Series becomes a column. The Series are aligned on their index, so the index labels must match for a clean horizontal concat.

s1 = pd.Series([1, 2, 3], name='x')
s2 = pd.Series([4, 5, 6], name='y')

# Vertical: one long Series
print(pd.concat([s1, s2]))
# 0    1
# 1    2
# ...

# Horizontal: a DataFrame with two columns
print(pd.concat([s1, s2], axis=1))
#    x  y
# 0  1  4
# 1  2  5
# 2  3  6

Verifying the Concatenated Result

After concatenating, always verify the result has the expected shape and no unexpected NaN values. A quick sanity check pattern: compare the combined row count to the sum of individual counts, and call isna().sum() to spot any unintended missing values introduced by column misalignment. These checks prevent silent data quality issues from propagating into your analysis.

combined = pd.concat([jan, feb], ignore_index=True)

# Sanity checks
assert len(combined) == len(jan) + len(feb), 'Row count mismatch'
print('Missing values per column:')
print(combined.isna().sum())
print('Shape:', combined.shape)

concat vs append (deprecated)

Older Pandas code may use df.append(other), which was a convenience wrapper around pd.concat(). This method was deprecated in Pandas 1.4 and removed in Pandas 2.0. Always use pd.concat([df, other], ignore_index=True) in modern code. The behaviour is identical but pd.concat is more explicit and supports concatenating more than two DataFrames at once.

# Old (removed in Pandas 2.0):
# combined = jan.append(feb, ignore_index=True)

# Modern equivalent:
combined = pd.concat([jan, feb], ignore_index=True)
print(combined)

Practical Example: Yearly Sales Report

Here is a realistic pattern: load monthly DataFrames, add a source label, concatenate, and compute a full-year summary. The keys parameter makes it easy to trace each row back to its month, and ignore_index=True combined with a month column gives a clean, flat DataFrame for group-by analysis.

months = ['jan', 'feb', 'mar']
frames = []
for m in months:
    df = pd.DataFrame({'product': ['A', 'B'], 'sales': [100, 200]})
    df['month'] = m
    frames.append(df)

yearly = pd.concat(frames, ignore_index=True)
print(yearly.groupby('month')['sales'].sum())

Quick Check

Test your understanding of pd.concat for stacking DataFrames from this lesson.

Lesson Recap

In this lesson you learned: how to stack DataFrames vertically with pd.concat() and ignore_index=True, how to track sources using the keys parameter, and how join='inner' vs 'outer' controls column handling when schemas differ. Next up we explore pd.merge() for SQL-style inner and outer joins on shared key columns.

Часто задаваемые вопросы

Урок «pd.concat для объединения DataFrames» бесплатный?

Да — полный текст урока «pd.concat для объединения DataFrames» бесплатно доступен здесь в веб-версии. Чтобы практиковать его интерактивно (встроенный редактор кода и ИИ-репетитор 24/7) и разблокировать остальной курс Pandas & NumPy Academy, подпишись на CoddyKit PRO. Курс Pandas & NumPy Academy содержит 4 уроков всего.

Чему я научусь в уроке «pd.concat для объединения DataFrames»?

Объединяйте DataFrames вертикально или горизонтально с помощью pd.concat, выравнивая их по индексу или столбцам и обрабатывая дублирующиеся индексы. Ты практикуешь Pandas & NumPy Academy с помощью реального кода, который запускаешь прямо в браузере, и ИИ-репетитор 24/7 отвечает на твои вопросы во время урока.

Нужен ли мне опыт, чтобы начать Pandas & NumPy Academy?

Предыдущий опыт не требуется. Pandas & NumPy Academy на CoddyKit структурирован для всех уровней — от новичков до продвинутых, поэтому ты можешь начать отсюда или с самого начала и учиться в своем темпе. Это урок 1 из 4.

Сколько времени занимает урок «pd.concat для объединения DataFrames»?

Большинство уроков CoddyKit занимают около 5–10 минут. Каждый из них компактный и интерактивный, поэтому ты постоянно делаешь прогресс и продолжаешь с того же места в веб-версии и приложении.

Можно ли писать и запускать код в этом уроке Pandas & NumPy Academy?

Да. Каждый урок Pandas & NumPy Academy включает встроенный редактор кода, поэтому ты пишешь и запускаешь реальный код прямо в браузере и получаешь моментальную обратную связь от AI — локальная установка не требуется.

Все уроки этого курса

  1. pd.concat для объединения DataFrames
  2. pd.merge: внутреннее и внешнее объединение
  3. Левое и правое объединение
  4. Объединение по индексу
← Назад к Pandas & NumPy Academy