0Pricing
Pandas & NumPy Academy · レッスン

DataFrameの結合に使うpd.concat

pd.concatでDataFrameを縦方向または横方向に積み重ね、インデックスや列に揃えて結合し、重複インデックスを処理します。

「DataFrameの結合に使うpd.concat」はCoddyKit上の無料Pandas & NumPy Academyレッスンです。 これはレッスン1/4です。 下記で完全なレッスンを無料で読むことができます。その後、ブラウザ内の組み込みコードエディタと24時間対応のAIチューターでハンズオン演習できます。 これはPandas & NumPy Academy学習パスの一部であり、ウェブとCoddyKitアプリ全体で進捗が同期されます。 Pandas & NumPy Academyコースには全4レッスンが含まれています。

このレッスンの一部はまだ翻訳されておらず、英語で表示されています。

Why Concatenate DataFrames?

In real projects, data rarely arrives in a single file. You might have monthly sales files, regional survey exports, or database query results split across multiple queries. Concatenation stacks these separate DataFrames into one combined table. Pandas provides pd.concat() for this purpose, handling both vertical (row-wise) and horizontal (column-wise) stacking.

Basic Vertical Concatenation

The most common use of pd.concat() is stacking DataFrames vertically (adding more rows). Pass a list of DataFrames and the function appends them one below the other. Both DataFrames must have compatible columns for a clean result. Pandas aligns on column names automatically, filling missing columns with NaN.

import pandas as pd

jan = pd.DataFrame({'product': ['A', 'B'], 'sales': [100, 200]})
feb = pd.DataFrame({'product': ['A', 'C'], 'sales': [150, 120]})

combined = pd.concat([jan, feb])
print(combined)
#   product  sales
# 0       A    100
# 1       B    200
# 0       A    150
# 1       C    120

Resetting the Index After Concat

Notice that pd.concat() preserves the original indices from each DataFrame, which can result in duplicate index values (as seen above with two rows at index 0 and index 1). Pass ignore_index=True to create a fresh sequential integer index in the combined result, which is almost always what you want.

combined = pd.concat([jan, feb], ignore_index=True)
print(combined)
#   product  sales
# 0       A    100
# 1       B    200
# 2       A    150
# 3       C    120

print(combined.index)
# RangeIndex(start=0, stop=4, step=1)

Tracking Source with keys

When concatenating data from multiple sources, you may want to know which original DataFrame each row came from. Pass the keys parameter with a list of labels. Pandas creates a MultiIndex where the outer level identifies the source. You can then use .loc['label'] to access rows from a specific source.

combined = pd.concat([jan, feb], keys=['January', 'February'])
print(combined)
#                product  sales
# January   0         A    100
#           1         B    200
# February  0         A    150
#           1         C    120

# Access only February rows
print(combined.loc['February'])

Horizontal Concatenation with axis=1

Pass axis=1 to concatenate DataFrames side by side (adding more columns). Pandas aligns on the row index, so both DataFrames should have the same index for a clean result. If indices differ, non-matching rows will be filled with NaN. This is useful for combining features computed from the same dataset in separate steps.

names = pd.DataFrame({'name': ['Alice', 'Bob', 'Carol']}, index=[1, 2, 3])
scores = pd.DataFrame({'score': [95, 88, 72]}, index=[1, 2, 3])

combined = pd.concat([names, scores], axis=1)
print(combined)
#     name  score
# 1  Alice     95
# 2    Bob     88
# 3  Carol     72

Handling Mismatched Columns

If the DataFrames being concatenated have different columns, Pandas keeps all columns and fills missing values with NaN. This is the default join='outer' behaviour. If you want to keep only the columns that appear in all DataFrames, pass join='inner', which drops any column not present in every source.

df1 = pd.DataFrame({'a': [1, 2], 'b': [3, 4]})
df2 = pd.DataFrame({'b': [5, 6], 'c': [7, 8]})

# Outer join (default): keeps all columns
print(pd.concat([df1, df2], ignore_index=True))
#      a  b    c
# 0  1.0  3  NaN
# 1  2.0  4  NaN
# 2  NaN  5  7.0
# 3  NaN  6  8.0

# Inner join: keeps only shared columns
print(pd.concat([df1, df2], join='inner', ignore_index=True))
#    b
# 0  3
# 1  4
# 2  5
# 3  6

Concatenating Many Files in a Loop

A typical pattern when loading multiple files is to collect all DataFrames in a list and call pd.concat() once at the end. Avoid concatenating inside a loop (e.g., df = pd.concat([df, new_chunk])) because this creates a new DataFrame copy on every iteration, leading to quadratic time complexity for large collections.

import glob

# Efficient: collect first, concat once
files = glob.glob('data/sales_*.csv')
frames = [pd.read_csv(f) for f in files]
combined = pd.concat(frames, ignore_index=True)

# Inefficient (avoid):
# result = pd.DataFrame()
# for f in files:
#     result = pd.concat([result, pd.read_csv(f)])  # slow!

Concatenating Series

pd.concat() works with Series as well as DataFrames. Concatenating a list of Series vertically gives a longer Series. Concatenating with axis=1 produces a DataFrame where each Series becomes a column. The Series are aligned on their index, so the index labels must match for a clean horizontal concat.

s1 = pd.Series([1, 2, 3], name='x')
s2 = pd.Series([4, 5, 6], name='y')

# Vertical: one long Series
print(pd.concat([s1, s2]))
# 0    1
# 1    2
# ...

# Horizontal: a DataFrame with two columns
print(pd.concat([s1, s2], axis=1))
#    x  y
# 0  1  4
# 1  2  5
# 2  3  6

Verifying the Concatenated Result

After concatenating, always verify the result has the expected shape and no unexpected NaN values. A quick sanity check pattern: compare the combined row count to the sum of individual counts, and call isna().sum() to spot any unintended missing values introduced by column misalignment. These checks prevent silent data quality issues from propagating into your analysis.

combined = pd.concat([jan, feb], ignore_index=True)

# Sanity checks
assert len(combined) == len(jan) + len(feb), 'Row count mismatch'
print('Missing values per column:')
print(combined.isna().sum())
print('Shape:', combined.shape)

concat vs append (deprecated)

Older Pandas code may use df.append(other), which was a convenience wrapper around pd.concat(). This method was deprecated in Pandas 1.4 and removed in Pandas 2.0. Always use pd.concat([df, other], ignore_index=True) in modern code. The behaviour is identical but pd.concat is more explicit and supports concatenating more than two DataFrames at once.

# Old (removed in Pandas 2.0):
# combined = jan.append(feb, ignore_index=True)

# Modern equivalent:
combined = pd.concat([jan, feb], ignore_index=True)
print(combined)

Practical Example: Yearly Sales Report

Here is a realistic pattern: load monthly DataFrames, add a source label, concatenate, and compute a full-year summary. The keys parameter makes it easy to trace each row back to its month, and ignore_index=True combined with a month column gives a clean, flat DataFrame for group-by analysis.

months = ['jan', 'feb', 'mar']
frames = []
for m in months:
    df = pd.DataFrame({'product': ['A', 'B'], 'sales': [100, 200]})
    df['month'] = m
    frames.append(df)

yearly = pd.concat(frames, ignore_index=True)
print(yearly.groupby('month')['sales'].sum())

Quick Check

Test your understanding of pd.concat for stacking DataFrames from this lesson.

Lesson Recap

In this lesson you learned: how to stack DataFrames vertically with pd.concat() and ignore_index=True, how to track sources using the keys parameter, and how join='inner' vs 'outer' controls column handling when schemas differ. Next up we explore pd.merge() for SQL-style inner and outer joins on shared key columns.

よくある質問

「DataFrameの結合に使うpd.concat」レッスンは無料ですか?

はい。「DataFrameの結合に使うpd.concat」の完全なテキストはこのウェブで無料で読めます。インタラクティブに演習し(組み込みコードエディタと24時間対応のAIチューター)、Pandas & NumPy Academyコースの残りをアンロックするには、CoddyKit PROにアップグレードしてください。 Pandas & NumPy Academyコースには全4レッスンが含まれています。

「DataFrameの結合に使うpd.concat」で何を学びますか?

pd.concatでDataFrameを縦方向または横方向に積み重ね、インデックスや列に揃えて結合し、重複インデックスを処理します。 ブラウザで直接実行するハンズオンコードでPandas & NumPy Academyを演習し、24時間対応のAIチューターがレッスンを進める中での質問に答えます。

Pandas & NumPy Academyを始めるのに経験は必要ですか?

事前経験は必要ありません。CoddyKitのPandas & NumPy Academyは初級者から上級者向けに構成されているため、ここから始めるか最初から始めて、自分のペースで進むことができます。 これはレッスン1/4です。

「DataFrameの結合に使うpd.concat」レッスンにはどのくらい時間がかかりますか?

ほとんどのCoddyKitレッスンは約5~10分かかります。各レッスンはコンパクトでインタラクティブなので、着実に進歩し、ウェブとアプリ全体で正確に前回の場所から再開できます。

このPandas & NumPy Academyレッスンでコードを書いて実行できますか?

はい。すべてのPandas & NumPy Academyレッスンに組み込みコードエディタが含まれているため、ブラウザでリアルコードを書いて実行し、即座のAIフィードバックを取得できます。ローカル設定は不要です。

このコースのすべてのレッスン

  1. DataFrameの結合に使うpd.concat
  2. pd.merge:内部結合と外部結合
  3. 左結合と右結合
  4. インデックスによる結合
← Pandas & NumPy Academyに戻る