0Pricing
Pandas & NumPy Academy · Lesson

pd.concat for Stacking DataFrames

Stack DataFrames vertically or horizontally with pd.concat, align on index or columns, and handle duplicate indices.

pd.concat for Stacking DataFrames is a free Pandas & NumPy Academy lesson on CoddyKit — lesson 1 of 4. You can read the complete lesson below for free — then practise it hands-on in the browser with a built-in code editor and a 24/7 AI tutor. It is part of the Pandas & NumPy Academy learning path, one of 4 lessons in the course, and your progress syncs across the web and the CoddyKit app.

Why Concatenate DataFrames?

In real projects, data rarely arrives in a single file. You might have monthly sales files, regional survey exports, or database query results split across multiple queries. Concatenation stacks these separate DataFrames into one combined table. Pandas provides pd.concat() for this purpose, handling both vertical (row-wise) and horizontal (column-wise) stacking.

Basic Vertical Concatenation

The most common use of pd.concat() is stacking DataFrames vertically (adding more rows). Pass a list of DataFrames and the function appends them one below the other. Both DataFrames must have compatible columns for a clean result. Pandas aligns on column names automatically, filling missing columns with NaN.

import pandas as pd

jan = pd.DataFrame({'product': ['A', 'B'], 'sales': [100, 200]})
feb = pd.DataFrame({'product': ['A', 'C'], 'sales': [150, 120]})

combined = pd.concat([jan, feb])
print(combined)
#   product  sales
# 0       A    100
# 1       B    200
# 0       A    150
# 1       C    120

Resetting the Index After Concat

Notice that pd.concat() preserves the original indices from each DataFrame, which can result in duplicate index values (as seen above with two rows at index 0 and index 1). Pass ignore_index=True to create a fresh sequential integer index in the combined result, which is almost always what you want.

combined = pd.concat([jan, feb], ignore_index=True)
print(combined)
#   product  sales
# 0       A    100
# 1       B    200
# 2       A    150
# 3       C    120

print(combined.index)
# RangeIndex(start=0, stop=4, step=1)

Tracking Source with keys

When concatenating data from multiple sources, you may want to know which original DataFrame each row came from. Pass the keys parameter with a list of labels. Pandas creates a MultiIndex where the outer level identifies the source. You can then use .loc['label'] to access rows from a specific source.

combined = pd.concat([jan, feb], keys=['January', 'February'])
print(combined)
#                product  sales
# January   0         A    100
#           1         B    200
# February  0         A    150
#           1         C    120

# Access only February rows
print(combined.loc['February'])

Horizontal Concatenation with axis=1

Pass axis=1 to concatenate DataFrames side by side (adding more columns). Pandas aligns on the row index, so both DataFrames should have the same index for a clean result. If indices differ, non-matching rows will be filled with NaN. This is useful for combining features computed from the same dataset in separate steps.

names = pd.DataFrame({'name': ['Alice', 'Bob', 'Carol']}, index=[1, 2, 3])
scores = pd.DataFrame({'score': [95, 88, 72]}, index=[1, 2, 3])

combined = pd.concat([names, scores], axis=1)
print(combined)
#     name  score
# 1  Alice     95
# 2    Bob     88
# 3  Carol     72

Handling Mismatched Columns

If the DataFrames being concatenated have different columns, Pandas keeps all columns and fills missing values with NaN. This is the default join='outer' behaviour. If you want to keep only the columns that appear in all DataFrames, pass join='inner', which drops any column not present in every source.

df1 = pd.DataFrame({'a': [1, 2], 'b': [3, 4]})
df2 = pd.DataFrame({'b': [5, 6], 'c': [7, 8]})

# Outer join (default): keeps all columns
print(pd.concat([df1, df2], ignore_index=True))
#      a  b    c
# 0  1.0  3  NaN
# 1  2.0  4  NaN
# 2  NaN  5  7.0
# 3  NaN  6  8.0

# Inner join: keeps only shared columns
print(pd.concat([df1, df2], join='inner', ignore_index=True))
#    b
# 0  3
# 1  4
# 2  5
# 3  6

Concatenating Many Files in a Loop

A typical pattern when loading multiple files is to collect all DataFrames in a list and call pd.concat() once at the end. Avoid concatenating inside a loop (e.g., df = pd.concat([df, new_chunk])) because this creates a new DataFrame copy on every iteration, leading to quadratic time complexity for large collections.

import glob

# Efficient: collect first, concat once
files = glob.glob('data/sales_*.csv')
frames = [pd.read_csv(f) for f in files]
combined = pd.concat(frames, ignore_index=True)

# Inefficient (avoid):
# result = pd.DataFrame()
# for f in files:
#     result = pd.concat([result, pd.read_csv(f)])  # slow!

Concatenating Series

pd.concat() works with Series as well as DataFrames. Concatenating a list of Series vertically gives a longer Series. Concatenating with axis=1 produces a DataFrame where each Series becomes a column. The Series are aligned on their index, so the index labels must match for a clean horizontal concat.

s1 = pd.Series([1, 2, 3], name='x')
s2 = pd.Series([4, 5, 6], name='y')

# Vertical: one long Series
print(pd.concat([s1, s2]))
# 0    1
# 1    2
# ...

# Horizontal: a DataFrame with two columns
print(pd.concat([s1, s2], axis=1))
#    x  y
# 0  1  4
# 1  2  5
# 2  3  6

Verifying the Concatenated Result

After concatenating, always verify the result has the expected shape and no unexpected NaN values. A quick sanity check pattern: compare the combined row count to the sum of individual counts, and call isna().sum() to spot any unintended missing values introduced by column misalignment. These checks prevent silent data quality issues from propagating into your analysis.

combined = pd.concat([jan, feb], ignore_index=True)

# Sanity checks
assert len(combined) == len(jan) + len(feb), 'Row count mismatch'
print('Missing values per column:')
print(combined.isna().sum())
print('Shape:', combined.shape)

concat vs append (deprecated)

Older Pandas code may use df.append(other), which was a convenience wrapper around pd.concat(). This method was deprecated in Pandas 1.4 and removed in Pandas 2.0. Always use pd.concat([df, other], ignore_index=True) in modern code. The behaviour is identical but pd.concat is more explicit and supports concatenating more than two DataFrames at once.

# Old (removed in Pandas 2.0):
# combined = jan.append(feb, ignore_index=True)

# Modern equivalent:
combined = pd.concat([jan, feb], ignore_index=True)
print(combined)

Practical Example: Yearly Sales Report

Here is a realistic pattern: load monthly DataFrames, add a source label, concatenate, and compute a full-year summary. The keys parameter makes it easy to trace each row back to its month, and ignore_index=True combined with a month column gives a clean, flat DataFrame for group-by analysis.

months = ['jan', 'feb', 'mar']
frames = []
for m in months:
    df = pd.DataFrame({'product': ['A', 'B'], 'sales': [100, 200]})
    df['month'] = m
    frames.append(df)

yearly = pd.concat(frames, ignore_index=True)
print(yearly.groupby('month')['sales'].sum())

Quick Check

Test your understanding of pd.concat for stacking DataFrames from this lesson.

Lesson Recap

In this lesson you learned: how to stack DataFrames vertically with pd.concat() and ignore_index=True, how to track sources using the keys parameter, and how join='inner' vs 'outer' controls column handling when schemas differ. Next up we explore pd.merge() for SQL-style inner and outer joins on shared key columns.

Frequently asked questions

Is the “pd.concat for Stacking DataFrames” lesson free?

Yes — the full text of “pd.concat for Stacking DataFrames” is free to read here on the web, and the Pandas & NumPy Academy course includes 4 lessons in total. To practise it interactively (a built-in code editor and a 24/7 AI tutor) and unlock the rest of the Pandas & NumPy Academy course, upgrade to CoddyKit PRO.

What will I learn in “pd.concat for Stacking DataFrames”?

Stack DataFrames vertically or horizontally with pd.concat, align on index or columns, and handle duplicate indices. You practise Pandas & NumPy Academy with hands-on code you run directly in the browser, and a 24/7 AI tutor answers your questions as you work through the lesson.

Do I need any experience to start Pandas & NumPy Academy?

No prior experience is required. Pandas & NumPy Academy on CoddyKit is structured for beginners through advanced learners; this is — lesson 1 of 4, so you can start here or from the beginning and move at your own pace.

How long does the “pd.concat for Stacking DataFrames” lesson take?

Most CoddyKit lessons take about 5–10 minutes. Each one is bite-sized and interactive, so you make steady progress and pick up exactly where you left off across the web and the app.

Can I write and run code in this Pandas & NumPy Academy lesson?

Yes. Every Pandas & NumPy Academy lesson includes a built-in code editor, so you write and run real code right in your browser and get instant AI feedback — no local setup required.

All lessons in this course

  1. pd.concat for Stacking DataFrames
  2. pd.merge: Inner and Outer Joins
  3. Left and Right Joins
  4. Joining on the Index
← Back to Pandas & NumPy Academy