0Pricing
Pandas & NumPy Academy · Lektion

pd.concat zum Stapeln von DataFrames

Stapeln Sie DataFrames mit pd.concat vertikal oder horizontal, richten Sie sie am Index oder an Spalten aus und behandeln Sie doppelte Indizes.

pd.concat zum Stapeln von DataFrames ist eine kostenlose Pandas & NumPy Academy-Lektion auf CoddyKit. Dies ist Lektion 1 von 4. Du kannst die komplette Lektion unten kostenlos lesen – dann übst du sie direkt im Browser mit einem integrierten Code-Editor und einem KI-Tutor rund um die Uhr. Sie ist Teil des Pandas & NumPy Academy-Lernpfads, und dein Fortschritt wird über Web und CoddyKit-App synchronisiert. Der Pandas & NumPy Academy-Kurs umfasst insgesamt 4 Lektionen.

Teile dieser Lektion wurden noch nicht übersetzt und werden auf Englisch angezeigt.

Why Concatenate DataFrames?

In real projects, data rarely arrives in a single file. You might have monthly sales files, regional survey exports, or database query results split across multiple queries. Concatenation stacks these separate DataFrames into one combined table. Pandas provides pd.concat() for this purpose, handling both vertical (row-wise) and horizontal (column-wise) stacking.

Basic Vertical Concatenation

The most common use of pd.concat() is stacking DataFrames vertically (adding more rows). Pass a list of DataFrames and the function appends them one below the other. Both DataFrames must have compatible columns for a clean result. Pandas aligns on column names automatically, filling missing columns with NaN.

import pandas as pd

jan = pd.DataFrame({'product': ['A', 'B'], 'sales': [100, 200]})
feb = pd.DataFrame({'product': ['A', 'C'], 'sales': [150, 120]})

combined = pd.concat([jan, feb])
print(combined)
#   product  sales
# 0       A    100
# 1       B    200
# 0       A    150
# 1       C    120

Resetting the Index After Concat

Notice that pd.concat() preserves the original indices from each DataFrame, which can result in duplicate index values (as seen above with two rows at index 0 and index 1). Pass ignore_index=True to create a fresh sequential integer index in the combined result, which is almost always what you want.

combined = pd.concat([jan, feb], ignore_index=True)
print(combined)
#   product  sales
# 0       A    100
# 1       B    200
# 2       A    150
# 3       C    120

print(combined.index)
# RangeIndex(start=0, stop=4, step=1)

Tracking Source with keys

When concatenating data from multiple sources, you may want to know which original DataFrame each row came from. Pass the keys parameter with a list of labels. Pandas creates a MultiIndex where the outer level identifies the source. You can then use .loc['label'] to access rows from a specific source.

combined = pd.concat([jan, feb], keys=['January', 'February'])
print(combined)
#                product  sales
# January   0         A    100
#           1         B    200
# February  0         A    150
#           1         C    120

# Access only February rows
print(combined.loc['February'])

Horizontal Concatenation with axis=1

Pass axis=1 to concatenate DataFrames side by side (adding more columns). Pandas aligns on the row index, so both DataFrames should have the same index for a clean result. If indices differ, non-matching rows will be filled with NaN. This is useful for combining features computed from the same dataset in separate steps.

names = pd.DataFrame({'name': ['Alice', 'Bob', 'Carol']}, index=[1, 2, 3])
scores = pd.DataFrame({'score': [95, 88, 72]}, index=[1, 2, 3])

combined = pd.concat([names, scores], axis=1)
print(combined)
#     name  score
# 1  Alice     95
# 2    Bob     88
# 3  Carol     72

Handling Mismatched Columns

If the DataFrames being concatenated have different columns, Pandas keeps all columns and fills missing values with NaN. This is the default join='outer' behaviour. If you want to keep only the columns that appear in all DataFrames, pass join='inner', which drops any column not present in every source.

df1 = pd.DataFrame({'a': [1, 2], 'b': [3, 4]})
df2 = pd.DataFrame({'b': [5, 6], 'c': [7, 8]})

# Outer join (default): keeps all columns
print(pd.concat([df1, df2], ignore_index=True))
#      a  b    c
# 0  1.0  3  NaN
# 1  2.0  4  NaN
# 2  NaN  5  7.0
# 3  NaN  6  8.0

# Inner join: keeps only shared columns
print(pd.concat([df1, df2], join='inner', ignore_index=True))
#    b
# 0  3
# 1  4
# 2  5
# 3  6

Concatenating Many Files in a Loop

A typical pattern when loading multiple files is to collect all DataFrames in a list and call pd.concat() once at the end. Avoid concatenating inside a loop (e.g., df = pd.concat([df, new_chunk])) because this creates a new DataFrame copy on every iteration, leading to quadratic time complexity for large collections.

import glob

# Efficient: collect first, concat once
files = glob.glob('data/sales_*.csv')
frames = [pd.read_csv(f) for f in files]
combined = pd.concat(frames, ignore_index=True)

# Inefficient (avoid):
# result = pd.DataFrame()
# for f in files:
#     result = pd.concat([result, pd.read_csv(f)])  # slow!

Concatenating Series

pd.concat() works with Series as well as DataFrames. Concatenating a list of Series vertically gives a longer Series. Concatenating with axis=1 produces a DataFrame where each Series becomes a column. The Series are aligned on their index, so the index labels must match for a clean horizontal concat.

s1 = pd.Series([1, 2, 3], name='x')
s2 = pd.Series([4, 5, 6], name='y')

# Vertical: one long Series
print(pd.concat([s1, s2]))
# 0    1
# 1    2
# ...

# Horizontal: a DataFrame with two columns
print(pd.concat([s1, s2], axis=1))
#    x  y
# 0  1  4
# 1  2  5
# 2  3  6

Verifying the Concatenated Result

After concatenating, always verify the result has the expected shape and no unexpected NaN values. A quick sanity check pattern: compare the combined row count to the sum of individual counts, and call isna().sum() to spot any unintended missing values introduced by column misalignment. These checks prevent silent data quality issues from propagating into your analysis.

combined = pd.concat([jan, feb], ignore_index=True)

# Sanity checks
assert len(combined) == len(jan) + len(feb), 'Row count mismatch'
print('Missing values per column:')
print(combined.isna().sum())
print('Shape:', combined.shape)

concat vs append (deprecated)

Older Pandas code may use df.append(other), which was a convenience wrapper around pd.concat(). This method was deprecated in Pandas 1.4 and removed in Pandas 2.0. Always use pd.concat([df, other], ignore_index=True) in modern code. The behaviour is identical but pd.concat is more explicit and supports concatenating more than two DataFrames at once.

# Old (removed in Pandas 2.0):
# combined = jan.append(feb, ignore_index=True)

# Modern equivalent:
combined = pd.concat([jan, feb], ignore_index=True)
print(combined)

Practical Example: Yearly Sales Report

Here is a realistic pattern: load monthly DataFrames, add a source label, concatenate, and compute a full-year summary. The keys parameter makes it easy to trace each row back to its month, and ignore_index=True combined with a month column gives a clean, flat DataFrame for group-by analysis.

months = ['jan', 'feb', 'mar']
frames = []
for m in months:
    df = pd.DataFrame({'product': ['A', 'B'], 'sales': [100, 200]})
    df['month'] = m
    frames.append(df)

yearly = pd.concat(frames, ignore_index=True)
print(yearly.groupby('month')['sales'].sum())

Quick Check

Test your understanding of pd.concat for stacking DataFrames from this lesson.

Lesson Recap

In this lesson you learned: how to stack DataFrames vertically with pd.concat() and ignore_index=True, how to track sources using the keys parameter, and how join='inner' vs 'outer' controls column handling when schemas differ. Next up we explore pd.merge() for SQL-style inner and outer joins on shared key columns.

Häufig gestellte Fragen

Ist die Lektion „pd.concat zum Stapeln von DataFrames“ kostenlos?

Ja — der vollständige Text von „pd.concat zum Stapeln von DataFrames“ ist hier im Web kostenlos zu lesen. Um sie interaktiv zu üben (integrierter Code-Editor und 24/7 KI-Tutor) und den Rest des Pandas & NumPy Academy-Kurses freizuschalten, upgrade auf CoddyKit PRO. Der Pandas & NumPy Academy-Kurs umfasst insgesamt 4 Lektionen.

Was lerne ich in „pd.concat zum Stapeln von DataFrames“?

Stapeln Sie DataFrames mit pd.concat vertikal oder horizontal, richten Sie sie am Index oder an Spalten aus und behandeln Sie doppelte Indizes. Du übst Pandas & NumPy Academy mit praktischem Code, den du direkt im Browser ausführst, und ein 24/7 KI-Tutor beantwortet deine Fragen während du die Lektion bearbeitest.

Brauche ich Erfahrung, um Pandas & NumPy Academy zu starten?

Keine Vorkenntnisse erforderlich. Pandas & NumPy Academy auf CoddyKit ist für Anfänger bis fortgeschrittene Lernende strukturiert, sodass du hier starten oder von Anfang an beginnen und in deinem eigenen Tempo voranschreiten kannst. Dies ist Lektion 1 von 4.

Wie lange dauert die Lektion „pd.concat zum Stapeln von DataFrames“?

Die meisten CoddyKit-Lektionen dauern etwa 5–10 Minuten. Jede ist kompakt und interaktiv, sodass du stetig Fortschritte machst und genau dort weitermachst, wo du aufgehört hast – im Web und in der App.

Kann ich in dieser Pandas & NumPy Academy-Lektion Code schreiben und ausführen?

Ja. Jede Pandas & NumPy Academy-Lektion enthält einen integrierten Code-Editor, sodass du echten Code direkt in deinem Browser schreibst und ausführst und sofort KI-Feedback erhältst — ohne lokale Einrichtung erforderlich.

Alle Lektionen in diesem Kurs

  1. pd.concat zum Stapeln von DataFrames
  2. pd.merge: Inner- und Outer-Joins
  3. Left- und Right-Joins
  4. Über den Index verknüpfen
← Zurück zu Pandas & NumPy Academy