GroupBy mit einem und mehreren Schlüsseln
Gruppieren Sie mit groupby() nach einer oder mehreren Spalten und wenden Sie die Aggregationen sum, mean, count sowie min/max an.
GroupBy mit einem und mehreren Schlüsseln ist eine kostenlose Pandas & NumPy Academy-Lektion auf CoddyKit. Dies ist Lektion 2 von 4. Du kannst die komplette Lektion unten kostenlos lesen – dann übst du sie direkt im Browser mit einem integrierten Code-Editor und einem KI-Tutor rund um die Uhr. Sie ist Teil des Pandas & NumPy Academy-Lernpfads, und dein Fortschritt wird über Web und CoddyKit-App synchronisiert. Der Pandas & NumPy Academy-Kurs umfasst insgesamt 4 Lektionen.
Teile dieser Lektion wurden noch nicht übersetzt und werden auf Englisch angezeigt.
GroupBy with a Single Key
The simplest GroupBy call uses a single column as the grouping key. You call df.groupby('column_name') and chain an aggregation. Pandas creates one group for each unique value in that column and applies the aggregation to every numeric column (or the selected column). This is the most common form of GroupBy in day-to-day analysis.
import pandas as pd
df = pd.DataFrame({
'dept': ['Eng', 'HR', 'Eng', 'HR', 'Eng'],
'salary': [90000, 60000, 95000, 62000, 88000],
'years': [3, 5, 7, 2, 4]
})
print(df.groupby('dept')['salary'].mean())
# dept
# Eng 91000.0
# HR 61000.0Choosing Which Columns to Aggregate
After calling groupby() you can select one column with bracket notation to get a SeriesGroupBy, or select multiple columns with a list to get a DataFrameGroupBy. If you skip selection entirely, Pandas aggregates all numeric columns, which is convenient but can produce unexpected results when you have irrelevant numeric columns.
# Single column result -> SeriesGroupBy
print(df.groupby('dept')['salary'].sum())
# Multiple columns result -> DataFrameGroupBy
print(df.groupby('dept')[['salary', 'years']].mean())
# salary years
# dept
# Eng 91000.0 4.667
# HR 61000.0 3.500All Common Aggregations
Pandas supports a full set of built-in aggregation methods on GroupBy objects: sum(), mean(), median(), min(), max(), count(), std(), var(), first(), and last(). Each returns a scalar per group, giving you a clean summary table indexed by the grouping key.
g = df.groupby('dept')['salary']
print('sum: ', g.sum())
print('mean: ', g.mean())
print('min: ', g.min())
print('max: ', g.max())
print('count: ', g.count())
print('std: ', g.std().round(2))GroupBy with Multiple Keys
Pass a list of column names to groupby() to group by more than one dimension at once. Each unique combination of values in those columns becomes its own group. The result has a MultiIndex with one level per grouping column, allowing you to drill into cross-dimensional summaries in a single step.
df2 = pd.DataFrame({
'dept': ['Eng', 'Eng', 'HR', 'HR', 'Eng', 'HR'],
'level': ['L1', 'L2', 'L1', 'L2', 'L1', 'L1'],
'salary': [80000, 100000, 55000, 70000, 82000, 58000]
})
result = df2.groupby(['dept', 'level'])['salary'].mean()
print(result)
# dept level
# Eng L1 81000.0
# L2 100000.0
# HR L1 56500.0
# L2 70000.0Accessing MultiIndex Results
When you group by multiple keys the result index is a MultiIndex. You can access a specific outer level with .loc['value'] and navigate inner levels with tuples. Calling reset_index() flattens the MultiIndex into regular columns, which is often more convenient for further operations or export.
result = df2.groupby(['dept', 'level'])['salary'].mean()
# Access Engineering rows only
print(result.loc['Eng'])
# level
# L1 81000.0
# L2 100000.0
# Flatten to a regular DataFrame
print(result.reset_index())
# dept level salary
# 0 Eng L1 81000.0
# 1 Eng L2 100000.0Using as_index=False
By default, the grouping columns become the index of the result. Passing as_index=False keeps them as regular columns instead, giving you a flat DataFrame that is easier to pass to further Pandas operations or to export. This is equivalent to calling reset_index() on the result.
flat = df2.groupby(['dept', 'level'], as_index=False)['salary'].mean()
print(flat)
# dept level salary
# 0 Eng L1 81000.0
# 1 Eng L2 100000.0
# 2 HR L1 56500.0
# 3 HR L2 70000.0Counting Rows Per Group
count() returns the number of non-null values in each group. If you want the total number of rows per group regardless of nulls, use size() instead. This distinction matters when your data has missing values, because count() will undercount groups that have NaN entries.
import numpy as np
df3 = pd.DataFrame({
'dept': ['Eng', 'Eng', 'HR', 'HR'],
'bonus': [5000, np.nan, 3000, 4000]
})
print(df3.groupby('dept')['bonus'].count()) # ignores NaN
# dept
# Eng 1
# HR 2
print(df3.groupby('dept')['bonus'].size()) # includes NaN rows
# dept
# Eng 2
# HR 2Sorting GroupBy Results
By default, GroupBy sorts the result by the group key. You can disable sorting with sort=False to preserve the original order of first appearance, which is slightly faster on large datasets. Once you have the result as a DataFrame, you can sort it further with sort_values() on any column.
result = (df.groupby('dept', sort=False)['salary']
.mean()
.reset_index()
.sort_values('salary', ascending=False))
print(result)
# dept salary
# 0 Eng 91000.0
# 1 HR 61000.0Chaining GroupBy with Other Methods
GroupBy results are regular Series or DataFrames, so you can chain further Pandas methods on them immediately. A common pattern is to aggregate, then reset_index(), then rename columns, then sort_values() — all in a single readable method chain without storing intermediate variables.
summary = (
df
.groupby('dept')['salary']
.agg(['mean', 'count'])
.rename(columns={'mean': 'avg_salary', 'count': 'headcount'})
.reset_index()
.sort_values('avg_salary', ascending=False)
)
print(summary)Applying sum, mean, count Together
You frequently need more than one statistic per group. Pass a list of function names to .agg() to compute several at once. The result is a DataFrame with a column for each function. This is more efficient than calling each aggregation separately because Pandas processes all of them in a single pass over the data.
result = df.groupby('dept')['salary'].agg(['sum', 'mean', 'count', 'max'])
print(result)
# sum mean count max
# dept
# Eng 273000 91000.0 3 95000
# HR 122000 61000.0 2 62000Grouping by Categorical Columns
When a grouping column has the Categorical dtype, Pandas includes all categories in the result by default, even those with no rows. This can be useful to ensure your summary table always shows every category, but it creates rows with NaN or 0 for empty groups. Control this with the observed parameter (set to True to skip empty categories).
df['dept'] = df['dept'].astype('category')
# With observed=True, only groups that appear in data are included
print(df.groupby('dept', observed=True)['salary'].mean())
# dept
# Eng 91000.0
# HR 61000.0Quick Check
Test your understanding of GroupBy with single and multiple keys from this lesson.
Lesson Recap
In this lesson you learned: how to group by a single column and apply common aggregations, how to group by multiple columns to produce cross-dimensional summaries, and the difference between count() and size() when nulls are present. Next up we explore the powerful agg() method for computing multiple different statistics in a single call.
Häufig gestellte Fragen
Ist die Lektion „GroupBy mit einem und mehreren Schlüsseln“ kostenlos?
Ja — der vollständige Text von „GroupBy mit einem und mehreren Schlüsseln“ ist hier im Web kostenlos zu lesen. Um sie interaktiv zu üben (integrierter Code-Editor und 24/7 KI-Tutor) und den Rest des Pandas & NumPy Academy-Kurses freizuschalten, upgrade auf CoddyKit PRO. Der Pandas & NumPy Academy-Kurs umfasst insgesamt 4 Lektionen.
Was lerne ich in „GroupBy mit einem und mehreren Schlüsseln“?
Gruppieren Sie mit groupby() nach einer oder mehreren Spalten und wenden Sie die Aggregationen sum, mean, count sowie min/max an. Du übst Pandas & NumPy Academy mit praktischem Code, den du direkt im Browser ausführst, und ein 24/7 KI-Tutor beantwortet deine Fragen während du die Lektion bearbeitest.
Brauche ich Erfahrung, um Pandas & NumPy Academy zu starten?
Keine Vorkenntnisse erforderlich. Pandas & NumPy Academy auf CoddyKit ist für Anfänger bis fortgeschrittene Lernende strukturiert, sodass du hier starten oder von Anfang an beginnen und in deinem eigenen Tempo voranschreiten kannst. Dies ist Lektion 2 von 4.
Wie lange dauert die Lektion „GroupBy mit einem und mehreren Schlüsseln“?
Die meisten CoddyKit-Lektionen dauern etwa 5–10 Minuten. Jede ist kompakt und interaktiv, sodass du stetig Fortschritte machst und genau dort weitermachst, wo du aufgehört hast – im Web und in der App.
Kann ich in dieser Pandas & NumPy Academy-Lektion Code schreiben und ausführen?
Ja. Jede Pandas & NumPy Academy-Lektion enthält einen integrierten Code-Editor, sodass du echten Code direkt in deinem Browser schreibst und ausführst und sofort KI-Feedback erhältst — ohne lokale Einrichtung erforderlich.
Alle Lektionen in diesem Kurs
- Das Split-Apply-Combine-Muster
- GroupBy mit einem und mehreren Schlüsseln
- Die Methode agg()
- GroupBy-Transformation und -Filterung