0Pricing
Pandas & NumPy Academy · Lektion

DataFrames erstellen

Erstellen Sie DataFrames aus Dictionaries von Listen, Listen von Dictionaries und NumPy-Arrays und untersuchen Sie anschließend shape, columns und dtypes.

DataFrames erstellen ist eine kostenlose Pandas & NumPy Academy-Lektion auf CoddyKit. Dies ist Lektion 1 von 4. Du kannst die komplette Lektion unten kostenlos lesen – dann übst du sie direkt im Browser mit einem integrierten Code-Editor und einem KI-Tutor rund um die Uhr. Sie ist Teil des Pandas & NumPy Academy-Lernpfads, und dein Fortschritt wird über Web und CoddyKit-App synchronisiert. Der Pandas & NumPy Academy-Kurs umfasst insgesamt 4 Lektionen.

Teile dieser Lektion wurden noch nicht übersetzt und werden auf Englisch angezeigt.

What Is a DataFrame?

A DataFrame is a two-dimensional, size-mutable, labelled data structure — essentially a table with rows and columns, similar to a spreadsheet or SQL table. Each column is a Pandas Series sharing the same row index. DataFrames can hold columns of different dtypes, making them perfect for real-world datasets that mix numbers, text, and dates.

import pandas as pd

df = pd.DataFrame({'name': ['Alice', 'Bob', 'Carol'],
                   'age': [30, 25, 35],
                   'score': [88.5, 72.0, 95.3]})
print(df)

Creating from a Dict of Lists

The most common way to construct a DataFrame is from a dict of lists: the keys become column names and the lists become column values. All lists must be the same length. The row index defaults to 0, 1, 2, ... unless you specify one with index=. This pattern mirrors how CSV data is often represented in Python.

import pandas as pd

df = pd.DataFrame({
    'city': ['Berlin', 'Paris', 'Rome'],
    'pop': [3.6, 2.2, 2.8],
    'country': ['DE', 'FR', 'IT']
})
print(df.shape)    # (3, 3)
print(df.dtypes)

Creating from a List of Dicts

You can also pass a list of dicts, where each dict represents one row. Missing keys in any dict produce NaN for that column. This format is common when consuming JSON APIs that return a list of record objects. Pandas aligns all keys across dicts to form the column set automatically.

import pandas as pd

rows = [
    {'name': 'Alice', 'age': 30, 'dept': 'Eng'},
    {'name': 'Bob',   'age': 25},              # 'dept' missing -> NaN
    {'name': 'Carol', 'age': 35, 'dept': 'HR'}
]
df = pd.DataFrame(rows)
print(df)

Creating from a NumPy Array

Passing a 2-D NumPy array creates a DataFrame with integer column names (0, 1, 2, ...) and a RangeIndex by default. Supply columns= and index= to add meaningful labels. This is the bridge between NumPy numerical computation and Pandas labelled data analysis.

import pandas as pd
import numpy as np

arr = np.array([[1, 2, 3], [4, 5, 6], [7, 8, 9]])
df = pd.DataFrame(arr,
                  columns=['x', 'y', 'z'],
                  index=['r1', 'r2', 'r3'])
print(df)

Setting a Custom Index at Creation

The index= parameter sets the row labels at creation time. Common choices include date strings, IDs, or category names that make row lookups meaningful. A well-chosen index makes .loc access intuitive and enables powerful time-series operations when the index is a DatetimeIndex.

import pandas as pd

df = pd.DataFrame(
    {'revenue': [100, 200, 150], 'costs': [80, 160, 90]},
    index=['Jan', 'Feb', 'Mar']
)
print(df)
print(df.loc['Feb'])  # select row 'Feb'

Inspecting columns, dtypes, and index

Three attributes immediately tell you the structure of a DataFrame: df.columns gives an Index of column names, df.dtypes returns a Series mapping each column to its dtype, and df.index describes the row labels. These are the first checks to run on any new DataFrame to understand what data you have before transforming it.

import pandas as pd

df = pd.DataFrame({'a': [1, 2], 'b': [3.0, 4.0], 'c': ['x', 'y']})
print(df.columns)  # Index(['a', 'b', 'c'], dtype='object')
print(df.dtypes)
# a      int64
# b    float64
# c     object
print(df.index)    # RangeIndex(start=0, stop=2, step=1)

Specifying Column Order

When constructing from a dict, column order follows dict insertion order (Python 3.7+). If you need a specific order, pass the columns= parameter with a list of column names. Any name not in the data will produce a NaN column; any name in the data but not in the list will be excluded. This lets you select and order columns at construction time.

import pandas as pd

data = {'c': [3, 6], 'a': [1, 4], 'b': [2, 5]}
df = pd.DataFrame(data, columns=['a', 'b', 'c'])
print(df.columns.tolist())  # ['a', 'b', 'c']

Creating from a Dict of Series

Passing a dict of Pandas Series aligns on the union of all indices. Where a Series is missing a label that another has, the result cell is NaN. This is the most index-aware construction method and is used when combining separately computed columns that may have different row counts or labels.

import pandas as pd

s1 = pd.Series([1, 2, 3], index=['a', 'b', 'c'])
s2 = pd.Series([10, 20], index=['b', 'c'])
df = pd.DataFrame({'col1': s1, 'col2': s2})
print(df)
#    col1  col2
# a   1.0   NaN
# b   2.0  10.0
# c   3.0  20.0

shape, len, and size

df.shape returns a tuple (nrows, ncols). len(df) returns the number of rows. df.size returns the total number of cells (rows × columns). These are the quickest sanity checks after loading data to confirm you received the expected number of records and that no columns were silently dropped.

import pandas as pd

df = pd.DataFrame({'a': range(5), 'b': range(5), 'c': range(5)})
print(df.shape)  # (5, 3)
print(len(df))   # 5
print(df.size)   # 15  (5 rows x 3 cols)

Creating an Empty DataFrame

You can create an empty DataFrame with specific column names and dtypes for use as a template or accumulator. Append rows with pd.concat([df, new_row])m. Specifying dtypes at creation avoids expensive type inference when rows are added later. An empty DataFrame is also useful as a seed in iterative data collection loops.

import pandas as pd

df = pd.DataFrame(columns=['name', 'score'])
new_row = pd.DataFrame([{'name': 'Alice', 'score': 90}])
df = pd.concat([df, new_row], ignore_index=True)
print(df)

Copying a DataFrame

Assigning a DataFrame to a new variable creates a reference, not a copy — modifying one modifies the other. Use df.copy() to create an independent copy. This is important before making transformations you do not want to apply to the original, or when you want to keep a pre-cleaning backup while building a cleaned version.

import pandas as pd

original = pd.DataFrame({'a': [1, 2, 3]})
ref = original          # same object!
copy = original.copy()  # independent copy

ref['a'] = 99
print(original['a'].tolist())  # [99 99 99] -- ref modified original
print(copy['a'].tolist())      # [1, 2, 3]  -- copy unchanged

Quick Check

Test your understanding of creating DataFrames from this lesson.

Lesson Recap

In this lesson you learned: DataFrames can be created from dicts of lists, lists of dicts, or NumPy arrays, the columns, dtypes, and index attributes describe the table structure, and df.copy() is required to get an independent copy rather than a reference. Next up we select specific columns and rows using bracket notation, .loc, and .iloc.

Häufig gestellte Fragen

Ist die Lektion „DataFrames erstellen“ kostenlos?

Ja — der vollständige Text von „DataFrames erstellen“ ist hier im Web kostenlos zu lesen. Um sie interaktiv zu üben (integrierter Code-Editor und 24/7 KI-Tutor) und den Rest des Pandas & NumPy Academy-Kurses freizuschalten, upgrade auf CoddyKit PRO. Der Pandas & NumPy Academy-Kurs umfasst insgesamt 4 Lektionen.

Was lerne ich in „DataFrames erstellen“?

Erstellen Sie DataFrames aus Dictionaries von Listen, Listen von Dictionaries und NumPy-Arrays und untersuchen Sie anschließend shape, columns und dtypes. Du übst Pandas & NumPy Academy mit praktischem Code, den du direkt im Browser ausführst, und ein 24/7 KI-Tutor beantwortet deine Fragen während du die Lektion bearbeitest.

Brauche ich Erfahrung, um Pandas & NumPy Academy zu starten?

Keine Vorkenntnisse erforderlich. Pandas & NumPy Academy auf CoddyKit ist für Anfänger bis fortgeschrittene Lernende strukturiert, sodass du hier starten oder von Anfang an beginnen und in deinem eigenen Tempo voranschreiten kannst. Dies ist Lektion 1 von 4.

Wie lange dauert die Lektion „DataFrames erstellen“?

Die meisten CoddyKit-Lektionen dauern etwa 5–10 Minuten. Jede ist kompakt und interaktiv, sodass du stetig Fortschritte machst und genau dort weitermachst, wo du aufgehört hast – im Web und in der App.

Kann ich in dieser Pandas & NumPy Academy-Lektion Code schreiben und ausführen?

Ja. Jede Pandas & NumPy Academy-Lektion enthält einen integrierten Code-Editor, sodass du echten Code direkt in deinem Browser schreibst und ausführst und sofort KI-Feedback erhältst — ohne lokale Einrichtung erforderlich.

Alle Lektionen in diesem Kurs

  1. DataFrames erstellen
  2. Spalten und Zeilen auswählen
  3. Spalten hinzufügen und entfernen
  4. Grundlegende DataFrame-Inspektion
← Zurück zu Pandas & NumPy Academy