创建 DataFrames
从列表字典、字典列表和 NumPy 数组构建 DataFrames,然后查看其形状、列和数据类型。
创建 DataFrames 是 CoddyKit 上的免费 Pandas & NumPy Academy 课时。 这是第 1 节课,共 4 节。 你可以在下方免费阅读本课时的完整内容 — 然后在浏览器中使用内置代码编辑器和全天候 AI 导师进行实践。 这是 Pandas & NumPy Academy 学习路径的一部分,你的进度在网页和 CoddyKit 应用中同步。 Pandas & NumPy Academy 课程共包含 4 节课。
本课时的部分内容尚未翻译,以英文显示。
What Is a DataFrame?
A DataFrame is a two-dimensional, size-mutable, labelled data structure — essentially a table with rows and columns, similar to a spreadsheet or SQL table. Each column is a Pandas Series sharing the same row index. DataFrames can hold columns of different dtypes, making them perfect for real-world datasets that mix numbers, text, and dates.
import pandas as pd
df = pd.DataFrame({'name': ['Alice', 'Bob', 'Carol'],
'age': [30, 25, 35],
'score': [88.5, 72.0, 95.3]})
print(df)Creating from a Dict of Lists
The most common way to construct a DataFrame is from a dict of lists: the keys become column names and the lists become column values. All lists must be the same length. The row index defaults to 0, 1, 2, ... unless you specify one with index=. This pattern mirrors how CSV data is often represented in Python.
import pandas as pd
df = pd.DataFrame({
'city': ['Berlin', 'Paris', 'Rome'],
'pop': [3.6, 2.2, 2.8],
'country': ['DE', 'FR', 'IT']
})
print(df.shape) # (3, 3)
print(df.dtypes)Creating from a List of Dicts
You can also pass a list of dicts, where each dict represents one row. Missing keys in any dict produce NaN for that column. This format is common when consuming JSON APIs that return a list of record objects. Pandas aligns all keys across dicts to form the column set automatically.
import pandas as pd
rows = [
{'name': 'Alice', 'age': 30, 'dept': 'Eng'},
{'name': 'Bob', 'age': 25}, # 'dept' missing -> NaN
{'name': 'Carol', 'age': 35, 'dept': 'HR'}
]
df = pd.DataFrame(rows)
print(df)Creating from a NumPy Array
Passing a 2-D NumPy array creates a DataFrame with integer column names (0, 1, 2, ...) and a RangeIndex by default. Supply columns= and index= to add meaningful labels. This is the bridge between NumPy numerical computation and Pandas labelled data analysis.
import pandas as pd
import numpy as np
arr = np.array([[1, 2, 3], [4, 5, 6], [7, 8, 9]])
df = pd.DataFrame(arr,
columns=['x', 'y', 'z'],
index=['r1', 'r2', 'r3'])
print(df)Setting a Custom Index at Creation
The index= parameter sets the row labels at creation time. Common choices include date strings, IDs, or category names that make row lookups meaningful. A well-chosen index makes .loc access intuitive and enables powerful time-series operations when the index is a DatetimeIndex.
import pandas as pd
df = pd.DataFrame(
{'revenue': [100, 200, 150], 'costs': [80, 160, 90]},
index=['Jan', 'Feb', 'Mar']
)
print(df)
print(df.loc['Feb']) # select row 'Feb'Inspecting columns, dtypes, and index
Three attributes immediately tell you the structure of a DataFrame: df.columns gives an Index of column names, df.dtypes returns a Series mapping each column to its dtype, and df.index describes the row labels. These are the first checks to run on any new DataFrame to understand what data you have before transforming it.
import pandas as pd
df = pd.DataFrame({'a': [1, 2], 'b': [3.0, 4.0], 'c': ['x', 'y']})
print(df.columns) # Index(['a', 'b', 'c'], dtype='object')
print(df.dtypes)
# a int64
# b float64
# c object
print(df.index) # RangeIndex(start=0, stop=2, step=1)Specifying Column Order
When constructing from a dict, column order follows dict insertion order (Python 3.7+). If you need a specific order, pass the columns= parameter with a list of column names. Any name not in the data will produce a NaN column; any name in the data but not in the list will be excluded. This lets you select and order columns at construction time.
import pandas as pd
data = {'c': [3, 6], 'a': [1, 4], 'b': [2, 5]}
df = pd.DataFrame(data, columns=['a', 'b', 'c'])
print(df.columns.tolist()) # ['a', 'b', 'c']Creating from a Dict of Series
Passing a dict of Pandas Series aligns on the union of all indices. Where a Series is missing a label that another has, the result cell is NaN. This is the most index-aware construction method and is used when combining separately computed columns that may have different row counts or labels.
import pandas as pd
s1 = pd.Series([1, 2, 3], index=['a', 'b', 'c'])
s2 = pd.Series([10, 20], index=['b', 'c'])
df = pd.DataFrame({'col1': s1, 'col2': s2})
print(df)
# col1 col2
# a 1.0 NaN
# b 2.0 10.0
# c 3.0 20.0shape, len, and size
df.shape returns a tuple (nrows, ncols). len(df) returns the number of rows. df.size returns the total number of cells (rows × columns). These are the quickest sanity checks after loading data to confirm you received the expected number of records and that no columns were silently dropped.
import pandas as pd
df = pd.DataFrame({'a': range(5), 'b': range(5), 'c': range(5)})
print(df.shape) # (5, 3)
print(len(df)) # 5
print(df.size) # 15 (5 rows x 3 cols)Creating an Empty DataFrame
You can create an empty DataFrame with specific column names and dtypes for use as a template or accumulator. Append rows with pd.concat([df, new_row])m. Specifying dtypes at creation avoids expensive type inference when rows are added later. An empty DataFrame is also useful as a seed in iterative data collection loops.
import pandas as pd
df = pd.DataFrame(columns=['name', 'score'])
new_row = pd.DataFrame([{'name': 'Alice', 'score': 90}])
df = pd.concat([df, new_row], ignore_index=True)
print(df)Copying a DataFrame
Assigning a DataFrame to a new variable creates a reference, not a copy — modifying one modifies the other. Use df.copy() to create an independent copy. This is important before making transformations you do not want to apply to the original, or when you want to keep a pre-cleaning backup while building a cleaned version.
import pandas as pd
original = pd.DataFrame({'a': [1, 2, 3]})
ref = original # same object!
copy = original.copy() # independent copy
ref['a'] = 99
print(original['a'].tolist()) # [99 99 99] -- ref modified original
print(copy['a'].tolist()) # [1, 2, 3] -- copy unchangedQuick Check
Test your understanding of creating DataFrames from this lesson.
Lesson Recap
In this lesson you learned: DataFrames can be created from dicts of lists, lists of dicts, or NumPy arrays, the columns, dtypes, and index attributes describe the table structure, and df.copy() is required to get an independent copy rather than a reference. Next up we select specific columns and rows using bracket notation, .loc, and .iloc.
常见问题解答
「创建 DataFrames」课时是免费的吗?
是的 — 「创建 DataFrames」的完整文本可在网页上免费阅读。要进行交互式练习(内置代码编辑器和全天候 AI 导师)并解锁 Pandas & NumPy Academy 课程的其余内容,请升级到 CoddyKit PRO。 Pandas & NumPy Academy 课程共包含 4 节课。
「创建 DataFrames」这节课中我会学到什么?
从列表字典、字典列表和 NumPy 数组构建 DataFrames,然后查看其形状、列和数据类型。 你通过在浏览器中直接运行的动手代码来练习 Pandas & NumPy Academy,全天候 AI 导师会在你学习这节课的过程中回答你的问题。
学习 Pandas & NumPy Academy 需要有经验吗?
无需任何先前经验。CoddyKit 上的 Pandas & NumPy Academy 课程适合初学者到高级学习者,你可以从这里开始或从头开始,按照自己的节奏学习。 这是第 1 节课,共 4 节。
「创建 DataFrames」课时需要多长时间?
大多数 CoddyKit 课程大约需要 5–10 分钟。每节课都很精短且互动,所以你能稳步进步,并在网页和应用中从离开的地方继续。
我能在这节 Pandas & NumPy Academy 课中编写并运行代码吗?
能。每节 Pandas & NumPy Academy 课都包含内置代码编辑器,你可以在浏览器中直接编写并运行真实代码,并获得即时 AI 反馈 — 无需本地设置。
此课程中的所有课时
- 创建 DataFrames
- 选择列与行
- 添加与删除列
- DataFrame 基础检查