0Pricing
Pandas & NumPy Academy · 课时

选择列与行

使用方括号表示法选择单列或多列,并通过 .loc 和 .iloc 按标签和位置获取行。

选择列与行 是 CoddyKit 上的免费 Pandas & NumPy Academy 课时。 这是第 2 节课,共 4 节。 你可以在下方免费阅读本课时的完整内容 — 然后在浏览器中使用内置代码编辑器和全天候 AI 导师进行实践。 这是 Pandas & NumPy Academy 学习路径的一部分,你的进度在网页和 CoddyKit 应用中同步。 Pandas & NumPy Academy 课程共包含 4 节课。

本课时的部分内容尚未翻译,以英文显示。

Selecting a Single Column

Access a single column by name using bracket notation df['col'], which returns a Series. You can also use dot notation df.col when the column name is a valid Python identifier with no spaces. Bracket notation is always safe; dot notation fails for names that clash with DataFrame methods like count or mean.

import pandas as pd

df = pd.DataFrame({'name': ['Alice', 'Bob'], 'score': [90, 75]})
print(df['name'])     # Series
print(type(df['name']))  # pandas.core.series.Series

Selecting Multiple Columns

To select multiple columns, pass a list of column names inside the brackets: df[['col1', 'col2']]. Note the double brackets — the outer pair is the indexing operator, the inner pair creates a Python list. The result is a DataFrame, not a Series. This is commonly used to extract a feature matrix for machine learning.

import pandas as pd

df = pd.DataFrame({'a': [1,2], 'b': [3,4], 'c': [5,6]})
subset = df[['a', 'c']]
print(type(subset))      # pandas.core.frame.DataFrame
print(subset)

.loc for Row and Column Selection

df.loc[row_label, col_label] selects data by label. Provide a single label, a list of labels, or a slice for both rows and columns. df.loc[:, 'score'] selects all rows of the 'score' column. df.loc['r1':'r3', ['a', 'b']] selects rows r1 to r3 (inclusive) for columns a and b.

import pandas as pd

df = pd.DataFrame({'x': [10,20,30], 'y': [40,50,60]},
                  index=['r1', 'r2', 'r3'])
print(df.loc['r2', 'x'])       # 20
print(df.loc['r1':'r2', 'y'])  # r1:40  r2:50

.iloc for Positional Selection

df.iloc[row_pos, col_pos] selects by integer position ignoring labels. Both arguments follow Python slice conventions with exclusive stops. df.iloc[:, 0] selects the first column as a Series. df.iloc[0:2, 1:3] selects a rectangular 2×2 sub-DataFrame from the top-left.

import pandas as pd

df = pd.DataFrame({'a': [1,2,3], 'b': [4,5,6], 'c': [7,8,9]})
print(df.iloc[1, 2])        # 8  -- row 1, col 2
print(df.iloc[0:2, 0:2])   # top-left 2x2 sub-table

Boolean Row Selection

Pass a boolean Series (with the same index as the DataFrame) to .loc to filter rows. Build the boolean Series from a column condition. You can combine conditions with & and |. This is the standard filtering pattern in Pandas data analysis and is far more expressive than SQL WHERE clauses for complex multi-column conditions.

import pandas as pd

df = pd.DataFrame({'name': ['A','B','C','D'], 'score': [80,55,92,71]})
high = df.loc[df['score'] >= 70]
print(high)
#   name  score
# 0    A     80
# 2    C     92
# 3    D     71

Compound Filtering Conditions

Multiple conditions must each be in parentheses and combined with & (AND) or | (OR). Using Python's and/or keywords on Series raises an error. To negate a condition, use ~ (tilde) instead of not. Always test conditions individually before combining to isolate the source of unexpected results.

import pandas as pd

df = pd.DataFrame({'name': ['A','B','C','D'],
                   'score': [80, 55, 92, 71],
                   'dept': ['Eng', 'HR', 'Eng', 'HR']})
filtered = df.loc[(df['score'] >= 70) & (df['dept'] == 'Eng')]
print(filtered)

Selecting Rows by Integer Position

df.iloc[n] returns the n-th row as a Series. df.iloc[n:m] returns rows n through m-1 as a DataFrame. df.iloc[[0, 2, 4]] selects specific rows by position using fancy indexing. These are equivalent to NumPy array row selection and are often used in train/test splitting before applying machine learning models.

import pandas as pd

df = pd.DataFrame({'a': range(5), 'b': range(5, 10)})
print(df.iloc[0])         # first row as Series
print(df.iloc[1:3])       # rows 1 and 2 as DataFrame
print(df.iloc[[0, 4]])    # first and last row

Combining Row and Column Selection

Both .loc and .iloc accept two arguments: df.loc[row_selector, col_selector]. This enables precise rectangular selection in a single expression. A colon : alone selects all rows or all columns. This pattern is essential for extracting a feature matrix with specific columns for only the training rows of a dataset.

import pandas as pd

df = pd.DataFrame({'x': [1,2,3], 'y': [4,5,6], 'z': [7,8,9]})
# Filter rows where x > 1, keep columns y and z
result = df.loc[df['x'] > 1, ['y', 'z']]
print(result)
#    y  z
# 1  5  8
# 2  6  9

Selecting a Single Row with .loc

When you pass a single scalar label to df.loc[label], the result is a Series where the index is the column names. This is useful for inspecting a specific record. To get a single-row DataFrame instead, wrap the label in a list: df.loc[[label]]. The distinction matters when passing results to functions expecting a DataFrame.

import pandas as pd

df = pd.DataFrame({'name': ['A','B'], 'score': [80, 90]},
                  index=['r1', 'r2'])
print(type(df.loc['r1']))    # Series
print(type(df.loc[['r1']])) # DataFrame

at and iat for Fast Scalar Access

df.at[row_label, col_name] and df.iat[row_pos, col_pos] retrieve or set a single scalar value with lower overhead than .loc/.iloc. They are designed for loops that update individual cells. Always prefer vectorized operations over cell-by-cell updates in production code, but use at/iat when you genuinely need to iterate.

import pandas as pd

df = pd.DataFrame({'x': [1, 2, 3], 'y': [4, 5, 6]},
                  index=['a', 'b', 'c'])
print(df.at['b', 'y'])    # 5
print(df.iat[2, 0])       # 3

Chained Indexing Warning

Chained indexing like df['col'][condition] or df[mask]['col'] = val can produce a SettingWithCopyWarning because Pandas may operate on a copy instead of the original. Always use a single .loc expression for any modification: df.loc[mask, 'col'] = val. This is the correct, warning-free pattern.

import pandas as pd

df = pd.DataFrame({'score': [80, 55, 92]})
# WRONG - may not modify original:
# df[df['score'] > 60]['score'] = 100

# CORRECT:
df.loc[df['score'] > 60, 'score'] = 100
print(df)

Quick Check

Test your understanding of selecting columns and rows from this lesson.

Lesson Recap

In this lesson you learned: .loc selects rows and columns by label with inclusive slice stops, .iloc selects by integer position with exclusive stops like Python slices, and boolean conditions applied through .loc filter rows without the SettingWithCopyWarning pitfall of chained indexing. Next up we add new computed columns, rename existing ones, and remove unwanted columns with drop().

常见问题解答

「选择列与行」课时是免费的吗?

是的 — 「选择列与行」的完整文本可在网页上免费阅读。要进行交互式练习(内置代码编辑器和全天候 AI 导师)并解锁 Pandas & NumPy Academy 课程的其余内容,请升级到 CoddyKit PRO。 Pandas & NumPy Academy 课程共包含 4 节课。

「选择列与行」这节课中我会学到什么?

使用方括号表示法选择单列或多列,并通过 .loc 和 .iloc 按标签和位置获取行。 你通过在浏览器中直接运行的动手代码来练习 Pandas & NumPy Academy,全天候 AI 导师会在你学习这节课的过程中回答你的问题。

学习 Pandas & NumPy Academy 需要有经验吗?

无需任何先前经验。CoddyKit 上的 Pandas & NumPy Academy 课程适合初学者到高级学习者,你可以从这里开始或从头开始,按照自己的节奏学习。 这是第 2 节课,共 4 节。

「选择列与行」课时需要多长时间?

大多数 CoddyKit 课程大约需要 5–10 分钟。每节课都很精短且互动,所以你能稳步进步,并在网页和应用中从离开的地方继续。

我能在这节 Pandas & NumPy Academy 课中编写并运行代码吗?

能。每节 Pandas & NumPy Academy 课都包含内置代码编辑器,你可以在浏览器中直接编写并运行真实代码,并获得即时 AI 反馈 — 无需本地设置。

此课程中的所有课时

  1. 创建 DataFrames
  2. 选择列与行
  3. 添加与删除列
  4. DataFrame 基础检查
← 返回 Pandas & NumPy Academy