0Pricing
Pandas & NumPy Academy · 강의

열과 행 선택

대괄호 표기법으로 하나 이상의 열을 선택하고 .loc 및 .iloc으로 레이블과 위치에 따라 행을 가져옵니다.

열과 행 선택은(는) CoddyKit의 무료 Pandas & NumPy Academy 강의입니다. 이것은 4개 중 2번째 강의입니다. 아래에서 전체 강의를 무료로 읽을 수 있으며, 내장 코드 에디터와 24/7 AI 튜터와 함께 브라우저에서 직접 실습할 수 있습니다. 이 강의는 Pandas & NumPy Academy 학습 경로의 일부이며, 진행 상황이 웹과 CoddyKit 앱에 동기화됩니다. Pandas & NumPy Academy 강의에는 총 4개의 강의가 포함되어 있습니다.

이 강의의 일부는 아직 번역되지 않았으며 영어로 표시됩니다.

Selecting a Single Column

Access a single column by name using bracket notation df['col'], which returns a Series. You can also use dot notation df.col when the column name is a valid Python identifier with no spaces. Bracket notation is always safe; dot notation fails for names that clash with DataFrame methods like count or mean.

import pandas as pd

df = pd.DataFrame({'name': ['Alice', 'Bob'], 'score': [90, 75]})
print(df['name'])     # Series
print(type(df['name']))  # pandas.core.series.Series

Selecting Multiple Columns

To select multiple columns, pass a list of column names inside the brackets: df[['col1', 'col2']]. Note the double brackets — the outer pair is the indexing operator, the inner pair creates a Python list. The result is a DataFrame, not a Series. This is commonly used to extract a feature matrix for machine learning.

import pandas as pd

df = pd.DataFrame({'a': [1,2], 'b': [3,4], 'c': [5,6]})
subset = df[['a', 'c']]
print(type(subset))      # pandas.core.frame.DataFrame
print(subset)

.loc for Row and Column Selection

df.loc[row_label, col_label] selects data by label. Provide a single label, a list of labels, or a slice for both rows and columns. df.loc[:, 'score'] selects all rows of the 'score' column. df.loc['r1':'r3', ['a', 'b']] selects rows r1 to r3 (inclusive) for columns a and b.

import pandas as pd

df = pd.DataFrame({'x': [10,20,30], 'y': [40,50,60]},
                  index=['r1', 'r2', 'r3'])
print(df.loc['r2', 'x'])       # 20
print(df.loc['r1':'r2', 'y'])  # r1:40  r2:50

.iloc for Positional Selection

df.iloc[row_pos, col_pos] selects by integer position ignoring labels. Both arguments follow Python slice conventions with exclusive stops. df.iloc[:, 0] selects the first column as a Series. df.iloc[0:2, 1:3] selects a rectangular 2×2 sub-DataFrame from the top-left.

import pandas as pd

df = pd.DataFrame({'a': [1,2,3], 'b': [4,5,6], 'c': [7,8,9]})
print(df.iloc[1, 2])        # 8  -- row 1, col 2
print(df.iloc[0:2, 0:2])   # top-left 2x2 sub-table

Boolean Row Selection

Pass a boolean Series (with the same index as the DataFrame) to .loc to filter rows. Build the boolean Series from a column condition. You can combine conditions with & and |. This is the standard filtering pattern in Pandas data analysis and is far more expressive than SQL WHERE clauses for complex multi-column conditions.

import pandas as pd

df = pd.DataFrame({'name': ['A','B','C','D'], 'score': [80,55,92,71]})
high = df.loc[df['score'] >= 70]
print(high)
#   name  score
# 0    A     80
# 2    C     92
# 3    D     71

Compound Filtering Conditions

Multiple conditions must each be in parentheses and combined with & (AND) or | (OR). Using Python's and/or keywords on Series raises an error. To negate a condition, use ~ (tilde) instead of not. Always test conditions individually before combining to isolate the source of unexpected results.

import pandas as pd

df = pd.DataFrame({'name': ['A','B','C','D'],
                   'score': [80, 55, 92, 71],
                   'dept': ['Eng', 'HR', 'Eng', 'HR']})
filtered = df.loc[(df['score'] >= 70) & (df['dept'] == 'Eng')]
print(filtered)

Selecting Rows by Integer Position

df.iloc[n] returns the n-th row as a Series. df.iloc[n:m] returns rows n through m-1 as a DataFrame. df.iloc[[0, 2, 4]] selects specific rows by position using fancy indexing. These are equivalent to NumPy array row selection and are often used in train/test splitting before applying machine learning models.

import pandas as pd

df = pd.DataFrame({'a': range(5), 'b': range(5, 10)})
print(df.iloc[0])         # first row as Series
print(df.iloc[1:3])       # rows 1 and 2 as DataFrame
print(df.iloc[[0, 4]])    # first and last row

Combining Row and Column Selection

Both .loc and .iloc accept two arguments: df.loc[row_selector, col_selector]. This enables precise rectangular selection in a single expression. A colon : alone selects all rows or all columns. This pattern is essential for extracting a feature matrix with specific columns for only the training rows of a dataset.

import pandas as pd

df = pd.DataFrame({'x': [1,2,3], 'y': [4,5,6], 'z': [7,8,9]})
# Filter rows where x > 1, keep columns y and z
result = df.loc[df['x'] > 1, ['y', 'z']]
print(result)
#    y  z
# 1  5  8
# 2  6  9

Selecting a Single Row with .loc

When you pass a single scalar label to df.loc[label], the result is a Series where the index is the column names. This is useful for inspecting a specific record. To get a single-row DataFrame instead, wrap the label in a list: df.loc[[label]]. The distinction matters when passing results to functions expecting a DataFrame.

import pandas as pd

df = pd.DataFrame({'name': ['A','B'], 'score': [80, 90]},
                  index=['r1', 'r2'])
print(type(df.loc['r1']))    # Series
print(type(df.loc[['r1']])) # DataFrame

at and iat for Fast Scalar Access

df.at[row_label, col_name] and df.iat[row_pos, col_pos] retrieve or set a single scalar value with lower overhead than .loc/.iloc. They are designed for loops that update individual cells. Always prefer vectorized operations over cell-by-cell updates in production code, but use at/iat when you genuinely need to iterate.

import pandas as pd

df = pd.DataFrame({'x': [1, 2, 3], 'y': [4, 5, 6]},
                  index=['a', 'b', 'c'])
print(df.at['b', 'y'])    # 5
print(df.iat[2, 0])       # 3

Chained Indexing Warning

Chained indexing like df['col'][condition] or df[mask]['col'] = val can produce a SettingWithCopyWarning because Pandas may operate on a copy instead of the original. Always use a single .loc expression for any modification: df.loc[mask, 'col'] = val. This is the correct, warning-free pattern.

import pandas as pd

df = pd.DataFrame({'score': [80, 55, 92]})
# WRONG - may not modify original:
# df[df['score'] > 60]['score'] = 100

# CORRECT:
df.loc[df['score'] > 60, 'score'] = 100
print(df)

Quick Check

Test your understanding of selecting columns and rows from this lesson.

Lesson Recap

In this lesson you learned: .loc selects rows and columns by label with inclusive slice stops, .iloc selects by integer position with exclusive stops like Python slices, and boolean conditions applied through .loc filter rows without the SettingWithCopyWarning pitfall of chained indexing. Next up we add new computed columns, rename existing ones, and remove unwanted columns with drop().

자주 묻는 질문

“열과 행 선택” 강의는 무료인가요?

네 — “열과 행 선택” 전체 내용을 이 웹사이트에서 무료로 읽을 수 있습니다. 인터랙티브하게 실습하려면(내장 코드 에디터와 24/7 AI 튜터), CoddyKit PRO로 업그레이드하면 Pandas & NumPy Academy 강의 전체를 잠금 해제할 수 있습니다. Pandas & NumPy Academy 강의에는 총 4개의 강의가 포함되어 있습니다.

“열과 행 선택”에서 뭘 배우나요?

대괄호 표기법으로 하나 이상의 열을 선택하고 .loc 및 .iloc으로 레이블과 위치에 따라 행을 가져옵니다. 브라우저에서 직접 실행하는 실습 코드로 Pandas & NumPy Academy을(를) 배우며, 24/7 AI 튜터가 강의를 진행하면서 질문에 답변해줍니다.

Pandas & NumPy Academy을(를) 시작하는 데 경험이 필요한가요?

사전 경험은 필요하지 않습니다. CoddyKit의 Pandas & NumPy Academy은(는) 초급자부터 고급 학습자까지를 위해 구성되어 있으므로, 여기서 시작하거나 처음부터 시작할 수 있으며 자신의 속도대로 진행할 수 있습니다. 이것은 4개 중 2번째 강의입니다.

“열과 행 선택” 강의는 얼마나 걸리나요?

대부분의 CoddyKit 강의는 약 5~10분이 소요됩니다. 각 강의는 간결하고 인터랙티브하여 꾸준한 진행이 가능하며, 웹과 앱에서 중단한 부분부터 바로 시작할 수 있습니다.

이 Pandas & NumPy Academy 강의에서 코드를 작성하고 실행할 수 있나요?

네. 모든 Pandas & NumPy Academy 강의에는 내장 코드 에디터가 포함되어 있으므로, 브라우저에서 바로 실제 코드를 작성하고 실행한 후 즉시 AI 피드백을 받을 수 있습니다 — 로컬 설정이 필요 없습니다.

이 강의의 모든 강의

  1. DataFrames 만들기
  2. 열과 행 선택
  3. 열 추가와 삭제
  4. 기본 DataFrame 검사
← Pandas & NumPy Academy(으)로 돌아가기