0Pricing
Data Science Academy · 강의

자료형과 중복 행 수정하기

자료형을 변환하고 반복 행을 삭제합니다

자료형과 중복 행 수정하기은(는) CoddyKit의 무료 Data Science Academy 강의입니다. 이것은 4개 중 4번째 강의입니다. 아래에서 전체 강의를 무료로 읽을 수 있으며, 내장 코드 에디터와 24/7 AI 튜터와 함께 브라우저에서 직접 실습할 수 있습니다. 이 강의는 Data Science Academy 학습 경로의 일부이며, 진행 상황이 웹과 CoddyKit 앱에 동기화됩니다. Data Science Academy 강의에는 총 4개의 강의가 포함되어 있습니다.

이 강의의 일부는 아직 번역되지 않았으며 영어로 표시됩니다.

Clean Goes Beyond Gaps

Filling missing values is only half the job. Clean data also needs correct dtypes and no accidental duplicate rows muddying your counts. 🧹

What a dtype Is

Every column has a dtype, its data type, like int, float, object, or datetime. The dtype controls which operations the column even allows.

df.dtypes

Numbers Stuck as Text

A frequent mess: numbers loaded as object strings because one stray symbol crept in. You cannot do math on them until the type is fixed.

Coerce With to_numeric

Convert a text column to numbers with to_numeric. Pass errors='coerce' to turn anything unparseable into NaN instead of crashing.

df['price'] = pd.to_numeric(df['price'], errors='coerce')

Cast With astype

When a column is already clean, astype changes its type directly, for example turning whole-number floats back into compact integers.

df['count'] = df['count'].astype(int)

Fix Date Columns

Dates often arrive as text. Convert them with to_datetime so you can sort, filter by range, and extract parts like month later.

df['date'] = pd.to_datetime(df['date'])

Category to Save Memory

Columns with few repeated text values shrink dramatically as the category dtype, which stores each label once and references it.

df['city'] = df['city'].astype('category')

Spotting Duplicate Rows

Repeated records inflate totals and averages. Find them with duplicated(), which flags each row that has appeared before as True.

df.duplicated().sum()

Dropping Duplicates

Remove repeats with drop_duplicates(). By default it keeps the first occurrence of each row and discards the rest.

df = df.drop_duplicates()

Duplicates by Key Columns

Sometimes only certain columns define a duplicate. Pass subset to detect repeats based on a key like an id, ignoring the other columns.

df.drop_duplicates(subset=['user_id'])

Verify Before You Move On

After fixing types and repeats, run info() once more. Confirming dtypes and row counts protects you from cleaning errors slipping downstream.

df.info()

Quick Check

A price column loaded as object text. Which call safely turns bad entries into NaN?

Recap: Types and Dupes Tamed

You now fix dtypes with to_numeric, astype, and to_datetime, and clear repeats with drop_duplicates. Your table is finally analysis-ready.

자주 묻는 질문

“자료형과 중복 행 수정하기” 강의는 무료인가요?

네 — “자료형과 중복 행 수정하기” 전체 내용을 이 웹사이트에서 무료로 읽을 수 있습니다. 인터랙티브하게 실습하려면(내장 코드 에디터와 24/7 AI 튜터), CoddyKit PRO로 업그레이드하면 Data Science Academy 강의 전체를 잠금 해제할 수 있습니다. Data Science Academy 강의에는 총 4개의 강의가 포함되어 있습니다.

“자료형과 중복 행 수정하기”에서 뭘 배우나요?

자료형을 변환하고 반복 행을 삭제합니다 브라우저에서 직접 실행하는 실습 코드로 Data Science Academy을(를) 배우며, 24/7 AI 튜터가 강의를 진행하면서 질문에 답변해줍니다.

Data Science Academy을(를) 시작하는 데 경험이 필요한가요?

사전 경험은 필요하지 않습니다. CoddyKit의 Data Science Academy은(는) 초급자부터 고급 학습자까지를 위해 구성되어 있으므로, 여기서 시작하거나 처음부터 시작할 수 있으며 자신의 속도대로 진행할 수 있습니다. 이것은 4개 중 4번째 강의입니다.

“자료형과 중복 행 수정하기” 강의는 얼마나 걸리나요?

대부분의 CoddyKit 강의는 약 5~10분이 소요됩니다. 각 강의는 간결하고 인터랙티브하여 꾸준한 진행이 가능하며, 웹과 앱에서 중단한 부분부터 바로 시작할 수 있습니다.

이 Data Science Academy 강의에서 코드를 작성하고 실행할 수 있나요?

네. 모든 Data Science Academy 강의에는 내장 코드 에디터가 포함되어 있으므로, 브라우저에서 바로 실제 코드를 작성하고 실행한 후 즉시 AI 피드백을 받을 수 있습니다 — 로컬 설정이 필요 없습니다.

이 강의의 모든 강의

  1. 테이블에 숨은 NaNs 찾기
  2. 삭제와 채우기 중 현명하게 선택하기
  3. 평균, 중앙값, 최빈값으로 대체하기
  4. 자료형과 중복 행 수정하기
← Data Science Academy(으)로 돌아가기