0Pricing
Pandas & NumPy Academy · 강의

일관되지 않은 범주 표준화

별칭을 매핑하고 퍼지 매칭으로 오탈자를 수정하며 표준 목록을 적용해 자유 입력 텍스트 범주 열을 정규화합니다.

일관되지 않은 범주 표준화은(는) CoddyKit의 무료 Pandas & NumPy Academy 강의입니다. 이것은 4개 중 3번째 강의입니다. 아래에서 전체 강의를 무료로 읽을 수 있으며, 내장 코드 에디터와 24/7 AI 튜터와 함께 브라우저에서 직접 실습할 수 있습니다. 이 강의는 Pandas & NumPy Academy 학습 경로의 일부이며, 진행 상황이 웹과 CoddyKit 앱에 동기화됩니다. Pandas & NumPy Academy 강의에는 총 4개의 강의가 포함되어 있습니다.

이 강의의 일부는 아직 번역되지 않았으며 영어로 표시됩니다.

The Problem of Inconsistent Categories

When category data is entered by humans, the same concept appears under many different spellings: Electronics, electronics, ELECTRONICS, Electronicss. A groupby on this column produces dozens of tiny groups instead of one meaningful group. Standardising categories into a canonical list is one of the most important cleaning steps before any aggregation or machine learning feature creation.

import pandas as pd

df = pd.read_csv('products.csv')
print(df['category'].value_counts().head(20))

Normalising Case and Whitespace

The first and easiest standardisation step is normalising case and whitespace. Apply .str.strip().str.lower() to remove leading/trailing spaces and convert everything to lowercase before any other comparison. This single step collapses many variants: Electronics, electronics, and ' Electronics ' all become electronics after normalisation.

df['category_clean'] = df['category'].str.strip().str.lower()

print('Before:', df['category'].nunique())
print('After:', df['category_clean'].nunique())

Mapping Aliases with a Dict

After case normalisation, many variants are still distinct due to abbreviations, synonyms, or legacy names. Build an alias mapping dictionary where keys are variant spellings and values are the canonical form. Apply it with df['category_clean'].map(alias_map).fillna(df['category_clean']) — the fillna preserves values not in the dictionary rather than replacing them with NaN.

alias_map = {
    'elect': 'electronics',
    'elec': 'electronics',
    'tech': 'electronics',
    'clothing': 'apparel',
    'clothes': 'apparel',
    'garments': 'apparel'
}

df['category_clean'] = df['category_clean'].map(alias_map).fillna(df['category_clean'])
print(df['category_clean'].value_counts().head())

Using str.replace for Pattern Fixes

Some category names have consistent formatting errors like double spaces or trailing numbers. Use .str.replace() with a regular expression to fix these patterns across all rows at once. For example, .str.replace(r'\s+', ' ', regex=True) collapses multiple spaces into one, and .str.replace(r'\d+$', '', regex=True) strips trailing digits from category names.

df['category_clean'] = (
    df['category_clean']
    .str.replace(r'\s+', ' ', regex=True)  # collapse spaces
    .str.replace(r'\d+$', '', regex=True)   # strip trailing numbers
    .str.strip()
)

print(df['category_clean'].value_counts())

Fuzzy Matching with difflib

When typos are unpredictable, use fuzzy matching to find the closest canonical category for each variant. Python's built-in difflib.get_close_matches() returns the best matches from a list of valid categories based on string similarity. Apply it as a helper function via df['category'].apply() to resolve variants that map() alone cannot catch.

from difflib import get_close_matches

CANONICAL = ['electronics', 'apparel', 'home', 'sports', 'beauty']

def fuzzy_fix(val):
    matches = get_close_matches(str(val).lower().strip(), CANONICAL, n=1, cutoff=0.7)
    return matches[0] if matches else val

df['category_fuzzy'] = df['category_clean'].apply(fuzzy_fix)
print(df[['category_clean', 'category_fuzzy']].head(10))

Building a Canonical Category List

Define your canonical categories explicitly rather than inferring them from the data. A hard-coded list forces every value through a validation gate; anything not in the list is flagged as unknown. Maintain the canonical list in a config file or a separate DataFrame column so it is easy to update when the business adds a new product category without touching the cleaning code.

CANONICAL_CATEGORIES = {
    'electronics', 'apparel', 'home', 'sports',
    'beauty', 'food', 'toys', 'automotive'
}

df['is_known_category'] = df['category_fuzzy'].isin(CANONICAL_CATEGORIES)
unknown = df[~df['is_known_category']]['category_fuzzy'].unique()
print('Unknown categories remaining:', unknown)

Handling Unknown Categories

Unknown categories that do not match any canonical value after fuzzy fixing should be grouped under an Other label rather than dropped, so you do not silently lose rows. Set them to 'other' with np.where() or a simple conditional assignment. Log how many rows were mapped to 'other' and review them for patterns that might warrant a new canonical category.

import numpy as np

df['category_final'] = np.where(
    df['category_fuzzy'].isin(CANONICAL_CATEGORIES),
    df['category_fuzzy'],
    'other'
)

print(df['category_final'].value_counts())

Enforcing a Categorical Dtype

After standardisation, convert the clean category column to a Pandas Categorical dtype with an explicit list of valid categories. This enforces the schema: any attempt to assign an invalid category raises an error. It also reduces memory by storing category strings as integer codes internally, and speeds up groupby operations on large DataFrames.

df['category_final'] = pd.Categorical(
    df['category_final'],
    categories=list(CANONICAL_CATEGORIES) + ['other']
)

print(df['category_final'].dtype)
print(df['category_final'].cat.categories)

Validating the Clean Column

After all standardisation steps, run a final validation: assert that no value outside the canonical set exists in the cleaned column. Use assert df['category_final'].isin(valid_set).all(). Place this assertion at the end of the cleaning function so it runs every time the pipeline executes, catching regressions when new raw data contains previously unseen category labels.

valid_set = set(CANONICAL_CATEGORIES) | {'other'}

assert df['category_final'].isin(valid_set).all(), 'Invalid category found'
print('All categories valid. Distribution:')
print(df['category_final'].value_counts())

Tracking the Standardisation Changes

Build a diff table showing the original value and its cleaned replacement for every row that changed. This audit trail lets data owners review and approve or reject specific mappings. Use a boolean mask to select changed rows and compare the original and final columns side by side. Export the diff as a CSV for non-technical stakeholders to review.

changed = df['category'] != df['category_final']
diff_table = df[changed][['category', 'category_final']].drop_duplicates()
diff_table.columns = ['original', 'mapped_to']
print(f'Unique mappings applied: {len(diff_table)}')
print(diff_table)

Saving the Standardised Dataset

Replace the original messy category column with the standardised one and drop intermediate working columns before saving. Store the alias mapping dictionary and the fuzzy-fix cutoff threshold in the pipeline config so the standardisation is fully reproducible. A cleaning run two months later on a new data dump should produce identical results for the same input values.

df_out = df.drop(columns=['category_clean', 'category_fuzzy', 'is_known_category'])
df_out = df_out.rename(columns={'category_final': 'category'})

df_out.to_parquet('products_clean.parquet', index=False)
print('Saved. Category distribution:')
print(df_out['category'].value_counts())

Quick Check

Test your understanding of Data Analysis concepts from this lesson.

Lesson Recap

In this lesson you learned: normalising case and whitespace as a first standardisation step, mapping aliases and applying fuzzy matching for typo correction, and enforcing a canonical category list with Categorical dtype and assertions. Next up we explore schema validation and runtime assertions to guard every pipeline stage.

자주 묻는 질문

“일관되지 않은 범주 표준화” 강의는 무료인가요?

네 — “일관되지 않은 범주 표준화” 전체 내용을 이 웹사이트에서 무료로 읽을 수 있습니다. 인터랙티브하게 실습하려면(내장 코드 에디터와 24/7 AI 튜터), CoddyKit PRO로 업그레이드하면 Pandas & NumPy Academy 강의 전체를 잠금 해제할 수 있습니다. Pandas & NumPy Academy 강의에는 총 4개의 강의가 포함되어 있습니다.

“일관되지 않은 범주 표준화”에서 뭘 배우나요?

별칭을 매핑하고 퍼지 매칭으로 오탈자를 수정하며 표준 목록을 적용해 자유 입력 텍스트 범주 열을 정규화합니다. 브라우저에서 직접 실행하는 실습 코드로 Pandas & NumPy Academy을(를) 배우며, 24/7 AI 튜터가 강의를 진행하면서 질문에 답변해줍니다.

Pandas & NumPy Academy을(를) 시작하는 데 경험이 필요한가요?

사전 경험은 필요하지 않습니다. CoddyKit의 Pandas & NumPy Academy은(는) 초급자부터 고급 학습자까지를 위해 구성되어 있으므로, 여기서 시작하거나 처음부터 시작할 수 있으며 자신의 속도대로 진행할 수 있습니다. 이것은 4개 중 3번째 강의입니다.

“일관되지 않은 범주 표준화” 강의는 얼마나 걸리나요?

대부분의 CoddyKit 강의는 약 5~10분이 소요됩니다. 각 강의는 간결하고 인터랙티브하여 꾸준한 진행이 가능하며, 웹과 앱에서 중단한 부분부터 바로 시작할 수 있습니다.

이 Pandas & NumPy Academy 강의에서 코드를 작성하고 실행할 수 있나요?

네. 모든 Pandas & NumPy Academy 강의에는 내장 코드 에디터가 포함되어 있으므로, 브라우저에서 바로 실제 코드를 작성하고 실행한 후 즉시 AI 피드백을 받을 수 있습니다 — 로컬 설정이 필요 없습니다.

이 강의의 모든 강의

  1. 중복 감지와 제거
  2. 이상값 감지와 처리
  3. 일관되지 않은 범주 표준화
  4. 스키마 검증과 단언
← Pandas & NumPy Academy(으)로 돌아가기