असंतुलन का पता लगाना: वर्ग वितरण और आधाररेखा की समस्याएँ
शिक्षार्थी वर्ग आवृत्तियों की गणना करेंगे, असंतुलित डेटासेट में डमी-वर्गीकारक के जाल को उजागर करेंगे और पुष्टि करेंगे कि यहाँ सटीकता भ्रामक माप है।
असंतुलन का पता लगाना: वर्ग वितरण और आधाररेखा की समस्याएँ, CoddyKit पर Machine Learning Academy का एक निःशुल्क पाठ है। यह 4 में से 1वाँ पाठ है। आप नीचे पूरा पाठ निःशुल्क पढ़ सकते हैं—फिर अंतर्निहित कोड संपादक और 24/7 एआई ट्यूटर के साथ ब्राउज़र में इसका व्यावहारिक अभ्यास कर सकते हैं। यह Machine Learning Academy सीखने के मार्ग का हिस्सा है और आपकी प्रगति वेब तथा CoddyKit ऐप पर सिंक होती रहती है। Machine Learning Academy पाठ्यक्रम में कुल 4 पाठ शामिल हैं।
Class Imbalance क्या है
Class imbalance तब होता है जब dataset में कोई एक class हावी हो जाती है। Fraud detection में fraudulent transactions सभी records का केवल 0.1% हो सकती हैं; medical diagnosis में कोई दुर्लभ बीमारी 1000 में से केवल 1 patient को प्रभावित कर सकती है। ऐसा model जो हमेशा majority class की prediction करता है, 99.9% accuracy प्राप्त कर सकता है, जबकि दुर्लभ और महत्वपूर्ण class को पूरी तरह अनदेखा करता है—इसलिए इन परिस्थितियों में accuracy एक खतरनाक रूप से भ्रामक metric बन जाती है।
Class Distribution को मापना
किसी भी model को train करने से पहले, pd.Series(y).value_counts() या np.bincount(y) से class distribution की जाँच करें। एक उपयोगी summary statistic imbalance ratio है—majority और minority class की संख्या का अनुपात। 10:1 से अधिक ratio को imbalanced माना जाता है; 100:1 से अधिक ratio severe imbalance दर्शाता है और इसके लिए विशेष techniques की आवश्यकता होती है।
import numpy as np
import pandas as pd
# Simulate imbalanced binary classification dataset
np.random.seed(0)
y = np.array([0] * 950 + [1] * 50) # 95% negative, 5% positive
counts = pd.Series(y).value_counts()
print('Class counts:\n', counts)
print()
print('Class proportions:\n', counts / len(y))
print()
print('Imbalance ratio:', counts[0] / counts[1])Imbalanced Data में Accuracy का जाल
95% negative examples वाले dataset में ऐसा model जो हमेशा majority class (negative) की prediction करता है, बिना कुछ उपयोगी सीखे 95% accuracy प्राप्त कर लेता है। यही accuracy trap है। Model में positive cases का पता लगाने की कोई क्षमता नहीं होती (यही वे class हैं जिनकी आपको वास्तव में परवाह है), फिर भी इसकी accuracy प्रभावशाली दिखाई देती है। इसी कारण imbalanced datasets में accuracy के अलावा अन्य metrics की आवश्यकता होती है।
from sklearn.dummy import DummyClassifier
from sklearn.metrics import accuracy_score, classification_report
import numpy as np
from sklearn.model_selection import train_test_split
from sklearn.datasets import make_classification
X, y = make_classification(n_samples=1000, weights=[0.95, 0.05],
random_state=42, n_features=10)
X_train, X_test, y_train, y_test = train_test_split(X, y, random_state=42)
# Always predict majority class
dummy = DummyClassifier(strategy='most_frequent')
dummy.fit(X_train, y_train)
y_pred = dummy.predict(X_test)
print('Accuracy:', accuracy_score(y_test, y_pred).round(4))
print()
print(classification_report(y_test, y_pred))Dummy Classifier Baseline क्यों महत्वपूर्ण है
strategy='most_frequent' वाला DummyClassifier न्यूनतम स्वीकार्य baseline है। हर वास्तविक model को इसे सार्थक रूप से बेहतर करना चाहिए—सिर्फ accuracy में नहीं, बल्कि महत्वपूर्ण metrics (precision, recall, F1 या AUC-ROC) में भी। यदि आपका वास्तविक model dummy से केवल थोड़ा ही बेहतर है, तो संभवतः उसने भी minority class को अनदेखा करना सीख लिया है।
from sklearn.linear_model import LogisticRegression
from sklearn.dummy import DummyClassifier
from sklearn.metrics import classification_report
from sklearn.preprocessing import StandardScaler
from sklearn.datasets import make_classification
from sklearn.model_selection import train_test_split
X, y = make_classification(n_samples=1000, weights=[0.95, 0.05], random_state=42)
X_train, X_test, y_train, y_test = train_test_split(X, y, random_state=42)
X_train_s = StandardScaler().fit_transform(X_train)
X_test_s = StandardScaler().fit_transform(X_test)
lr = LogisticRegression().fit(X_train_s, y_train)
print('--- Logistic Regression ---')
print(classification_report(y_test, lr.predict(X_test_s)))Imbalanced Data पर Precision और Recall
Imbalanced problems के लिए दो सबसे उपयोगी metrics हैं: Precision = TP / (TP + FP)—जिन सभी positives की prediction की गई, उनमें से वास्तव में कितने positive थे? Recall = TP / (TP + FN)—सभी वास्तविक positives में से model ने कितनों का पता लगाया? Fraud detection में recall सर्वोपरि है (fraud छूटनी नहीं चाहिए); spam filtering में precision महत्वपूर्ण है (वैध emails को गलत तरीके से classify नहीं करना चाहिए)।
from sklearn.metrics import precision_score, recall_score, f1_score
from sklearn.linear_model import LogisticRegression
from sklearn.preprocessing import StandardScaler
from sklearn.datasets import make_classification
from sklearn.model_selection import train_test_split
X, y = make_classification(n_samples=1000, weights=[0.95, 0.05], random_state=42)
X_train, X_test, y_train, y_test = train_test_split(X, y, random_state=42)
X_train_s = StandardScaler().fit_transform(X_train)
X_test_s = StandardScaler().fit_transform(X_test)
lr = LogisticRegression().fit(X_train_s, y_train)
y_pred = lr.predict(X_test_s)
print(f'Precision: {precision_score(y_test, y_pred):.4f}')
print(f'Recall: {recall_score(y_test, y_pred):.4f}')
print(f'F1-Score: {f1_score(y_test, y_pred):.4f}')ROC-AUC: Threshold से स्वतंत्र Evaluation
ROC-AUC सभी संभावित decision thresholds पर model का evaluation करता है और परिणामस्वरूप बनने वाली curve के नीचे का area मापता है। 0.5 का AUC random performance दर्शाता है; 1.0 perfect performance है; 0.9+ excellent माना जाता है। Accuracy के विपरीत, AUC class imbalance से प्रभावित नहीं होता, क्योंकि यह fixed threshold वाली prediction के बजाय model की ranking ability का evaluation करता है। अधिकांश imbalanced binary classification tasks के लिए AUC recommended primary metric है।
from sklearn.metrics import roc_auc_score
from sklearn.linear_model import LogisticRegression
from sklearn.preprocessing import StandardScaler
from sklearn.datasets import make_classification
from sklearn.model_selection import train_test_split
X, y = make_classification(n_samples=1000, weights=[0.95, 0.05], random_state=42)
X_train, X_test, y_train, y_test = train_test_split(X, y, random_state=42)
sc = StandardScaler()
X_train_s = sc.fit_transform(X_train)
X_test_s = sc.transform(X_test)
lr = LogisticRegression().fit(X_train_s, y_train)
proba = lr.predict_proba(X_test_s)[:, 1]
print('ROC-AUC:', roc_auc_score(y_test, proba).round(4))Class Imbalance को Visualise करना
Class frequencies का एक सरल bar chart stakeholders को imbalance तुरंत समझा देता है। सामान्य practice के रूप में इसे अपने exploratory data analysis notebook में शामिल करें। यदि आप Pandas DataFrame के साथ काम कर रहे हैं, तो data collection bias से उत्पन्न छिपे हुए imbalance की भी जाँच करें, जो वास्तविक दुनिया में मौजूद rarity के कारण नहीं होता।
import matplotlib.pyplot as plt
import numpy as np
y = np.array([0] * 950 + [1] * 50)
classes, counts = np.unique(y, return_counts=True)
plt.bar(['Negative (0)', 'Positive (1)'], counts, color=['#2196F3', '#F44336'])
plt.ylabel('Count')
plt.title('Class Distribution (95:5 imbalance)')
for i, c in enumerate(counts):
plt.text(i, c + 5, str(c), ha='center', fontweight='bold')
plt.show()Ratios बनाए रखने के लिए Stratified Splitting
Imbalanced dataset को split करते समय train_test_split में stratify=y का उपयोग करें। Stratification के बिना, संयोगवश minority class के सभी samples एक ही split में जा सकते हैं, जिससे training या evaluation निरर्थक हो जाएगा। Stratified splitting यह सुनिश्चित करता है कि train और test sets में original dataset के समान class ratio हो।
from sklearn.model_selection import train_test_split
import numpy as np
y = np.array([0] * 950 + [1] * 50)
X = np.random.randn(1000, 5)
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.2, stratify=y, random_state=42
)
print('Train class ratio:', (y_train == 1).mean().round(4))
print('Test class ratio:', (y_test == 1).mean().round(4))
print('Expected ratio:', (y == 1).mean().round(4))Cross-Validation के लिए Stratified K-Fold
इसी तरह, StratifiedKFold का उपयोग करें (या cross_val_score में केवल cv=5 दें—classifiers के लिए scikit-learn अपने-आप StratifiedKFold का उपयोग करता है), ताकि प्रत्येक fold में original class ratio बना रहे। छोटी minority classes वाले regular KFold में ऐसे folds बन सकते हैं जिनमें कोई positive example न हो, जिससे errors या भ्रामक metrics उत्पन्न होते हैं।
from sklearn.model_selection import StratifiedKFold, cross_val_score
from sklearn.linear_model import LogisticRegression
from sklearn.preprocessing import StandardScaler
from sklearn.pipeline import Pipeline
import numpy as np
y = np.array([0] * 950 + [1] * 50)
X = np.random.randn(1000, 5)
pipe = Pipeline([('sc', StandardScaler()), ('lr', LogisticRegression())])
skf = StratifiedKFold(n_splits=5)
scores = cross_val_score(pipe, X, y, cv=skf, scoring='roc_auc')
print(f'Stratified CV AUC: {np.mean(scores):.4f} +/- {np.std(scores):.4f}')Precision-Recall Curve
अत्यधिक imbalanced data के लिए precision-recall (PR) curve अक्सर ROC curve से अधिक informative होती है। PR curve सभी thresholds पर y-axis पर precision और x-axis पर recall दिखाती है। PR curve के नीचे का area (AP score) 0 से 1 तक होता है और minority-class performance के प्रति संवेदनशील होता है; majority class के अत्यधिक प्रभावी होने पर ROC-AUC इस performance को छिपा सकता है।
from sklearn.metrics import precision_recall_curve, average_precision_score
from sklearn.linear_model import LogisticRegression
from sklearn.preprocessing import StandardScaler
from sklearn.datasets import make_classification
from sklearn.model_selection import train_test_split
import matplotlib.pyplot as plt
X, y = make_classification(n_samples=1000, weights=[0.95, 0.05], random_state=42)
X_train, X_test, y_train, y_test = train_test_split(X, y, stratify=y, random_state=42)
sc = StandardScaler()
X_train_s = sc.fit_transform(X_train)
X_test_s = sc.transform(X_test)
lr = LogisticRegression().fit(X_train_s, y_train)
proba = lr.predict_proba(X_test_s)[:, 1]
precision, recall, _ = precision_recall_curve(y_test, proba)
ap = average_precision_score(y_test, proba)
plt.plot(recall, precision, label=f'AP={ap:.3f}')
plt.xlabel('Recall')
plt.ylabel('Precision')
plt.title('Precision-Recall Curve')
plt.legend()
plt.show()Modelling से पहले Imbalance का Documentation
Best practice: हर classification project की शुरुआत में imbalance report बनाएँ। Class counts, ratios और baseline dummy accuracy दर्ज करें। इससे अपेक्षाएँ स्पष्ट होती हैं और कोई model train करने से पहले team को सही success metric पर सहमत होना पड़ता है। 95% accuracy वाला fraud detection model आमतौर पर बेकार होता है—आपके stakeholders को यह बात शुरुआत में ही पता होनी चाहिए।
import numpy as np
from sklearn.dummy import DummyClassifier
from sklearn.metrics import accuracy_score, roc_auc_score
y = np.array([0] * 950 + [1] * 50)
X = np.random.randn(1000, 5)
dummy = DummyClassifier(strategy='most_frequent').fit(X, y)
dummy_acc = accuracy_score(y, dummy.predict(X))
dummy_auc = roc_auc_score(y, dummy.predict_proba(X)[:, 1])
print('=== Imbalance Report ===')
print(f'Class 0: {(y==0).sum()} ({(y==0).mean():.1%})')
print(f'Class 1: {(y==1).sum()} ({(y==1).mean():.1%})')
print(f'Imbalance ratio: {(y==0).sum()/(y==1).sum():.0f}:1')
print(f'Dummy accuracy: {dummy_acc:.4f}')
print(f'Dummy AUC: {dummy_auc:.4f}')
print('Recommended metric: ROC-AUC or Precision-Recall AUC')त्वरित जाँच
इस पाठ से class imbalance और baseline की कमियों के बारे में अपनी समझ जाँचें।
पाठ का पुनरावलोकन
इस पाठ में आपने सीखा: class imbalance accuracy को भ्रामक metric बना देता है—केवल majority class की prediction करने वाला model बहुत अधिक accuracy प्राप्त कर सकता है, हमेशा DummyClassifier baseline से तुलना करें ताकि यह सुनिश्चित हो कि आपका model trivial तरीका अपनाने से आगे कुछ सीख रहा है, और imbalanced datasets के लिए ROC-AUC या precision-recall AUC को primary metrics के रूप में उपयोग करें। अब हम minority class को upsample करने के लिए SMOTE और random oversampling लागू करेंगे।
एआई शिक्षक के साथ Python सीखें — निःशुल्क
अपने ब्राउज़र में वास्तविक कोड लिखें और चलाएँ, चौबीसों घंटे एआई शिक्षक से तुरंत सहायता पाएँ, और वेब या ऐप पर वहीं से शुरू करें जहाँ आपने छोड़ा था।
- पाठ्यक्रम
- 30
- पाठ
- 120
अक्सर पूछे जाने वाले प्रश्न
क्या “असंतुलन का पता लगाना: वर्ग वितरण और आधाररेखा की समस्याएँ” पाठ निःशुल्क है?
हाँ—“असंतुलन का पता लगाना: वर्ग वितरण और आधाररेखा की समस्याएँ” का पूरा पाठ यहाँ वेब पर निःशुल्क पढ़ा जा सकता है। इंटरैक्टिव अभ्यास (अंतर्निहित कोड संपादक और 24/7 एआई ट्यूटर) करने और Machine Learning Academy पाठ्यक्रम का बाकी हिस्सा अनलॉक करने के लिए CoddyKit PRO लें। Machine Learning Academy पाठ्यक्रम में कुल 4 पाठ शामिल हैं।
“असंतुलन का पता लगाना: वर्ग वितरण और आधाररेखा की समस्याएँ” में मैं क्या सीखूँगा?
शिक्षार्थी वर्ग आवृत्तियों की गणना करेंगे, असंतुलित डेटासेट में डमी-वर्गीकारक के जाल को उजागर करेंगे और पुष्टि करेंगे कि यहाँ सटीकता भ्रामक माप है। आप ब्राउज़र में सीधे चलाए जाने वाले व्यावहारिक कोड के साथ Machine Learning Academy का अभ्यास करते हैं, और पाठ पूरा करते समय 24/7 एआई ट्यूटर आपके प्रश्नों के उत्तर देता है।
क्या Machine Learning Academy शुरू करने के लिए मुझे किसी अनुभव की आवश्यकता है?
पहले के अनुभव की आवश्यकता नहीं है। CoddyKit पर Machine Learning Academy शुरुआती से लेकर उन्नत शिक्षार्थियों तक सभी के लिए व्यवस्थित किया गया है, इसलिए आप यहीं से या शुरुआत से सीखना शुरू कर सकते हैं और अपनी गति से आगे बढ़ सकते हैं। यह 4 में से 1वाँ पाठ है।
“असंतुलन का पता लगाना: वर्ग वितरण और आधाररेखा की समस्याएँ” पाठ पूरा करने में कितना समय लगता है?
CoddyKit का अधिकांश पाठ लगभग 5–10 मिनट में पूरा हो जाता है। हर पाठ छोटा और संवादात्मक है, इसलिए आप लगातार प्रगति करते हैं और वेब या ऐप पर वहीं से सीखना जारी रख सकते हैं जहाँ आपने छोड़ा था।
क्या मैं इस Machine Learning Academy पाठ में कोड लिख और चला सकता हूँ?
हाँ। हर Machine Learning Academy पाठ में एक अंतर्निर्मित कोड संपादक शामिल है, जिससे आप सीधे अपने ब्राउज़र में वास्तविक कोड लिख और चला सकते हैं और तुरंत एआई प्रतिक्रिया पा सकते हैं—स्थानीय सेटअप की आवश्यकता नहीं है।
इस पाठ्यक्रम के सभी पाठ
- असंतुलन का पता लगाना: वर्ग वितरण और आधाररेखा की समस्याएँ
- रैंडम ओवरसैंपलिंग और SMOTE
- रैंडम अंडरसैंपलिंग और ClusterCentroids
- वर्ग भार और सीमा स्थानांतरण