Pandas & NumPy Academy · पाठ

स्वतंत्रता के लिए Chi-Squared परीक्षण

क्रॉसटैब आवृत्ति तालिका पर chi2_contingency का उपयोग करके जाँचिए कि दो श्रेणीबद्ध चर स्वतंत्र हैं या नहीं।

पाठ 3, कुल 4 में से13 चरण

स्वतंत्रता के लिए Chi-Squared परीक्षण, CoddyKit पर Pandas & NumPy Academy का एक निःशुल्क पाठ है। यह 4 में से 3वाँ पाठ है। आप नीचे पूरा पाठ निःशुल्क पढ़ सकते हैं—फिर अंतर्निहित कोड संपादक और 24/7 एआई ट्यूटर के साथ ब्राउज़र में इसका व्यावहारिक अभ्यास कर सकते हैं। यह Pandas & NumPy Academy सीखने के मार्ग का हिस्सा है और आपकी प्रगति वेब तथा CoddyKit ऐप पर सिंक होती रहती है। Pandas & NumPy Academy पाठ्यक्रम में कुल 4 पाठ शामिल हैं।

Categorical Relationships की जाँच

independence के लिए chi-squared test यह जाँचता है कि दो categorical variables statistically independent हैं या उनके बीच association है। उदाहरण के लिए: 'क्या customer churn subscription tier से independent है?' या 'क्या product preference age group से independent है?' Numeric means की तुलना करने वाले t-tests के विपरीत, chi-squared tests contingency table में मौजूद observed frequencies की तुलना उन frequencies से करते हैं जिनकी हम अपेक्षा करते, यदि variables independent होते।

Pandas से Contingency Table बनाना

contingency table (जिसे cross-tabulation भी कहते हैं) दो categorical variables के प्रत्येक combination के लिए observations की count दिखाती है। pd.crosstab(df['var1'], df['var2']) सीधे DataFrame से यह table बनाता है। प्रत्येक cell में उन observations की count होती है जिनमें row category और column category साथ-साथ मौजूद हैं। यही scipy.stats.chi2_contingency() का input है।

import pandas as pd
import numpy as np

np.random.seed(0)
df = pd.DataFrame({
    'subscription': np.random.choice(['free', 'basic', 'pro'], 300),
    'churned': np.random.choice(['yes', 'no'], 300, p=[0.3, 0.7])
})

# Build contingency table
ct = pd.crosstab(df['subscription'], df['churned'])
print(ct)
print()
print('Row totals:', ct.sum(axis=1).to_dict())

chi2_contingency() चलाना

scipy.stats.chi2_contingency(observed) observed contingency table (NumPy array या Pandas DataFrame के रूप में) लेता है और चार values लौटाता है: chi-squared statistic, p-value, degrees of freedom और expected frequency table। Null hypothesis है H₀: the two variables are independent। यदि p ≤ 0.05, तो हम independence को reject करते हैं और निष्कर्ष निकालते हैं कि statistically significant association है।

import pandas as pd
import numpy as np
from scipy import stats

np.random.seed(0)
df = pd.DataFrame({
    'tier': np.random.choice(['free', 'basic', 'pro'], 300),
    'churned': np.random.choice(['yes', 'no'], 300, p=[0.3, 0.7])
})

ct = pd.crosstab(df['tier'], df['churned'])
chi2, p, dof, expected = stats.chi2_contingency(ct)

print(f'Chi-squared: {chi2:.3f}')
print(f'p-value: {p:.4f}')
print(f'Degrees of freedom: {dof}')
print('Association significant?', 'Yes' if p < 0.05 else 'No')

Expected Frequencies को समझना

expected frequencies दर्शाती हैं कि यदि दोनों variables पूरी तरह independent होते, तो cell counts कैसी दिखतीं। प्रत्येक cell के लिए, expected = (row_total × column_total) / grand_total। Chi-squared statistic observed और expected counts के बीच squared differences के योग को expected counts से normalise करके मापता है। बड़ा chi-squared बताता है कि observed counts independence से बहुत अधिक विचलित हैं; छोटा chi-squared बताता है कि डेटा independence के अनुरूप है।

import pandas as pd
import numpy as np
from scipy import stats

observed = pd.DataFrame({
    'converted': [50, 30],
    'not_converted': [150, 170]
}, index=['variant', 'control'])

chi2, p, dof, expected = stats.chi2_contingency(observed)
print('Observed:')
print(observed)
print()
print('Expected (under independence):')
print(pd.DataFrame(expected,
                   index=observed.index,
                   columns=observed.columns).round(1))

Chi-Squared में Degrees of Freedom

r rows और c columns वाली contingency table के लिए degrees of freedom df = (r-1) × (c-1) होता है। 2×2 table में df=1 और 3×4 table में df=6 होता है। Degrees of freedom बढ़ने पर chi-squared की critical value भी बढ़ती है: α=0.05 पर df=1 के लिए critical value ≈ 3.84; df=6 के लिए α=0.05 पर ≈ 12.6। chi2_contingency function यह अपने-आप संभालता है, लेकिन formula जानने से समझ आता है कि बड़ी tables से प्राप्त chi-squared को significant होने के लिए बड़े statistics की आवश्यकता क्यों होती है।

from scipy import stats

# Critical values of chi-squared for alpha=0.05
for df in [1, 2, 3, 4, 6, 9]:
    critical = stats.chi2.ppf(0.95, df=df)
    print(f'df={df}: critical value = {critical:.3f}')

Minimum Expected Frequency की मान्यता

Chi-squared test तभी विश्वसनीय होता है जब सभी expected frequencies कम-से-कम 5 हों (कुछ sources के अनुसार कम-से-कम 1 हो सकती है, बशर्ते 20% से अधिक values 5 से कम न हों)। छोटी expected counts chi-squared approximation को inaccurate बना देती हैं। इस मान्यता का उल्लंघन होने पर 2×2 tables के लिए Fisher's exact test (scipy.stats.fisher_exact()) का उपयोग करें, या expected counts बढ़ाने के लिए दुर्लभ categories को मिला दें। chi2_contingency द्वारा लौटाई गई expected frequency matrix की हमेशा जाँच करें।

import numpy as np
from scipy import stats

# Table with small expected counts
observed = np.array([[2, 3], [100, 95]])
chi2, p, dof, expected = stats.chi2_contingency(observed)

print('Expected frequencies:')
print(expected)
print('Min expected:', expected.min())

if expected.min() < 5:
    print('Warning: Expected frequency < 5. Use Fisher\'s exact test.')
    odds_ratio, p_fisher = stats.fisher_exact(observed)
    print(f'Fisher\'s exact p: {p_fisher:.4f}')

छोटे Samples के लिए Fisher's Exact Test

scipy.stats.fisher_exact(table) null hypothesis of independence के अंतर्गत दी गई 2×2 table (या उससे भी अधिक extreme table) के observe होने की exact probability निकालता है, बिना किसी approximation के। यह sample size या cell counts की परवाह किए बिना हमेशा valid होता है — p-value approximate नहीं, exact होती है। इसकी computational cost है: यह सभी possible tables को enumerate करता है, इसलिए बड़े totals के लिए slow होता है। 2×2 tables में किसी भी cell count के 5 से कम होने पर हमेशा chi-squared के बजाय Fisher's exact test को प्राथमिकता दें।

import numpy as np
from scipy import stats

# Small clinical trial: treatment vs. outcome
observed = np.array([
    [3, 12],   # treated: 3 improved, 12 did not
    [1, 18]    # control: 1 improved, 18 did not
])

odds_ratio, p = stats.fisher_exact(observed, alternative='two-sided')
print(f'Odds ratio: {odds_ratio:.3f}')
print(f'p-value: {p:.4f}')
print('Association significant?', 'Yes' if p < 0.05 else 'No')

Association Strength मापना: Cramér's V

t-tests के लिए Cohen's d की तरह, केवल chi-squared statistic association की strength नहीं मापता — sample size के कारण यह बढ़ जाता है। Cramér's V chi-squared को 0–1 scale पर normalise करता है: 0 का अर्थ है कोई association नहीं, जबकि 1 का अर्थ है perfect association। V = sqrt(chi2 / (n × min(r-1, c-1)))। दिशानिर्देश: < 0.1 weak, 0.1–0.3 moderate और > 0.3 strong होता है। पूरी तस्वीर के लिए p-value के साथ हमेशा Cramér's V report करें।

import pandas as pd
import numpy as np
from scipy import stats

np.random.seed(0)
df = pd.DataFrame({
    'tier': np.random.choice(['free', 'basic', 'pro'], 500),
    'churn': np.random.choice(['yes', 'no'], 500, p=[0.3, 0.7])
})
ct = pd.crosstab(df['tier'], df['churn'])

chi2, p, dof, _ = stats.chi2_contingency(ct)
n = ct.sum().sum()
min_dim = min(ct.shape[0]-1, ct.shape[1]-1)
cramers_v = np.sqrt(chi2 / (n * min_dim))

print(f'chi2={chi2:.3f}, p={p:.4f}')
print(f'Cramer\'s V: {cramers_v:.4f}')
print('Association strength:', 'strong' if cramers_v > 0.3 else 'moderate' if cramers_v > 0.1 else 'weak')

Chi-Squared Goodness of Fit

Chi-squared का एक अलग application goodness-of-fit test है: यह जाँचता है कि observed frequencies किसी hypothesised distribution से match करती हैं या नहीं। उदाहरण के लिए, 'क्या पासे के rolls uniformly distributed हैं?' या 'क्या हमारी website traffic सप्ताह के दिनों के expected pattern का पालन करती है?' scipy.stats.chisquare(f_obs, f_exp) का उपयोग करें, जहाँ f_exp expected frequencies हैं। Null hypothesis यह है कि डेटा specified distribution का पालन करता है।

import numpy as np
from scipy import stats

# Observed: counts for each weekday over 70 weeks
observed = np.array([980, 1050, 1020, 1080, 1120, 850, 900])
# Expected: uniform distribution
expected = np.full(7, observed.sum() / 7)

chi2, p = stats.chisquare(f_obs=observed, f_exp=expected)
print('Observed:', observed)
print('Expected (uniform):', expected.round(0))
print(f'chi2={chi2:.3f}, p={p:.4f}')
print('Uniform?', 'Yes' if p > 0.05 else 'No - some days significantly busier')

normalise के साथ pd.crosstab का उपयोग करना

Chi-squared results report करते समय raw counts के बजाय proportions दिखाना उपयोगी होता है, ताकि readers relative association देख सकें। pd.crosstab(var1, var2, normalize='index') row proportion दिखाता है (प्रत्येक row category का कितना fraction प्रत्येक column में आता है)। normalize='columns' column proportions दिखाता है। Test चलाने से पहले इन proportions की visual comparison करने से association की direction का अनुमान लगाने और result को context में interpret करने में सहायता मिलती है।

import pandas as pd
import numpy as np

np.random.seed(0)
df = pd.DataFrame({
    'tier': np.random.choice(['free', 'basic', 'pro'], 300),
    'churned': np.random.choice(['yes', 'no'], 300, p=[0.3, 0.7])
})

# Row-proportions: churn rate within each tier
prop_table = pd.crosstab(df['tier'], df['churned'], normalize='index')
print('Churn rate by tier:')
print((prop_table * 100).round(1))

A/B Testing में Chi-Squared Test

Chi-squared test A/B testing conversion rates के लिए standard test है। rows = (control, variant) और columns = (converted, not converted) वाली 2×2 contingency table बनाएँ। Chi-squared test बताता है कि groups के बीच conversion rate में significant difference है या नहीं। बड़े samples के लिए यह two-proportion z-test के equivalent है। छोटे samples में (किसी भी expected count < 5 होने पर) इसके बजाय Fisher's exact test का उपयोग करें।

import numpy as np
from scipy import stats

# A/B test: control vs. variant, conversions vs. non-conversions
control_conv = 45
control_total = 500
variant_conv = 63
variant_total = 500

contingency = np.array([
    [control_conv, control_total - control_conv],
    [variant_conv, variant_total - variant_conv]
])

chi2, p, dof, expected = stats.chi2_contingency(contingency)
print(f'Control rate: {control_conv/control_total:.1%}')
print(f'Variant rate: {variant_conv/variant_total:.1%}')
print(f'p-value: {p:.4f}')
print('Variant significantly better?', 'Yes' if p < 0.05 else 'No')

त्वरित जाँच

इस lesson में Data Analysis की अवधारणाओं की अपनी समझ जाँचें।

Lesson का पुनरावलोकन

इस lesson में आपने सीखा: pd.crosstab() + chi2_contingency() observed और expected frequencies के आधार पर जाँचते हैं कि दो categorical variables independent हैं या नहीं, expected counts 5 से कम होने पर Fisher's exact test सुरक्षित alternative है, और Cramér's V sample size से independent association strength को मापता है। अगले lesson में हम ANOVA और post-hoc tests से तीन या अधिक groups के means की तुलना करेंगे।

शुरुआत निःशुल्क

एआई शिक्षक के साथ Python सीखें — निःशुल्क

अपने ब्राउज़र में वास्तविक कोड लिखें और चलाएँ, चौबीसों घंटे एआई शिक्षक से तुरंत सहायता पाएँ, और वेब या ऐप पर वहीं से शुरू करें जहाँ आपने छोड़ा था।

पाठ्यक्रम
30
पाठ
120

अक्सर पूछे जाने वाले प्रश्न

क्या “स्वतंत्रता के लिए Chi-Squared परीक्षण” पाठ निःशुल्क है?

हाँ—“स्वतंत्रता के लिए Chi-Squared परीक्षण” का पूरा पाठ यहाँ वेब पर निःशुल्क पढ़ा जा सकता है। इंटरैक्टिव अभ्यास (अंतर्निहित कोड संपादक और 24/7 एआई ट्यूटर) करने और Pandas & NumPy Academy पाठ्यक्रम का बाकी हिस्सा अनलॉक करने के लिए CoddyKit PRO लें। Pandas & NumPy Academy पाठ्यक्रम में कुल 4 पाठ शामिल हैं।

“स्वतंत्रता के लिए Chi-Squared परीक्षण” में मैं क्या सीखूँगा?

क्रॉसटैब आवृत्ति तालिका पर chi2_contingency का उपयोग करके जाँचिए कि दो श्रेणीबद्ध चर स्वतंत्र हैं या नहीं। आप ब्राउज़र में सीधे चलाए जाने वाले व्यावहारिक कोड के साथ Pandas & NumPy Academy का अभ्यास करते हैं, और पाठ पूरा करते समय 24/7 एआई ट्यूटर आपके प्रश्नों के उत्तर देता है।

क्या Pandas & NumPy Academy शुरू करने के लिए मुझे किसी अनुभव की आवश्यकता है?

पहले के अनुभव की आवश्यकता नहीं है। CoddyKit पर Pandas & NumPy Academy शुरुआती से लेकर उन्नत शिक्षार्थियों तक सभी के लिए व्यवस्थित किया गया है, इसलिए आप यहीं से या शुरुआत से सीखना शुरू कर सकते हैं और अपनी गति से आगे बढ़ सकते हैं। यह 4 में से 3वाँ पाठ है।

“स्वतंत्रता के लिए Chi-Squared परीक्षण” पाठ पूरा करने में कितना समय लगता है?

CoddyKit का अधिकांश पाठ लगभग 5–10 मिनट में पूरा हो जाता है। हर पाठ छोटा और संवादात्मक है, इसलिए आप लगातार प्रगति करते हैं और वेब या ऐप पर वहीं से सीखना जारी रख सकते हैं जहाँ आपने छोड़ा था।

क्या मैं इस Pandas & NumPy Academy पाठ में कोड लिख और चला सकता हूँ?

हाँ। हर Pandas & NumPy Academy पाठ में एक अंतर्निर्मित कोड संपादक शामिल है, जिससे आप सीधे अपने ब्राउज़र में वास्तविक कोड लिख और चला सकते हैं और तुरंत एआई प्रतिक्रिया पा सकते हैं—स्थानीय सेटअप की आवश्यकता नहीं है।

इस पाठ्यक्रम के सभी पाठ

  1. वर्णनात्मक आँकड़े और सामान्यता परीक्षण
  2. माध्य की तुलना के लिए T-परीक्षण
  3. स्वतंत्रता के लिए Chi-Squared परीक्षण
  4. ANOVA और पोस्ट-हॉक परीक्षण
← Pandas & NumPy Academy पर वापस जाएँ