इंडेक्स सेट और रीसेट करना
set_index() से किसी कॉलम को पंक्ति इंडेक्स बनाएँ, reset_index() से उसे वापस बदलें और MultiIndex का अवलोकन करें।
इंडेक्स सेट और रीसेट करना, CoddyKit पर Pandas & NumPy Academy का एक निःशुल्क पाठ है। यह 4 में से 4वाँ पाठ है। आप नीचे पूरा पाठ निःशुल्क पढ़ सकते हैं—फिर अंतर्निहित कोड संपादक और 24/7 एआई ट्यूटर के साथ ब्राउज़र में इसका व्यावहारिक अभ्यास कर सकते हैं। यह Pandas & NumPy Academy सीखने के मार्ग का हिस्सा है और आपकी प्रगति वेब तथा CoddyKit ऐप पर सिंक होती रहती है। Pandas & NumPy Academy पाठ्यक्रम में कुल 4 पाठ शामिल हैं।
DataFrame Index क्या है
हर Pandas DataFrame में एक row index होता है — प्रत्येक row से जुड़े labels का एक समूह। डिफ़ॉल्ट index RangeIndex (0, 1, 2, …) होता है, लेकिन आप इसे किसी भी ऐसे column से बदल सकते हैं जो प्राकृतिक identifier का काम करता हो: कोई date, product ID, user ID या कोई unique key। अर्थपूर्ण index readability सुधारता है, label-based slicing सक्षम करता है और time series resampling के लिए आवश्यक होता है।
import pandas as pd
df = pd.DataFrame({
'date': ['2024-01-01', '2024-01-02', '2024-01-03'],
'sales': [100, 200, 150]
})
print('Default index:', df.index.tolist()) # [0, 1, 2]
print(df)set_index() — किसी Column को Index बनाना
DataFrame.set_index('col') दिए गए column को row index बना देता है और उसे regular columns से हटा देता है। उस column की values row labels बन जाती हैं। जब कोई column natural key का काम करता हो, तब DataFrame को अधिक अर्थपूर्ण बनाने का यह standard तरीका है। नया DataFrame लौटाया जाता है; inplace=True का उपयोग न करने पर मूल DataFrame नहीं बदलता।
import pandas as pd
df = pd.DataFrame({
'date': ['2024-01-01', '2024-01-02', '2024-01-03'],
'sales': [100, 200, 150]
})
df = df.set_index('date')
print(df)
# sales
# date
# 2024-01-01 100
# 2024-01-02 200
# 2024-01-03 150
print(df.index) # Index(['2024-01-01', ...], dtype='object', name='date')Custom Index से Rows चुनना
किसी अर्थपूर्ण column को index बनाने के बाद, आप साफ़ और पढ़ने योग्य syntax में label के आधार पर rows प्राप्त करने के लिए .loc[label] का उपयोग कर सकते हैं — boolean masks की आवश्यकता नहीं होती। DatetimeIndex के लिए आप df.loc['2024-01'] जैसे आंशिक date strings का उपयोग करके जनवरी 2024 की सभी rows चुन सकते हैं।
import pandas as pd
df = pd.DataFrame({
'product': ['A', 'B', 'C'],
'price': [10, 20, 30],
'stock': [100, 50, 75]
})
df = df.set_index('product')
# Access row by label
print(df.loc['B'])
# price 20
# stock 50
# Name: B, dtype: int64कई Columns के साथ set_index() — MultiIndex
set_index() में column names की list देने पर MultiIndex (hierarchical index) बनता है। परिणामी DataFrame को .loc[] में tuples का उपयोग करके access किया जा सकता है। MultiIndex grouped time series (region + date), panel data (subject + time) और दो-स्तरीय row grouping वाले किसी भी analysis के लिए आवश्यक है।
import pandas as pd
df = pd.DataFrame({
'region': ['East', 'East', 'West', 'West'],
'year': [2023, 2024, 2023, 2024],
'revenue': [100, 150, 200, 220]
})
df = df.set_index(['region', 'year'])
print(df)
# revenue
# region year
# East 2023 100
# 2024 150
# West 2023 200
# 2024 220
print(df.loc[('East', 2024)]) # revenue = 150set_index() में drop= Parameter
डिफ़ॉल्ट रूप से, set_index('col') किसी column के index बनने के बाद उसे DataFrame के regular columns से हटा देता है। Index के रूप में उपयोग करने के साथ-साथ column को वहीं बनाए रखने के लिए drop=False दें। यह तब उपयोगी होता है जब आपको label-based access चाहिए, लेकिन regular column space में computations के लिए भी वह column उपलब्ध रखना हो।
import pandas as pd
df = pd.DataFrame({'id': [1, 2, 3], 'value': [10, 20, 30]})
# Keep 'id' as both index and column
df_with_both = df.set_index('id', drop=False)
print(df_with_both)
# id value
# id
# 1 1 10
# 2 2 20
# 3 3 30reset_index() — Index को वापस Column में लाना
reset_index(), set_index() का उलटा है: यह वर्तमान row index को वापस regular column में ले आता है और index को डिफ़ॉल्ट RangeIndex से बदल देता है। groupby().agg() के बाद या उन operations के बाद इसकी अक्सर आवश्यकता होती है जिनसे आपको कोई अर्थपूर्ण index मिल जाता है जिसे आगे की processing में column के रूप में उपयोग करना हो।
import pandas as pd
df = pd.DataFrame({'sales': [100, 200]}, index=['East', 'West'])
df.index.name = 'region'
reset = df.reset_index()
print(reset)
# region sales
# 0 East 100
# 1 West 200
print(reset.index.tolist()) # [0, 1] — default RangeIndex restoredreset_index(drop=True)
reset_index() में drop=True देने पर वर्तमान index को column में ले जाने के बजाय पूरी तरह हटा दिया जाता है। इसका उपयोग तब करें जब वर्तमान index में कोई अर्थपूर्ण जानकारी न हो (जैसे rows filter करने के बाद index में 0, 3, 7 जैसे gaps हों) और आप केवल 0 से शुरू होने वाला साफ़ क्रमिक index चाहते हों।
import pandas as pd
df = pd.DataFrame({'x': [10, 20, 30, 40, 50]})
filtered = df[df['x'] > 15]
print('Filtered index:', filtered.index.tolist()) # [1, 2, 3, 4]
# Drop the old index — don't move it to a column
clean = filtered.reset_index(drop=True)
print('Clean index:', clean.index.tolist()) # [0, 1, 2, 3]groupby के बाद reset_index()
groupby().agg() call करने के बाद groupby keys index बन जाती हैं। reset_index() call करने पर वे वापस regular columns में बदल जाती हैं, जिससे flat tabular result मिलता है जिसके साथ काम करना, merge करना या export करना आसान होता है। व्यवहार में reset_index() का यह सबसे आम उपयोग है।
import pandas as pd
df = pd.DataFrame({
'dept': ['Eng', 'HR', 'Eng', 'Sales', 'HR'],
'salary': [90000, 50000, 80000, 70000, 55000]
})
# After groupby, dept becomes the index
agg = df.groupby('dept')['salary'].mean()
print(type(agg), agg.index.tolist())
# dept is the index
# Reset makes dept a regular column again
result = agg.reset_index()
result.columns = ['dept', 'avg_salary']
print(result)Index के नाम के लिए rename_axis()
जब row index का कोई नाम न हो या आप स्पष्टता के लिए उसका नाम बदलना चाहें, तब df.rename_axis('new_name') (row index के लिए) या df.rename_axis('col_name', axis=1) (column axis के लिए) का उपयोग करें। Index को अर्थपूर्ण नाम देने से output को समझना आसान होता है, विशेष रूप से CSV में export करते समय, जहाँ index name column header बन जाता है।
import pandas as pd
df = pd.DataFrame({'sales': [100, 200, 300]}, index=['A', 'B', 'C'])
print(df.index.name) # None
# Name the index
df = df.rename_axis('region')
print(df)
# sales
# region
# A 100
# B 200
# C 300Gaps भरने के लिए Reindexing
DataFrame.reindex(new_index) DataFrame का आकार नए index के अनुरूप बदलता है और नई index में मौजूद, लेकिन मूल data में अनुपस्थित positions को NaN से भरता है। इसका उपयोग DataFrames को किसी साझा पूर्ण index के साथ align करने के लिए किया जाता है — जैसे यह सुनिश्चित करना कि वर्ष का हर महीना दिखाई दे, भले ही कुछ महीनों के लिए data उपलब्ध न हो।
import pandas as pd
df = pd.DataFrame({
'month': ['Jan', 'Mar', 'Jun'],
'revenue': [100, 200, 300]
}).set_index('month')
all_months = ['Jan', 'Feb', 'Mar', 'Apr', 'May', 'Jun']
complete = df.reindex(all_months)
print(complete)
# revenue
# Jan 100.0
# Feb NaN <- filled with NaN
# Mar 200.0
# Apr NaN
# May NaN
# Jun 300.0MultiIndex reset_index और Level Selection
MultiIndex DataFrame के लिए reset_index() डिफ़ॉल्ट रूप से सभी levels को वापस columns में ले आता है। केवल कुछ levels को reset करने के लिए level= दें। यह तब उपयोगी है जब आप किसी एक level को index बनाए रखना चाहते हों (जैसे date को index बनाए रखना) और केवल दूसरे level को column में बदलना चाहते हों।
import pandas as pd
df = pd.DataFrame(
{'revenue': [100, 150, 200, 220]},
index=pd.MultiIndex.from_tuples(
[('East', 2023), ('East', 2024), ('West', 2023), ('West', 2024)],
names=['region', 'year']
)
)
# Reset only the 'year' level, keep 'region' as index
df2 = df.reset_index(level='year')
print(df2)
# year revenue
# region
# East 2023 100
# East 2024 150
# West 2023 200
# West 2024 220त्वरित जाँच
Index को set और reset करने की अपनी समझ जाँचिए।
पाठ का पुनरावलोकन
इस पाठ में आपने सीखा: set_index('col') label-based access के लिए किसी column को row index बनाता है, list देने पर MultiIndex बनता है, और reset_index() index को वापस column में ले आता है (उसे हटाने के लिए drop=True का उपयोग करें)। Missing positions को NaN से भरते हुए किसी पूर्ण target index के साथ align करने के लिए reindex() का उपयोग करें। आगे आने वाले GroupBy और Merging lessons के लिए ये techniques बुनियादी हैं।
एआई शिक्षक के साथ Python सीखें — निःशुल्क
अपने ब्राउज़र में वास्तविक कोड लिखें और चलाएँ, चौबीसों घंटे एआई शिक्षक से तुरंत सहायता पाएँ, और वेब या ऐप पर वहीं से शुरू करें जहाँ आपने छोड़ा था।
- पाठ्यक्रम
- 30
- पाठ
- 120
अक्सर पूछे जाने वाले प्रश्न
क्या “इंडेक्स सेट और रीसेट करना” पाठ निःशुल्क है?
हाँ—“इंडेक्स सेट और रीसेट करना” का पूरा पाठ यहाँ वेब पर निःशुल्क पढ़ा जा सकता है। इंटरैक्टिव अभ्यास (अंतर्निहित कोड संपादक और 24/7 एआई ट्यूटर) करने और Pandas & NumPy Academy पाठ्यक्रम का बाकी हिस्सा अनलॉक करने के लिए CoddyKit PRO लें। Pandas & NumPy Academy पाठ्यक्रम में कुल 4 पाठ शामिल हैं।
“इंडेक्स सेट और रीसेट करना” में मैं क्या सीखूँगा?
set_index() से किसी कॉलम को पंक्ति इंडेक्स बनाएँ, reset_index() से उसे वापस बदलें और MultiIndex का अवलोकन करें। आप ब्राउज़र में सीधे चलाए जाने वाले व्यावहारिक कोड के साथ Pandas & NumPy Academy का अभ्यास करते हैं, और पाठ पूरा करते समय 24/7 एआई ट्यूटर आपके प्रश्नों के उत्तर देता है।
क्या Pandas & NumPy Academy शुरू करने के लिए मुझे किसी अनुभव की आवश्यकता है?
पहले के अनुभव की आवश्यकता नहीं है। CoddyKit पर Pandas & NumPy Academy शुरुआती से लेकर उन्नत शिक्षार्थियों तक सभी के लिए व्यवस्थित किया गया है, इसलिए आप यहीं से या शुरुआत से सीखना शुरू कर सकते हैं और अपनी गति से आगे बढ़ सकते हैं। यह 4 में से 4वाँ पाठ है।
“इंडेक्स सेट और रीसेट करना” पाठ पूरा करने में कितना समय लगता है?
CoddyKit का अधिकांश पाठ लगभग 5–10 मिनट में पूरा हो जाता है। हर पाठ छोटा और संवादात्मक है, इसलिए आप लगातार प्रगति करते हैं और वेब या ऐप पर वहीं से सीखना जारी रख सकते हैं जहाँ आपने छोड़ा था।
क्या मैं इस Pandas & NumPy Academy पाठ में कोड लिख और चला सकता हूँ?
हाँ। हर Pandas & NumPy Academy पाठ में एक अंतर्निर्मित कोड संपादक शामिल है, जिससे आप सीधे अपने ब्राउज़र में वास्तविक कोड लिख और चला सकते हैं और तुरंत एआई प्रतिक्रिया पा सकते हैं—स्थानीय सेटअप की आवश्यकता नहीं है।
इस पाठ्यक्रम के सभी पाठ
- कॉलम मानों के आधार पर क्रमबद्ध करना
- इंडेक्स के आधार पर क्रमबद्ध करना
- मानों की रैंकिंग
- इंडेक्स सेट और रीसेट करना