أنواع بيانات فعّالة لتقليل الذاكرة
خفّض أنواع الأعمدة العددية إلى int32/float32، وحوّل أعمدة السلاسل النصية إلى Categorical لتقليل ذاكرة DataFrame بنسبة تصل إلى 70%.
أنواع بيانات فعّالة لتقليل الذاكرة درس مجاني في Pandas & NumPy Academy على CoddyKit. هذا هو الدرس 3 من أصل 4. يمكنك قراءة الدرس كاملاً أدناه مجاناً — ثم تمرن عليه مباشرة في المتصفح باستخدام محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7. هذا الدرس جزء من مسار التعلم في Pandas & NumPy Academy، وتقدمك يتزامن عبر الويب وتطبيق CoddyKit. تتضمن دورة Pandas & NumPy Academy 4 دروس في المجموع.
بعض أجزاء هذا الدرس لم تُترجم بعد وتظهر باللغة الإنجليزية.
Why Data Types Affect Performance
In Pandas, every column has a dtype (data type) that determines how values are stored in memory and how fast operations run. Pandas uses wide types by default when loading data: int64 (8 bytes per value), float64 (8 bytes), and object (variable, often 50-200 bytes per string). For millions of rows, choosing smaller types can reduce memory by 50-80% and speed up operations by 2-5x due to better cache utilisation.
import pandas as pd
import numpy as np
np.random.seed(0)
df = pd.DataFrame({
'id': np.arange(1000000, dtype='int64'),
'score': np.random.uniform(0, 100, 1000000).astype('float64'),
'category': np.random.choice(['A','B','C','D'], 1000000)
})
mem = df.memory_usage(deep=True)
print('Memory per column:')
print(mem)
print(f'Total: {mem.sum() / 1e6:.1f} MB')Downcasting Integer Columns
If a column contains integer values that fit in a smaller range, downcast it from int64 to a smaller integer type. pd.to_numeric(series, downcast='integer') automatically selects the smallest integer type that can hold all values: int8 (–128 to 127, 1 byte), int16 (–32768 to 32767, 2 bytes), int32 (±2 billion, 4 bytes), or stays as int64 if needed. An int8 column uses 8x less memory than int64.
import pandas as pd
import numpy as np
np.random.seed(0)
df = pd.DataFrame({
'age': np.random.randint(18, 90, 500000).astype('int64'),
'score': np.random.randint(0, 100, 500000).astype('int64'),
'large_id': np.random.randint(0, 2**31, 500000).astype('int64')
})
# Downcast integers
for col in df.select_dtypes('int64').columns:
df[col] = pd.to_numeric(df[col], downcast='integer')
print('Dtypes after downcasting:')
print(df.dtypes)
print(f'\nMemory: {df.memory_usage(deep=True).sum()/1e6:.2f} MB')Downcasting Float Columns
pd.to_numeric(series, downcast='float') converts float64 to float32 (4 bytes instead of 8 bytes) wherever precision allows. float32 has about 7 decimal digits of precision vs. 15 for float64. For most analytics tasks (percentages, prices, normalised features), float32 precision is sufficient. For scientific computing requiring high precision, stay with float64. Halving float precision halves memory and improves cache performance.
import pandas as pd
import numpy as np
np.random.seed(0)
df = pd.DataFrame({'lat': np.random.uniform(-90, 90, 1000000),
'lon': np.random.uniform(-180, 180, 1000000),
'temp': np.random.uniform(-40, 50, 1000000)})
print('Before:', df.dtypes.unique())
print(f'Memory: {df.memory_usage(deep=True).sum()/1e6:.1f} MB')
for col in df.select_dtypes('float64').columns:
df[col] = pd.to_numeric(df[col], downcast='float')
print('After:', df.dtypes.unique())
print(f'Memory: {df.memory_usage(deep=True).sum()/1e6:.1f} MB')The Categorical Data Type
The Categorical dtype is the single biggest memory win for string columns with repeated values. Instead of storing the same string (e.g. 'Electronics') thousands of times, Pandas stores a dictionary of unique values plus a compact integer code for each row. A column with 1 million rows and only 50 unique categories goes from ~50 MB (object dtype) to ~1 MB (Categorical) — a 50x reduction in memory.
import pandas as pd
import numpy as np
np.random.seed(0)
categories = ['Electronics', 'Clothing', 'Books', 'Food', 'Furniture',
'Sports', 'Toys', 'Health', 'Garden', 'Automotive']
df = pd.DataFrame({'product_cat': np.random.choice(categories, 1000000)})
mem_before = df.memory_usage(deep=True).sum()
df['product_cat'] = df['product_cat'].astype('category')
mem_after = df.memory_usage(deep=True).sum()
print(f'Object dtype: {mem_before/1e6:.1f} MB')
print(f'Categorical: {mem_after/1e6:.2f} MB')
print(f'Reduction: {mem_before/mem_after:.0f}x')When Categorical Saves Memory
Categorical dtype saves memory only when the column has many fewer unique values than total rows. The break-even point is when the ratio of unique values to total rows (cardinality) is above about 50%: if every row has a unique string, Categorical actually uses more memory than object dtype because it stores both the code array and the categories array. Always check df['col'].nunique() / len(df) before converting — a ratio below 0.5 (and ideally below 0.05) means Categorical will help.
import pandas as pd
import numpy as np
def memory_gain(df, col):
n = len(df)
n_unique = df[col].nunique()
ratio = n_unique / n
mem_obj = df[col].memory_usage(deep=True)
df2 = df.copy()
df2[col] = df2[col].astype('category')
mem_cat = df2[col].memory_usage(deep=True)
print(f'{col}: {n_unique} unique / {n} total (ratio={ratio:.3f})')
print(f' Object: {mem_obj/1e6:.2f} MB → Categorical: {mem_cat/1e6:.2f} MB')
print(f' {"Saves" if mem_cat < mem_obj else "Wastes"} {abs(mem_obj-mem_cat)/1e6:.2f} MB')
np.random.seed(0)
df = pd.DataFrame({
'low_card': np.random.choice(['A','B','C'], 500000), # low cardinality
'high_card': [f'id_{i}' for i in range(500000)] # high cardinality
})
memory_gain(df, 'low_card')
memory_gain(df, 'high_card')Categorical Boosts GroupBy Performance
Beyond memory savings, Categorical dtype also speeds up groupby operations because Pandas can use the integer codes directly rather than hashing strings to find group boundaries. For DataFrames with millions of rows grouped by low-cardinality string columns (like region, product category, or status), converting those columns to Categorical before groupby can yield 2-5x speedups.
import pandas as pd
import numpy as np
import timeit
np.random.seed(0)
df = pd.DataFrame({
'region': np.random.choice(['North','South','East','West'], 1000000),
'sales': np.random.randn(1000000)
})
# Object dtype groupby
t1 = timeit.timeit(lambda: df.groupby('region')['sales'].mean(), number=50)
# Categorical dtype groupby
df2 = df.copy()
df2['region'] = df2['region'].astype('category')
t2 = timeit.timeit(lambda: df2.groupby('region')['sales'].mean(), number=50)
print(f'Object dtype groupby: {t1/50*1000:.2f} ms')
print(f'Categorical groupby: {t2/50*1000:.2f} ms')
print(f'Speedup: {t1/t2:.1f}x')Using boolean dtype for Flag Columns
Binary columns (True/False, 0/1, yes/no) are often stored as object or int64. Converting them to bool dtype uses only 1 byte per value (vs. 8 bytes for int64 or 50+ bytes for string 'yes'/'no'). Use astype(bool) after ensuring the column only contains 0/1 or True/False values. Boolean columns also enable faster filtering because Pandas can use bitwise operations internally.
import pandas as pd
import numpy as np
np.random.seed(0)
df = pd.DataFrame({
'is_premium': np.random.choice([0, 1], 1000000).astype('int64'),
'has_discount': np.random.choice(['yes', 'no'], 1000000)
})
print('Before:')
print(df.dtypes)
print(f'Memory: {df.memory_usage(deep=True).sum()/1e6:.1f} MB')
df['is_premium'] = df['is_premium'].astype(bool)
df['has_discount'] = df['has_discount'].map({'yes': True, 'no': False}).astype(bool)
print('\nAfter:')
print(df.dtypes)
print(f'Memory: {df.memory_usage(deep=True).sum()/1e6:.2f} MB')Automating Type Optimisation
Write a reusable optimise_dtypes(df) function that automatically applies all optimisations: downcast integers, downcast floats, convert low-cardinality objects to Categorical, and convert 0/1 integers to bool. Run this function at data load time to minimise memory from the start. This ensures every team member who loads the same dataset gets the optimised version without remembering to apply each step manually.
import pandas as pd
import numpy as np
def optimise_dtypes(df, cat_threshold=0.5):
'''Reduce DataFrame memory by choosing smaller dtypes.'''
for col in df.columns:
col_type = df[col].dtype
if col_type == 'int64':
df[col] = pd.to_numeric(df[col], downcast='integer')
elif col_type == 'float64':
df[col] = pd.to_numeric(df[col], downcast='float')
elif col_type == 'object':
cardinality = df[col].nunique() / len(df)
if cardinality < cat_threshold:
df[col] = df[col].astype('category')
return df
# Test
np.random.seed(0)
df = pd.DataFrame({'a': np.random.randint(0,100,500000),
'b': np.random.randn(500000),
'c': np.random.choice(['X','Y','Z'],500000)})
before = df.memory_usage(deep=True).sum()
df = optimise_dtypes(df)
after = df.memory_usage(deep=True).sum()
print(f'Before: {before/1e6:.2f} MB → After: {after/1e6:.2f} MB ({before/after:.1f}x reduction)')Unsigned Integer Types for Non-Negative Data
When a column only contains non-negative integers (like IDs, counts, or quantities), use unsigned integer types: uint8 (0–255), uint16 (0–65535), uint32 (0–4 billion), or uint64. Unsigned types have the same byte size as their signed counterparts but can store values up to twice as large in the positive range. For example, a product ID that ranges from 0 to 60,000 fits in uint16 (2 bytes) rather than int32 (4 bytes). Use astype('uint16') after verifying the column has no negative values.
import pandas as pd
import numpy as np
np.random.seed(0)
df = pd.DataFrame({
'product_id': np.random.randint(0, 50000, 1000000).astype('int64'),
'quantity': np.random.randint(0, 255, 1000000).astype('int64')
})
print('Before (int64):', df.memory_usage(deep=True).sum() / 1e6, 'MB')
# product_id fits in uint16 (0-65535)
df['product_id'] = df['product_id'].astype('uint16')
# quantity fits in uint8 (0-255)
df['quantity'] = df['quantity'].astype('uint8')
print('After (uint16+uint8):', df.memory_usage(deep=True).sum() / 1e6, 'MB')
print('\nDtypes:', df.dtypes.to_dict())Checking Memory After Every Optimisation Step
Apply optimisations one at a time and check memory after each step. This reveals exactly how much each technique contributes. A typical large DataFrame might go from 2 GB (all default dtypes) to 600 MB (integer downcasting), then to 300 MB (float downcasting), then to 200 MB (Categorical for string columns). Documenting this breakdown helps justify the extra complexity to team members who see unfamiliar dtypes in the codebase.
import pandas as pd
import numpy as np
np.random.seed(0)
df = pd.DataFrame({
'user_id': np.random.randint(0, 99999, 1000000).astype('int64'),
'age': np.random.randint(18, 80, 1000000).astype('int64'),
'revenue': np.random.uniform(0, 5000, 1000000).astype('float64'),
'country': np.random.choice(['US','UK','DE','FR','JP'], 1000000),
'tier': np.random.choice(['free','basic','pro','enterprise'], 1000000)
})
def mem_mb(df):
return df.memory_usage(deep=True).sum() / 1e6
print(f'Start: {mem_mb(df):.1f} MB')
for c in ['user_id','age']:
df[c] = pd.to_numeric(df[c], downcast='integer')
print(f'After int downcast: {mem_mb(df):.1f} MB')
df['revenue'] = pd.to_numeric(df['revenue'], downcast='float')
print(f'After float downcast: {mem_mb(df):.1f} MB')
for c in ['country','tier']:
df[c] = df[c].astype('category')
print(f'After Categorical: {mem_mb(df):.1f} MB')Saving and Reloading Optimised Types
Optimised dtypes are lost if you save to CSV (which converts everything back to text). To preserve dtypes, save to Parquet format with df.to_parquet('file.parquet') — Parquet preserves int32, float32, and Categorical types. When reloaded with pd.read_parquet, the DataFrame has the same memory-efficient types without re-running the optimisation function. Parquet also loads significantly faster than CSV for large files.
import pandas as pd
import numpy as np
np.random.seed(0)
df = pd.DataFrame({
'age': np.random.randint(18, 80, 100000).astype('int8'), # already optimised
'score': np.random.randn(100000).astype('float32'),
'tier': pd.Categorical(np.random.choice(['free','pro'], 100000))
})
print('Dtypes:', df.dtypes.to_dict())
# Save to Parquet (preserves dtypes)
df.to_parquet('/tmp/optimised.parquet', index=False)
# Reload — dtypes preserved!
df_loaded = pd.read_parquet('/tmp/optimised.parquet')
print('\nLoaded dtypes:', df_loaded.dtypes.to_dict())
print(f'Memory: {df_loaded.memory_usage(deep=True).sum()/1e6:.2f} MB')Quick Check
Test your understanding of memory-efficient data types from this lesson.
Lesson Recap
In this lesson you learned: pd.to_numeric(downcast='integer') reduces integer columns from 8 bytes to 1-4 bytes, pd.to_numeric(downcast='float') halves float memory, Categorical dtype reduces low-cardinality string columns by up to 50x while also speeding up groupby, and Parquet format preserves optimised dtypes across file save/load cycles. Next up we explore chunked reading — processing files that exceed available RAM.
الأسئلة الشائعة
هل درس «أنواع بيانات فعّالة لتقليل الذاكرة» مجاني؟
نعم — نص درس «أنواع بيانات فعّالة لتقليل الذاكرة» كامل متاح مجاناً هنا على الويب. لتمرينه بشكل تفاعلي (محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7) وفتح باقي دورة Pandas & NumPy Academy، انتقل إلى CoddyKit PRO. تتضمن دورة Pandas & NumPy Academy 4 دروس في المجموع.
ماذا ستتعلم في «أنواع بيانات فعّالة لتقليل الذاكرة»؟
خفّض أنواع الأعمدة العددية إلى int32/float32، وحوّل أعمدة السلاسل النصية إلى Categorical لتقليل ذاكرة DataFrame بنسبة تصل إلى 70%. تتمرن على Pandas & NumPy Academy مع أكواد عملية تشغلها مباشرة في المتصفح، ومدرس ذكاء اصطناعي متاح 24/7 يجيب على أسئلتك أثناء عملك.
هل أحتاج إلى خبرة سابقة لأبدأ Pandas & NumPy Academy؟
لا تُشترط خبرة سابقة. Pandas & NumPy Academy على CoddyKit منظم للمبتدئين حتى المتقدمين، لذا يمكنك البدء من هنا أو من البداية والتقدم بسرعتك الخاصة. هذا هو الدرس 3 من أصل 4.
كم من الوقت يستغرق درس «أنواع بيانات فعّالة لتقليل الذاكرة»؟
معظم دروس CoddyKit تستغرق حوالي 5–10 دقائق. كل منها موجز وتفاعلي، لذا تحرز تقدماً مستمراً وتستأنف من حيث توقفت عبر الويب والتطبيق.
هل يمكنني كتابة وتشغيل أكواد في درس Pandas & NumPy Academy هذا؟
نعم. كل درس في Pandas & NumPy Academy يتضمن محرر أكواد مدمج، لذا تكتب وتشغل أكواداً حقيقية مباشرة في متصفحك وتحصل على تعليقات فورية من الذكاء الاصطناعي — بدون إعداد محلي.
جميع الدروس في هذه الدورة
- التوصيف باستخدام timeit وmemory_profiler
- تجنب iterrows وحلقات Python
- أنواع بيانات فعّالة لتقليل الذاكرة
- القراءة على دفعات للملفات الكبيرة