تحليل عنق الزجاجة
حدّد مواضع استهلاك الوقت والذاكرة
تحليل عنق الزجاجة درس مجاني في Deep Learning Academy على CoddyKit. هذا هو الدرس 3 من أصل 4. يمكنك قراءة الدرس كاملاً أدناه مجاناً — ثم تمرن عليه مباشرة في المتصفح باستخدام محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7. هذا الدرس جزء من مسار التعلم في Deep Learning Academy، وتقدمك يتزامن عبر الويب وتطبيق CoddyKit. تتضمن دورة Deep Learning Academy 4 دروس في المجموع.
بعض أجزاء هذا الدرس لم تُترجم بعد وتظهر باللغة الإنجليزية.
Why Profile First
Before optimizing, find out where time actually goes. Guessing wastes effort, while a quick profile shows you the real slow spots.
Two Common Bottlenecks
Training usually stalls in one of two places: the GPU compute doing math, or the data pipeline feeding it. Knowing which one matters.
Time It Crudely First
Start simple by timing a loop section with the clock. A rough perf_counter reading often points you to the right area in seconds.
import time
t = time.perf_counter()
# run one batch
print(time.perf_counter() - t)GPU Work Is Async
CUDA runs in the background, so naive timers lie. Call synchronize first to make sure the GPU has truly finished before you read the clock.
torch.cuda.synchronize()The Built-In Profiler
For real detail, use the torch.profiler context manager. It records how long every operation takes on both CPU and GPU.
from torch.profiler import profileWrap the Code to Profile
Run the part you care about inside a profile block. Choosing both CPU and CUDA activities captures the whole picture.
with profile(activities=[ProfilerActivity.CPU, ProfilerActivity.CUDA]) as prof:
model(x)Read the Table
Print results sorted by cost to see the heaviest ops at the top. The key_averages table groups identical operations together.
print(prof.key_averages().table(sort_by='cuda_time_total'))Spot a Data Bottleneck
If the GPU often sits idle waiting, your DataLoader is too slow. More workers or cached data usually fixes that gap.
Spot a Compute Bottleneck
If one matmul or conv dominates the table, the limit is raw compute. Mixed precision or a smaller model is the lever to pull.
Watch Memory Too
The profiler can also report peak memory. Tracking profile_memory reveals which layers eat the most, guiding what to trim.
with profile(profile_memory=True) as prof:
model(x)Measure, Change, Re-Measure
Optimization is a loop: profile, make one change, then profile again. Trust numbers, not hunches, to confirm a fix actually helped.
Quick Check
Your GPU often sits idle between batches. What is the likely bottleneck?
Recap
Profile before you tune: synchronize for honest timings, use torch.profiler to find the heaviest ops, then fix data or compute and measure again. 🔍
الأسئلة الشائعة
هل درس «تحليل عنق الزجاجة» مجاني؟
نعم — نص درس «تحليل عنق الزجاجة» كامل متاح مجاناً هنا على الويب. لتمرينه بشكل تفاعلي (محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7) وفتح باقي دورة Deep Learning Academy، انتقل إلى CoddyKit PRO. تتضمن دورة Deep Learning Academy 4 دروس في المجموع.
ماذا ستتعلم في «تحليل عنق الزجاجة»؟
حدّد مواضع استهلاك الوقت والذاكرة تتمرن على Deep Learning Academy مع أكواد عملية تشغلها مباشرة في المتصفح، ومدرس ذكاء اصطناعي متاح 24/7 يجيب على أسئلتك أثناء عملك.
هل أحتاج إلى خبرة سابقة لأبدأ Deep Learning Academy؟
لا تُشترط خبرة سابقة. Deep Learning Academy على CoddyKit منظم للمبتدئين حتى المتقدمين، لذا يمكنك البدء من هنا أو من البداية والتقدم بسرعتك الخاصة. هذا هو الدرس 3 من أصل 4.
كم من الوقت يستغرق درس «تحليل عنق الزجاجة»؟
معظم دروس CoddyKit تستغرق حوالي 5–10 دقائق. كل منها موجز وتفاعلي، لذا تحرز تقدماً مستمراً وتستأنف من حيث توقفت عبر الويب والتطبيق.
هل يمكنني كتابة وتشغيل أكواد في درس Deep Learning Academy هذا؟
نعم. كل درس في Deep Learning Academy يتضمن محرر أكواد مدمج، لذا تكتب وتشغل أكواداً حقيقية مباشرة في متصفحك وتحصل على تعليقات فورية من الذكاء الاصطناعي — بدون إعداد محلي.
جميع الدروس في هذه الدورة
- الدقة المختلطة باستخدام autocast وGradScaler
- تجميع التدرجات للدفعات الكبيرة
- تحليل عنق الزجاجة
- تقليل استخدام ذاكرة GPU