0Pricing
Deep Learning Academy · Lección

Perfilar el cuello de botella

Descubra en qué se emplean el tiempo y la memoria

Perfilar el cuello de botella es una lección gratuita de Deep Learning Academy en CoddyKit. Esta es la lección 3 de 4. Puedes leer la lección completa abajo gratuitamente — luego la practicas en el navegador con un editor de código integrado y un tutor de IA 24/7. Forma parte de la ruta de aprendizaje de Deep Learning Academy, y tu progreso se sincroniza en la web y la app de CoddyKit. El curso de Deep Learning Academy incluye 4 lecciones en total.

Partes de esta lección aún no han sido traducidas y se muestran en inglés.

Why Profile First

Before optimizing, find out where time actually goes. Guessing wastes effort, while a quick profile shows you the real slow spots.

Two Common Bottlenecks

Training usually stalls in one of two places: the GPU compute doing math, or the data pipeline feeding it. Knowing which one matters.

Time It Crudely First

Start simple by timing a loop section with the clock. A rough perf_counter reading often points you to the right area in seconds.

import time
t = time.perf_counter()
# run one batch
print(time.perf_counter() - t)

GPU Work Is Async

CUDA runs in the background, so naive timers lie. Call synchronize first to make sure the GPU has truly finished before you read the clock.

torch.cuda.synchronize()

The Built-In Profiler

For real detail, use the torch.profiler context manager. It records how long every operation takes on both CPU and GPU.

from torch.profiler import profile

Wrap the Code to Profile

Run the part you care about inside a profile block. Choosing both CPU and CUDA activities captures the whole picture.

with profile(activities=[ProfilerActivity.CPU, ProfilerActivity.CUDA]) as prof:
    model(x)

Read the Table

Print results sorted by cost to see the heaviest ops at the top. The key_averages table groups identical operations together.

print(prof.key_averages().table(sort_by='cuda_time_total'))

Spot a Data Bottleneck

If the GPU often sits idle waiting, your DataLoader is too slow. More workers or cached data usually fixes that gap.

Spot a Compute Bottleneck

If one matmul or conv dominates the table, the limit is raw compute. Mixed precision or a smaller model is the lever to pull.

Watch Memory Too

The profiler can also report peak memory. Tracking profile_memory reveals which layers eat the most, guiding what to trim.

with profile(profile_memory=True) as prof:
    model(x)

Measure, Change, Re-Measure

Optimization is a loop: profile, make one change, then profile again. Trust numbers, not hunches, to confirm a fix actually helped.

Quick Check

Your GPU often sits idle between batches. What is the likely bottleneck?

Recap

Profile before you tune: synchronize for honest timings, use torch.profiler to find the heaviest ops, then fix data or compute and measure again. 🔍

Preguntas frecuentes

¿La lección «Perfilar el cuello de botella» es gratis?

Sí — el texto completo de «Perfilar el cuello de botella» es gratis para leer aquí en la web. Para practicarla de forma interactiva (editor de código integrado y tutor de IA 24/7) y desbloquear el resto del curso de Deep Learning Academy, actualiza a CoddyKit PRO. El curso de Deep Learning Academy incluye 4 lecciones en total.

¿Qué aprenderé en «Perfilar el cuello de botella»?

Descubra en qué se emplean el tiempo y la memoria Practicas Deep Learning Academy con código real que ejecutas directamente en el navegador, y un tutor de IA 24/7 responde tus preguntas mientras trabajas en la lección.

¿Necesito experiencia previa para empezar Deep Learning Academy?

No se requiere experiencia previa. Deep Learning Academy en CoddyKit está estructurado para principiantes hasta estudiantes avanzados, así que puedes empezar aquí o desde el inicio y avanzar a tu ritmo. Esta es la lección 3 de 4.

¿Cuánto tiempo toma la lección «Perfilar el cuello de botella»?

La mayoría de las lecciones de CoddyKit toman alrededor de 5–10 minutos. Cada una es compacta e interactiva, así que avanzas constantemente y retomas exactamente por donde dejaste en la web y la app.

¿Puedo escribir y ejecutar código en esta lección de Deep Learning Academy?

Sí. Cada lección de Deep Learning Academy incluye un editor de código integrado, así que escribes y ejecutas código real directamente en tu navegador y obtienes retroalimentación instantánea de IA — sin configuración local necesaria.

Todas las lecciones de este curso

  1. Precisión mixta con autocast y GradScaler
  2. Acumulación de gradientes para lotes grandes
  3. Perfilar el cuello de botella
  4. Reduzca el uso de memoria de la GPU
← Volver a Deep Learning Academy