0Pricing
Deep Learning Academy · Lección

Programaciones de la tasa de aprendizaje y warmup

Estrategias por pasos, coseno y warmup

Programaciones de la tasa de aprendizaje y warmup es una lección gratuita de Deep Learning Academy en CoddyKit. Esta es la lección 4 de 4. Puedes leer la lección completa abajo gratuitamente — luego la practicas en el navegador con un editor de código integrado y un tutor de IA 24/7. Forma parte de la ruta de aprendizaje de Deep Learning Academy, y tu progreso se sincroniza en la web y la app de CoddyKit. El curso de Deep Learning Academy incluye 4 lecciones en total.

Partes de esta lección aún no han sido traducidas y se muestran en inglés.

A Rate That Changes

One fixed learning rate is rarely ideal for the whole run. A schedule changes it over time, usually large early and small near the end.

Why Decay Helps

Big early steps cover ground fast. Shrinking the rate later lets the model settle precisely into a minimum instead of bouncing around it.

Step Decay

The simplest schedule is step decay: drop the rate by a fixed factor every set number of epochs, like halving it every thirty epochs.

sched = torch.optim.lr_scheduler.StepLR(opt, step_size=30, gamma=0.5)

Cosine Annealing

Cosine schedules glide the rate down a smooth curve toward zero. The gentle, gradual decay is a favorite for training modern networks.

sched = torch.optim.lr_scheduler.CosineAnnealingLR(opt, T_max=100)

The Cold-Start Problem

Fresh weights are fragile. A full-size rate on step one can blow up the loss, especially with large batches or deep transformers.

Warmup Eases In

Warmup ramps the rate up from near zero over the first few hundred steps. This gentle start keeps early training stable before full speed.

Warmup Then Decay

The classic recipe is warmup followed by decay: climb to the peak rate, then ride a cosine curve down. It is the go-to for big models.

Step the Scheduler

A scheduler does nothing until you call step() on it, usually once per epoch right after the optimizer updates the weights.

opt.step()
sched.step()

Watch the Current Rate

Log the live learning rate while you train. Seeing it warm up and decay confirms the schedule fires when expected and helps you debug.

print(sched.get_last_lr())

Plateau-Based Decay

Prefer reacting to results? ReduceLROnPlateau drops the rate only when validation loss stops improving, no fixed timetable needed.

sched = torch.optim.lr_scheduler.ReduceLROnPlateau(opt)

Schedules Are Free Wins

A good schedule often boosts final accuracy with zero extra data. Warmup plus cosine decay is a strong, safe default to start from.

Quick Check

Confirm what warmup is for.

Recap

A schedule shrinks the learning rate over time so the model settles cleanly, while warmup ramps it up first for stable starts. Warmup plus cosine is a great default. 📉

Preguntas frecuentes

¿La lección «Programaciones de la tasa de aprendizaje y warmup» es gratis?

Sí — el texto completo de «Programaciones de la tasa de aprendizaje y warmup» es gratis para leer aquí en la web. Para practicarla de forma interactiva (editor de código integrado y tutor de IA 24/7) y desbloquear el resto del curso de Deep Learning Academy, actualiza a CoddyKit PRO. El curso de Deep Learning Academy incluye 4 lecciones en total.

¿Qué aprenderé en «Programaciones de la tasa de aprendizaje y warmup»?

Estrategias por pasos, coseno y warmup Practicas Deep Learning Academy con código real que ejecutas directamente en el navegador, y un tutor de IA 24/7 responde tus preguntas mientras trabajas en la lección.

¿Necesito experiencia previa para empezar Deep Learning Academy?

No se requiere experiencia previa. Deep Learning Academy en CoddyKit está estructurado para principiantes hasta estudiantes avanzados, así que puedes empezar aquí o desde el inicio y avanzar a tu ritmo. Esta es la lección 4 de 4.

¿Cuánto tiempo toma la lección «Programaciones de la tasa de aprendizaje y warmup»?

La mayoría de las lecciones de CoddyKit toman alrededor de 5–10 minutos. Cada una es compacta e interactiva, así que avanzas constantemente y retomas exactamente por donde dejaste en la web y la app.

¿Puedo escribir y ejecutar código en esta lección de Deep Learning Academy?

Sí. Cada lección de Deep Learning Academy incluye un editor de código integrado, así que escribes y ejecutas código real directamente en tu navegador y obtienes retroalimentación instantánea de IA — sin configuración local necesaria.

Todas las lecciones de este curso

  1. SGD con momentum
  2. Adam y AdamW explicados
  3. Weight decay frente a regularización L2
  4. Programaciones de la tasa de aprendizaje y warmup
← Volver a Deep Learning Academy