0Pricing
Deep Learning Academy · Aula

ReLU e Suas Parentes Leaky e GELU

A ativação padrão e suas variantes modernas

ReLU e Suas Parentes Leaky e GELU é uma aula grátis de Deep Learning Academy no CoddyKit. Esta é a aula 2 de 4. Você pode ler a aula completa abaixo gratuitamente — depois pratica ao vivo no navegador com um editor de código integrado e um tutor de IA 24/7. Faz parte do caminho de aprendizado de Deep Learning Academy, e seu progresso é sincronizado entre a web e o app CoddyKit. O curso de Deep Learning Academy inclui 4 aulas no total.

Partes desta aula ainda não foram traduzidas e aparecem em inglês.

Meet ReLU

The most popular activation is ReLU: it keeps positive values and turns every negative one into zero. Simple and fast. ⚡

import torch.nn.functional as F
y = F.relu(x)   # max(0, x), elementwise

Why It Caught On

ReLU is cheap to compute and its gradient is a clean 1 for positives. That keeps signals flowing and makes deep nets train quickly.

The Math

ReLU is just max(0, x). Positive inputs pass straight through; negatives flatten to zero. That single hinge is the whole trick.

The Dying ReLU Problem

If a neuron always outputs zero, its gradient is zero too, so it stops learning forever. We call this a dead neuron. 💀

Leaky ReLU to the Rescue

Leaky ReLU lets a tiny slope through for negatives instead of a hard zero. That small leak keeps dead neurons alive.

y = F.leaky_relu(x, negative_slope=0.01)

Parametric ReLU

PReLU goes further: it learns the negative slope during training instead of fixing it. The network tunes the leak itself.

Meet GELU

GELU smooths the ReLU corner into a soft curve. It gates inputs by how likely they are to be useful, not with a hard cutoff.

y = F.gelu(x)

Why Transformers Love GELU

Modern models like transformers favor GELU because its smooth shape gives gentler gradients. That often means steadier training. 🤖

SiLU and Friends

SiLU, also called Swish, multiplies the input by its own sigmoid. Like GELU, it is smooth and frequently edges out plain ReLU.

A Sensible Default

Start with ReLU for hidden layers; it is fast and reliable. Reach for Leaky ReLU or GELU only if you see dead neurons or want extra smoothness.

Use It as a Layer

You can drop these in as modules inside a model, not just as functions. That makes them easy to chain in nn.Sequential.

import torch.nn as nn
net = nn.Sequential(nn.Linear(4, 8), nn.ReLU())

Quick Check

Think about a neuron that always lands in the negative zone.

Recap

ReLU is the fast default, but it can let neurons die. Leaky ReLU, PReLU, and GELU smooth or leak the negatives to keep learning healthy. 🌟

Perguntas Frequentes

A aula “ReLU e Suas Parentes Leaky e GELU” é grátis?

Sim — o texto completo de “ReLU e Suas Parentes Leaky e GELU” é grátis para ler aqui na web. Para praticá-la interativamente (um editor de código integrado e um tutor de IA 24/7) e desbloquear o restante do curso de Deep Learning Academy, atualize para CoddyKit PRO. O curso de Deep Learning Academy inclui 4 aulas no total.

O que vou aprender em “ReLU e Suas Parentes Leaky e GELU”?

A ativação padrão e suas variantes modernas Você pratica Deep Learning Academy com código prático que executa diretamente no navegador, e um tutor de IA 24/7 responde suas dúvidas enquanto trabalha na aula.

Preciso ter experiência prévia para começar Deep Learning Academy?

Nenhuma experiência prévia é necessária. Deep Learning Academy no CoddyKit é estruturado para alunos iniciantes até avançados, então você pode começar aqui ou desde o início e aprender no seu ritmo. Esta é a aula 2 de 4.

Quanto tempo leva a aula “ReLU e Suas Parentes Leaky e GELU”?

A maioria das aulas CoddyKit leva cerca de 5–10 minutos. Cada uma é compacta e interativa, então você faz progresso constante e retoma exatamente de onde parou entre web e app.

Posso escrever e executar código nesta aula de Deep Learning Academy?

Sim. Cada aula de Deep Learning Academy inclui um editor de código integrado, então você escreve e executa código real direto no navegador e recebe feedback de IA instantaneamente — nenhuma configuração local necessária.

Todas as aulas deste curso

  1. Por Que a Não Linearidade Libera o Verdadeiro Poder
  2. ReLU e Suas Parentes Leaky e GELU
  3. Sigmoid e Tanh: Comprimindo para um Intervalo
  4. Softmax para Probabilidades
← Voltar para Deep Learning Academy