Fundindo filtros em um único kernel
Reduzindo lançamentos e tráfego global.
Fundindo filtros em um único kernel é uma aula grátis de CUDA Academy no CoddyKit. Esta é a aula 2 de 4. Você pode ler a aula completa abaixo gratuitamente — depois pratica ao vivo no navegador com um editor de código integrado e um tutor de IA 24/7. Faz parte do caminho de aprendizado de CUDA Academy, e seu progresso é sincronizado entre a web e o app CoddyKit. O curso de CUDA Academy inclui 4 aulas no total.
Partes desta aula ainda não foram traduzidas e aparecem em inglês.
Why Fuse At All
Running five kernels means five launches and five round-trips to global memory. Fusing filters into one kernel cuts both, often giving a big speedup. 🚀
Launch Overhead Adds Up
Every kernel launch costs a few microseconds. On a tiny image that fixed launch overhead can dwarf the real work, so fewer launches means more useful time.
Global Traffic Is the Enemy
Separate stages write a pixel to global memory then read it right back. Fusing keeps that value in a register, erasing the wasted round-trip entirely.
Load Once Per Thread
In a fused kernel each thread reads its pixel a single time. That one load then feeds every filter in sequence without touching global memory again.
float v = input[idx];Chain Operations in Registers
Apply each filter to the value already in hand. The pixel flows through brightness, gamma, and contrast as plain math, all kept in fast registers.
v = v * brightness;
v = powf(v, gamma);
v = (v - 0.5f) * contrast + 0.5f;Write Once At the End
After the whole chain runs, store the final result. One store replaces the many writes that separate kernels would have made.
output[idx] = v;Pointwise Filters Fuse Cleanly
Filters that touch only one pixel are pointwise and fuse with zero fuss. Brightness, gamma, and color tweaks are the easiest wins to combine.
Neighborhood Filters Are Harder
A blur reads nearby pixels, so fusing it needs shared memory tiles, not just registers. Fuse pointwise stages freely and treat stencils with extra care.
Watch Register Pressure
A big fused kernel uses more registers per thread. Too many and occupancy drops or values spill, so fuse aggressively but keep an eye on the cost.
Verify After Fusing
Fusing reorders work, so always compare the fused output against the staged version. A quick diff confirms the math still matches before you celebrate.
One Kernel, Many Filters
The payoff is real: a single launch reads, transforms, and writes each pixel exactly once. That is the heart of a well-tuned fused image kernel.
Quick Check
You merged three pointwise filters into one kernel. What is the main performance win?
Recap
You learned to fuse filters: load once, chain pointwise math in registers, write once. It cuts launches and global traffic, but watch register pressure and verify the output. ✅
Perguntas Frequentes
A aula “Fundindo filtros em um único kernel” é grátis?
Sim — o texto completo de “Fundindo filtros em um único kernel” é grátis para ler aqui na web. Para praticá-la interativamente (um editor de código integrado e um tutor de IA 24/7) e desbloquear o restante do curso de CUDA Academy, atualize para CoddyKit PRO. O curso de CUDA Academy inclui 4 aulas no total.
O que vou aprender em “Fundindo filtros em um único kernel”?
Reduzindo lançamentos e tráfego global. Você pratica CUDA Academy com código prático que executa diretamente no navegador, e um tutor de IA 24/7 responde suas dúvidas enquanto trabalha na aula.
Preciso ter experiência prévia para começar CUDA Academy?
Nenhuma experiência prévia é necessária. CUDA Academy no CoddyKit é estruturado para alunos iniciantes até avançados, então você pode começar aqui ou desde o início e aprender no seu ritmo. Esta é a aula 2 de 4.
Quanto tempo leva a aula “Fundindo filtros em um único kernel”?
A maioria das aulas CoddyKit leva cerca de 5–10 minutos. Cada uma é compacta e interativa, então você faz progresso constante e retoma exatamente de onde parou entre web e app.
Posso escrever e executar código nesta aula de CUDA Academy?
Sim. Cada aula de CUDA Academy inclui um editor de código integrado, então você escreve e executa código real direto no navegador e recebe feedback de IA instantaneamente — nenhuma configuração local necessária.
Todas as aulas deste curso
- Projetando o pipeline de processamento
- Fundindo filtros em um único kernel
- Transmitindo blocos para imagens grandes
- Perfilar, otimizar, entregar