0Pricing
CUDA Academy · Aula

Perfilar, otimizar, entregar

Fechando o ciclo com métricas do Nsight.

Perfilar, otimizar, entregar é uma aula grátis de CUDA Academy no CoddyKit. Esta é a aula 4 de 4. Você pode ler a aula completa abaixo gratuitamente — depois pratica ao vivo no navegador com um editor de código integrado e um tutor de IA 24/7. Faz parte do caminho de aprendizado de CUDA Academy, e seu progresso é sincronizado entre a web e o app CoddyKit. O curso de CUDA Academy inclui 4 aulas no total.

Partes desta aula ainda não foram traduzidas e aparecem em inglês.

Measure, Do Not Guess

Optimization starts with data, not hunches. Always profile first to learn where the time really goes before you change a single line. 🔍

Start With the Timeline

Open Nsight Systems to see the whole run: kernels, copies, and gaps. The timeline shows whether transfers and compute actually overlap.

Find the Hot Kernel

One stage usually dominates. The timeline reveals the longest-running kernel, and that is exactly where your tuning effort should go first.

Zoom In With Nsight Compute

For the hot kernel, switch to Nsight Compute. It reports per-kernel metrics like memory throughput, stalls, and achieved occupancy.

Read the Roofline

The roofline tells you if a kernel is memory-bound or compute-bound. That single answer decides which optimizations are even worth trying.

Fix Memory-Bound Kernels

If you are memory-bound, chase coalescing, shared-memory tiling, and vectorized loads. Feeding data faster is the only path to more speed.

Fix Compute-Bound Kernels

If you are compute-bound, raise ILP with unrolling, cut redundant math, or move to lower precision. Here the arithmetic units are the limit.

Label Your Code With NVTX

Wrap pipeline stages in NVTX ranges so the timeline shows named bars instead of anonymous kernels. Readable profiles are faster to debug.

nvtxRangePush("blur");
blurKernel<<<grid, block>>>(d_in, d_out);
nvtxRangePop();

Change One Thing at a Time

Apply a single optimization, then re-profile. This tight loop proves each change helped and stops you from chasing two effects at once.

Know When to Stop

When a kernel nears its roofline limit, more tuning yields little. Recognize diminishing returns and spend your time on the next bottleneck.

Validate, Then Ship

Before release, confirm correctness against a reference and check the speedup is stable. A fast but wrong pipeline is not shippable. 🚢

Quick Check

Nsight Compute says your hot kernel is memory-bound. Which fix should you try first?

Recap

You learned to profile with Nsight, read the roofline to classify kernels, apply targeted fixes one at a time, and validate before shipping a fast, correct pipeline. 🏁

Perguntas Frequentes

A aula “Perfilar, otimizar, entregar” é grátis?

Sim — o texto completo de “Perfilar, otimizar, entregar” é grátis para ler aqui na web. Para praticá-la interativamente (um editor de código integrado e um tutor de IA 24/7) e desbloquear o restante do curso de CUDA Academy, atualize para CoddyKit PRO. O curso de CUDA Academy inclui 4 aulas no total.

O que vou aprender em “Perfilar, otimizar, entregar”?

Fechando o ciclo com métricas do Nsight. Você pratica CUDA Academy com código prático que executa diretamente no navegador, e um tutor de IA 24/7 responde suas dúvidas enquanto trabalha na aula.

Preciso ter experiência prévia para começar CUDA Academy?

Nenhuma experiência prévia é necessária. CUDA Academy no CoddyKit é estruturado para alunos iniciantes até avançados, então você pode começar aqui ou desde o início e aprender no seu ritmo. Esta é a aula 4 de 4.

Quanto tempo leva a aula “Perfilar, otimizar, entregar”?

A maioria das aulas CoddyKit leva cerca de 5–10 minutos. Cada uma é compacta e interativa, então você faz progresso constante e retoma exatamente de onde parou entre web e app.

Posso escrever e executar código nesta aula de CUDA Academy?

Sim. Cada aula de CUDA Academy inclui um editor de código integrado, então você escreve e executa código real direto no navegador e recebe feedback de IA instantaneamente — nenhuma configuração local necessária.

Todas as aulas deste curso

  1. Projetando o pipeline de processamento
  2. Fundindo filtros em um único kernel
  3. Transmitindo blocos para imagens grandes
  4. Perfilar, otimizar, entregar
← Voltar para CUDA Academy