0Pricing
CUDA Academy · Aula

Cronometre seu primeiro ganho de velocidade

Meça a GPU e a CPU na mesma tarefa.

Cronometre seu primeiro ganho de velocidade é uma aula grátis de CUDA Academy no CoddyKit. Esta é a aula 4 de 4. Você pode ler a aula completa abaixo gratuitamente — depois pratica ao vivo no navegador com um editor de código integrado e um tutor de IA 24/7. Faz parte do caminho de aprendizado de CUDA Academy, e seu progresso é sincronizado entre a web e o app CoddyKit. O curso de CUDA Academy inclui 4 aulas no total.

Partes desta aula ainda não foram traduzidas e aparecem em inglês.

Now Measure the Win

You have a correct kernel, so the fun question is how much faster it is than the CPU. Measuring turns a guess into a real number. ⏱️

Time the CPU Baseline

Start with a fair baseline: time the same add running in a plain CPU loop. You need something to compare the GPU against.

Don't Time with the Clock Wrong

Avoid timing across a cudaMemcpy you forgot to wait for. The launch is asynchronous, so naive timers can report nonsense.

Use CUDA Events

The right tool is a pair of cudaEvent objects. They record GPU timestamps directly on the device for accurate kernel timing.

cudaEvent_t start, stop;
cudaEventCreate(&start);
cudaEventCreate(&stop);

Bracket the Kernel

Record start just before the launch and stop just after. The events queue in the stream alongside your kernel.

cudaEventRecord(start);
vecAdd<<<blocks, threads>>>(d_A, d_B, d_C, n);
cudaEventRecord(stop);

Wait, Then Read

Call cudaEventSynchronize on stop so the CPU waits for the GPU. Only then is the elapsed time ready to read.

cudaEventSynchronize(stop);
float ms = 0;
cudaEventElapsedTime(&ms, start, stop);

Warm Up First

The very first launch pays one-time setup costs. Run a throwaway warm-up launch before timing so those costs do not skew your number.

Average Several Runs

A single sample is noisy, so time the kernel a few times and take the average. Stable numbers make speedup claims trustworthy.

Count the Copies

Be honest about what you measure. Kernel-only time looks amazing, but real speedup must include the PCIe transfers too.

Compute the Speedup

Speedup is simply CPU time divided by GPU time. A small array may even be slower on the GPU once copies are counted.

float speedup = cpu_ms / gpu_ms;

Bigger Arrays Win More

The GPU shines when there is enough work to hide transfer cost. Grow n and watch the speedup climb as parallelism dominates. 📈

Quick Check

Which tool gives accurate GPU kernel timing?

Recap

You measured your first speedup with CUDA events: warm up, bracket the kernel, sync, and divide CPU by GPU time. Bigger problems, bigger wins. 🎉

Perguntas Frequentes

A aula “Cronometre seu primeiro ganho de velocidade” é grátis?

Sim — o texto completo de “Cronometre seu primeiro ganho de velocidade” é grátis para ler aqui na web. Para praticá-la interativamente (um editor de código integrado e um tutor de IA 24/7) e desbloquear o restante do curso de CUDA Academy, atualize para CoddyKit PRO. O curso de CUDA Academy inclui 4 aulas no total.

O que vou aprender em “Cronometre seu primeiro ganho de velocidade”?

Meça a GPU e a CPU na mesma tarefa. Você pratica CUDA Academy com código prático que executa diretamente no navegador, e um tutor de IA 24/7 responde suas dúvidas enquanto trabalha na aula.

Preciso ter experiência prévia para começar CUDA Academy?

Nenhuma experiência prévia é necessária. CUDA Academy no CoddyKit é estruturado para alunos iniciantes até avançados, então você pode começar aqui ou desde o início e aprender no seu ritmo. Esta é a aula 4 de 4.

Quanto tempo leva a aula “Cronometre seu primeiro ganho de velocidade”?

A maioria das aulas CoddyKit leva cerca de 5–10 minutos. Cada uma é compacta e interativa, então você faz progresso constante e retoma exatamente de onde parou entre web e app.

Posso escrever e executar código nesta aula de CUDA Academy?

Sim. Cada aula de CUDA Academy inclui um editor de código integrado, então você escreve e executa código real direto no navegador e recebe feedback de IA instantaneamente — nenhuma configuração local necessária.

Todas as aulas deste curso

  1. O núcleo de soma de vetores
  2. Conecte o lado do hospedeiro
  3. Verifique o resultado na CPU
  4. Cronometre seu primeiro ganho de velocidade
← Voltar para CUDA Academy