0Pricing
CUDA Academy · Aula

A armadilha do fluxo padrão

Veja como o fluxo 0 serializa tudo.

A armadilha do fluxo padrão é uma aula grátis de CUDA Academy no CoddyKit. Esta é a aula 1 de 4. Você pode ler a aula completa abaixo gratuitamente — depois pratica ao vivo no navegador com um editor de código integrado e um tutor de IA 24/7. Faz parte do caminho de aprendizado de CUDA Academy, e seu progresso é sincronizado entre a web e o app CoddyKit. O curso de CUDA Academy inclui 4 aulas no total.

Partes desta aula ainda não foram traduzidas e aparecem em inglês.

What Is a Stream?

A stream is an ordered queue of GPU work. Operations in the same stream run one after another, in the order you issued them. ⏳

The Default Stream

If you never name a stream, every call lands in the default stream, also called stream 0. It is where all your work has secretly been running.

Why It Is a Trap

The default stream is synchronizing: it blocks until other streams finish, and they wait for it. Nothing overlaps, so your GPU sits idle between steps.

Everything Serializes

Issue a copy, a kernel, then another copy on stream 0 and they run strictly back to back. The GPU can never start step two before step one ends.

cudaMemcpy(d_a, h_a, n, cudaMemcpyHostToDevice);
kernel<<<grid, block>>>(d_a);
cudaMemcpy(h_a, d_a, n, cudaMemcpyDeviceToHost);

Wasted Hardware

Modern GPUs have separate engines for compute and for copying. On the default stream those engines take turns instead of working together.

The Hidden Cost of Copies

Data transfers over PCIe are slow. When copies cannot overlap with compute, that transfer time is added directly onto your total runtime.

Implicit Synchronization

Many default-stream calls are blocking on the host too. cudaMemcpy returns only after the copy is done, stalling your CPU.

The Legacy Behavior

By default, work in any stream waits for stream 0, and stream 0 waits for them. This legacy rule quietly destroys concurrency you thought you had.

A Telltale Profile

In a profiler the default-stream trap shows up as a single busy lane with gaps, while copy and compute engines sit idle waiting for each other.

The Way Out

The fix is to issue independent work on separate streams. Then copies and kernels can run at the same time and fill those idle gaps.

Per-Thread Default Streams

Compiling with --default-stream per-thread gives each host thread its own default stream, easing the trap without rewriting every call.

nvcc --default-stream per-thread app.cu -o app

Quick Check

Why does putting all work on the default stream hurt performance?

Recap

The default stream serializes your GPU work and blocks overlap. To go faster, you will move independent tasks onto their own streams next. 🚀

Perguntas Frequentes

A aula “A armadilha do fluxo padrão” é grátis?

Sim — o texto completo de “A armadilha do fluxo padrão” é grátis para ler aqui na web. Para praticá-la interativamente (um editor de código integrado e um tutor de IA 24/7) e desbloquear o restante do curso de CUDA Academy, atualize para CoddyKit PRO. O curso de CUDA Academy inclui 4 aulas no total.

O que vou aprender em “A armadilha do fluxo padrão”?

Veja como o fluxo 0 serializa tudo. Você pratica CUDA Academy com código prático que executa diretamente no navegador, e um tutor de IA 24/7 responde suas dúvidas enquanto trabalha na aula.

Preciso ter experiência prévia para começar CUDA Academy?

Nenhuma experiência prévia é necessária. CUDA Academy no CoddyKit é estruturado para alunos iniciantes até avançados, então você pode começar aqui ou desde o início e aprender no seu ritmo. Esta é a aula 1 de 4.

Quanto tempo leva a aula “A armadilha do fluxo padrão”?

A maioria das aulas CoddyKit leva cerca de 5–10 minutos. Cada uma é compacta e interativa, então você faz progresso constante e retoma exatamente de onde parou entre web e app.

Posso escrever e executar código nesta aula de CUDA Academy?

Sim. Cada aula de CUDA Academy inclui um editor de código integrado, então você escreve e executa código real direto no navegador e recebe feedback de IA instantaneamente — nenhuma configuração local necessária.

Todas as aulas deste curso

  1. A armadilha do fluxo padrão
  2. Crie e use fluxos
  3. Eventos para cronometragem e sincronização
  4. Sobreponha cópia e cálculo
← Voltar para CUDA Academy