Kopieren und Berechnen überlappen
Verbergen Sie Transfers hinter Kernels.
Kopieren und Berechnen überlappen ist eine kostenlose CUDA Academy-Lektion auf CoddyKit. Dies ist Lektion 4 von 4. Du kannst die komplette Lektion unten kostenlos lesen – dann übst du sie direkt im Browser mit einem integrierten Code-Editor und einem KI-Tutor rund um die Uhr. Sie ist Teil des CUDA Academy-Lernpfads, und dein Fortschritt wird über Web und CoddyKit-App synchronisiert. Der CUDA Academy-Kurs umfasst insgesamt 4 Lektionen.
Teile dieser Lektion wurden noch nicht übersetzt und werden auf Englisch angezeigt.
The Big Idea
Real speedups come from doing two things at once: copying one chunk of data while the GPU computes on another. 🔀
Separate Hardware Engines
A GPU has distinct copy engines and compute units. Overlap works because these can run truly in parallel, not just take turns.
Pinned Memory Is Required
Async copies need pinned host memory from cudaMallocHost. Ordinary pageable malloc forces a blocking copy that cannot overlap.
cudaMallocHost(&h_data, bytes);Async Copies
Use cudaMemcpyAsync with a stream so the transfer returns immediately and joins that stream's queue alongside kernels.
cudaMemcpyAsync(d_in, h_in, bytes, cudaMemcpyHostToDevice, s);Chunk the Work
Split the array into pieces and give each piece its own stream. While stream 0 computes, stream 1 can already be copying.
Copy, Compute, Copy Back
Each chunk does the same three steps in its stream: upload input, run the kernel, download output, all without blocking the host.
cudaMemcpyAsync(d, h, n, H2D, s);
k<<<g, b, 0, s>>>(d);
cudaMemcpyAsync(h, d, n, D2H, s);How the Overlap Forms
Because chunks live in different streams, the GPU can copy chunk two while still computing chunk one. The timeline fills up. 📊
Hiding the PCIe Cost
Transfers no longer add to total time; they hide behind compute. Ideally your runtime shrinks to roughly the longer of copy or kernel.
Watch the Default Stream
One stray default-stream call mid-loop can serialize everything again. Keep every async op tied to a real, non-default stream.
Verify in a Profiler
Open Nsight Systems and look for copy and compute lanes that overlap in time. Stacked, busy lanes mean your concurrency is real.
Synchronize at the End
After issuing all chunks, call cudaDeviceSynchronize once so every stream finishes before you read the final results on the host.
cudaDeviceSynchronize();Quick Check
What does an async transfer require to overlap with compute?
Recap
By chunking work across streams with pinned memory and async copies, you hide transfers behind compute and keep the GPU truly busy. 🚀
Häufig gestellte Fragen
Ist die Lektion „Kopieren und Berechnen überlappen“ kostenlos?
Ja — der vollständige Text von „Kopieren und Berechnen überlappen“ ist hier im Web kostenlos zu lesen. Um sie interaktiv zu üben (integrierter Code-Editor und 24/7 KI-Tutor) und den Rest des CUDA Academy-Kurses freizuschalten, upgrade auf CoddyKit PRO. Der CUDA Academy-Kurs umfasst insgesamt 4 Lektionen.
Was lerne ich in „Kopieren und Berechnen überlappen“?
Verbergen Sie Transfers hinter Kernels. Du übst CUDA Academy mit praktischem Code, den du direkt im Browser ausführst, und ein 24/7 KI-Tutor beantwortet deine Fragen während du die Lektion bearbeitest.
Brauche ich Erfahrung, um CUDA Academy zu starten?
Keine Vorkenntnisse erforderlich. CUDA Academy auf CoddyKit ist für Anfänger bis fortgeschrittene Lernende strukturiert, sodass du hier starten oder von Anfang an beginnen und in deinem eigenen Tempo voranschreiten kannst. Dies ist Lektion 4 von 4.
Wie lange dauert die Lektion „Kopieren und Berechnen überlappen“?
Die meisten CoddyKit-Lektionen dauern etwa 5–10 Minuten. Jede ist kompakt und interaktiv, sodass du stetig Fortschritte machst und genau dort weitermachst, wo du aufgehört hast – im Web und in der App.
Kann ich in dieser CUDA Academy-Lektion Code schreiben und ausführen?
Ja. Jede CUDA Academy-Lektion enthält einen integrierten Code-Editor, sodass du echten Code direkt in deinem Browser schreibst und ausführst und sofort KI-Feedback erhältst — ohne lokale Einrichtung erforderlich.
Alle Lektionen in diesem Kurs
- Die Falle des Standard-Streams
- Streams erstellen und verwenden
- Events für Zeitmessung und Synchronisierung
- Kopieren und Berechnen überlappen