0Pricing
CUDA Academy · Lektion

Transfers vom Device zum Host

Laden Sie Ergebnisse zurück auf die CPU.

Transfers vom Device zum Host ist eine kostenlose CUDA Academy-Lektion auf CoddyKit. Dies ist Lektion 2 von 4. Du kannst die komplette Lektion unten kostenlos lesen – dann übst du sie direkt im Browser mit einem integrierten Code-Editor und einem KI-Tutor rund um die Uhr. Sie ist Teil des CUDA Academy-Lernpfads, und dein Fortschritt wird über Web und CoddyKit-App synchronisiert. Der CUDA Academy-Kurs umfasst insgesamt 4 Lektionen.

Teile dieser Lektion wurden noch nicht übersetzt und werden auf Englisch angezeigt.

Bringing Results Home

Your kernel just filled a device buffer with answers, but the CPU cannot read GPU memory directly. You must download the results back. 📥

The Download Direction

To pull data back you use cudaMemcpyDeviceToHost. Now the source is device memory and the destination is host RAM.

cudaMemcpy(h_c, d_c, bytes, cudaMemcpyDeviceToHost);

Destination Still First

The rule never changes: destination argument first, source second. For a download, the host pointer leads.

cudaMemcpy(h_dst, d_src, bytes, cudaMemcpyDeviceToHost);

Host Buffer Must Exist

The destination needs real CPU memory waiting for it. Allocate host storage with malloc or new before you copy back.

float* h_c = (float*)malloc(bytes);

Wait for the Kernel

A blocking cudaMemcpy waits for prior work in the default stream, so the kernel finishes before any results are read.

myKernel<<<blocks, threads>>>(d_c, n);
cudaMemcpy(h_c, d_c, bytes, cudaMemcpyDeviceToHost);

Same Byte Count

Download exactly as many bytes as you uploaded and computed. Mismatched sizes give you truncated or garbage results.

size_t bytes = n * sizeof(float);

Now You Can Read It

Once the copy returns, the values live in normal CPU memory. You can print, check, or save them like any host array.

printf("%f\n", h_c[0]);

The Round Trip

A full GPU job is a round trip: upload inputs, launch the kernel, then download outputs. Each leg uses cudaMemcpy.

Verify Before You Trust

After downloading, compare the GPU result against a quick CPU reference. Verification catches bugs before they spread. ✅

Free What You Allocated

When results are home, release both sides: cudaFree the device buffer and free the host buffer to avoid leaks.

cudaFree(d_c);
free(h_c);

Downloads Cost Time Too

The return trip crosses the same slow bus, so only copy back the results you actually need on the host.

Quick Check

Your kernel wrote results into device buffer d_c. How do you read them on the CPU?

Recap

You learned the return trip with cudaMemcpyDeviceToHost: allocate host storage, host pointer first, wait for the kernel, then verify and free. 🎉

Häufig gestellte Fragen

Ist die Lektion „Transfers vom Device zum Host“ kostenlos?

Ja — der vollständige Text von „Transfers vom Device zum Host“ ist hier im Web kostenlos zu lesen. Um sie interaktiv zu üben (integrierter Code-Editor und 24/7 KI-Tutor) und den Rest des CUDA Academy-Kurses freizuschalten, upgrade auf CoddyKit PRO. Der CUDA Academy-Kurs umfasst insgesamt 4 Lektionen.

Was lerne ich in „Transfers vom Device zum Host“?

Laden Sie Ergebnisse zurück auf die CPU. Du übst CUDA Academy mit praktischem Code, den du direkt im Browser ausführst, und ein 24/7 KI-Tutor beantwortet deine Fragen während du die Lektion bearbeitest.

Brauche ich Erfahrung, um CUDA Academy zu starten?

Keine Vorkenntnisse erforderlich. CUDA Academy auf CoddyKit ist für Anfänger bis fortgeschrittene Lernende strukturiert, sodass du hier starten oder von Anfang an beginnen und in deinem eigenen Tempo voranschreiten kannst. Dies ist Lektion 2 von 4.

Wie lange dauert die Lektion „Transfers vom Device zum Host“?

Die meisten CoddyKit-Lektionen dauern etwa 5–10 Minuten. Jede ist kompakt und interaktiv, sodass du stetig Fortschritte machst und genau dort weitermachst, wo du aufgehört hast – im Web und in der App.

Kann ich in dieser CUDA Academy-Lektion Code schreiben und ausführen?

Ja. Jede CUDA Academy-Lektion enthält einen integrierten Code-Editor, sodass du echten Code direkt in deinem Browser schreibst und ausführst und sofort KI-Feedback erhältst — ohne lokale Einrichtung erforderlich.

Alle Lektionen in diesem Kurs

  1. Transfers vom Host zum Device
  2. Transfers vom Device zum Host
  3. Das Enum für die Kopierrichtung
  4. Der PCIe-Transferengpass
← Zurück zu CUDA Academy