0Pricing
CUDA Academy · Lektion

Vektorisierte Ladevorgänge mit float4

Breitere Transaktionen für mehr Bandbreite

Vektorisierte Ladevorgänge mit float4 ist eine kostenlose CUDA Academy-Lektion auf CoddyKit. Dies ist Lektion 3 von 4. Du kannst die komplette Lektion unten kostenlos lesen – dann übst du sie direkt im Browser mit einem integrierten Code-Editor und einem KI-Tutor rund um die Uhr. Sie ist Teil des CUDA Academy-Lernpfads, und dein Fortschritt wird über Web und CoddyKit-App synchronisiert. Der CUDA Academy-Kurs umfasst insgesamt 4 Lektionen.

Teile dieser Lektion wurden noch nicht übersetzt und werden auf Englisch angezeigt.

Move More Per Instruction

A normal load fetches one value at a time. A vectorized load grabs several adjacent values in a single wider instruction, doing more work per request.

Meet float4

CUDA ships built-in vector types. A float4 packs four floats into one 16-byte bundle you can load or store together.

float4 v = make_float4(1, 2, 3, 4);
float first = v.x; // also .y .z .w

One Load, Four Floats

Reading a float4 from memory pulls four floats in a single transaction. That cuts the number of memory instructions a thread must issue by four.

float4 a = reinterpret_cast<float4*>(in)[i];

Reinterpret the Pointer

To use vector loads on a plain float array, you cast the pointer. reinterpret_cast lets you read the same bytes as float4 elements.

float4* in4 = reinterpret_cast<float4*>(in);
float4 chunk = in4[idx];

Alignment Is Required

A float4 must sit on a 16-byte boundary. If the data is not properly aligned, the load is illegal and your kernel can crash or read garbage.

cudaMalloc Aligns for You

Good news: buffers from cudaMalloc are aligned to at least 256 bytes. So freshly allocated device arrays are safe to read as float4 from the start.

Fewer Requests, More Bandwidth

Wider loads mean fewer outstanding requests for the same data. That can raise achieved bandwidth when a kernel is limited by memory, not math.

Index by Vectors, Not Floats

With float4 each thread now covers four elements. Your global index steps through float4 slots, so the total thread count drops by four.

int i = blockIdx.x * blockDim.x + threadIdx.x;
float4 d = in4[i]; // covers 4 floats

Other Vector Widths

float4 is not the only choice. Types like float2 and int4 give you 8- or 16-byte loads, so you can pick a width that fits your data.

Watch the Leftovers

If your array length is not a multiple of four, handle the extra remainder elements with a normal scalar loop after the vectorized part.

It Costs Registers

A float4 holds four values, so it uses more registers per thread. As always, confirm the bandwidth gain outweighs any drop in occupancy.

Quick Check

You reinterpret a float array as float4 but the kernel crashes. Why?

Recap: Wider Is Faster

You learned that float4 loads four floats at once, cutting instructions and lifting bandwidth. Keep data aligned and handle leftover elements. 📦

Häufig gestellte Fragen

Ist die Lektion „Vektorisierte Ladevorgänge mit float4“ kostenlos?

Ja — der vollständige Text von „Vektorisierte Ladevorgänge mit float4“ ist hier im Web kostenlos zu lesen. Um sie interaktiv zu üben (integrierter Code-Editor und 24/7 KI-Tutor) und den Rest des CUDA Academy-Kurses freizuschalten, upgrade auf CoddyKit PRO. Der CUDA Academy-Kurs umfasst insgesamt 4 Lektionen.

Was lerne ich in „Vektorisierte Ladevorgänge mit float4“?

Breitere Transaktionen für mehr Bandbreite Du übst CUDA Academy mit praktischem Code, den du direkt im Browser ausführst, und ein 24/7 KI-Tutor beantwortet deine Fragen während du die Lektion bearbeitest.

Brauche ich Erfahrung, um CUDA Academy zu starten?

Keine Vorkenntnisse erforderlich. CUDA Academy auf CoddyKit ist für Anfänger bis fortgeschrittene Lernende strukturiert, sodass du hier starten oder von Anfang an beginnen und in deinem eigenen Tempo voranschreiten kannst. Dies ist Lektion 3 von 4.

Wie lange dauert die Lektion „Vektorisierte Ladevorgänge mit float4“?

Die meisten CoddyKit-Lektionen dauern etwa 5–10 Minuten. Jede ist kompakt und interaktiv, sodass du stetig Fortschritte machst und genau dort weitermachst, wo du aufgehört hast – im Web und in der App.

Kann ich in dieser CUDA Academy-Lektion Code schreiben und ausführen?

Ja. Jede CUDA Academy-Lektion enthält einen integrierten Code-Editor, sodass du echten Code direkt in deinem Browser schreibst und ausführst und sofort KI-Feedback erhältst — ohne lokale Einrichtung erforderlich.

Alle Lektionen in diesem Kurs

  1. Instruction-Level Parallelism
  2. Schleifen mit #pragma unroll entrollen
  3. Vektorisierte Ladevorgänge mit float4
  4. Registerdruck und Spills
← Zurück zu CUDA Academy