Was eine Speichertransaktion ist
Cache-Zeilen und 32-/128-Byte-Segmente.
Was eine Speichertransaktion ist ist eine kostenlose CUDA Academy-Lektion auf CoddyKit. Dies ist Lektion 1 von 4. Du kannst die komplette Lektion unten kostenlos lesen – dann übst du sie direkt im Browser mit einem integrierten Code-Editor und einem KI-Tutor rund um die Uhr. Sie ist Teil des CUDA Academy-Lernpfads, und dein Fortschritt wird über Web und CoddyKit-App synchronisiert. Der CUDA Academy-Kurs umfasst insgesamt 4 Lektionen.
Teile dieser Lektion wurden noch nicht übersetzt und werden auf Englisch angezeigt.
Memory Comes in Chunks
The GPU never fetches a single byte by itself. It moves memory in fixed-size blocks called a memory transaction, even if you only asked for one value.
The Warp Asks Together
A warp of 32 threads issues its loads at the same time. The hardware gathers all 32 requests and tries to serve them with as few transactions as possible.
Cache Lines and Segments
Global memory is carved into aligned segments, commonly 32 or 128 bytes wide. Each transaction pulls in one whole segment, never a partial slice.
Why 128 Bytes Matters
A 128-byte line holds exactly 32 four-byte floats. That is one value for every thread in a warp, so a tidy warp can be served by a single transaction.
// 32 threads x 4-byte float = 128 bytes = one lineAlignment Is Key
Segments start at fixed, aligned addresses. If your data begins mid-segment, a single warp can spill across two lines and force an extra transaction.
Paying for Bytes You Skip
If a warp touches only a few bytes of a 128-byte line, the GPU still moves all 128. The unused bytes are wasted bandwidth you paid for anyway.
Counting Transactions
Performance often comes down to one number: how many transactions a warp needs. Fewer transactions means you read less and finish faster.
The Best Case
When 32 threads read 32 neighboring floats from an aligned address, the GPU answers with one clean 128-byte line. That is the ideal single transaction.
float v = data[blockIdx.x * blockDim.x + threadIdx.x];The Worst Case
If each thread grabs a value far from its neighbors, the warp may trigger up to 32 separate transactions, dragging in 32 full lines for tiny use.
Bandwidth Is the Limit
Most kernels are bound by how fast they move memory, not by math. Shrinking transaction count directly raises your effective bandwidth. 🚀
Set the Stage
Now you know the unit of cost: the segment-sized transaction. Every coalescing trick ahead is really about packing warps into as few of these as you can.
Quick Check
Test your grasp of transaction size.
Recap
You learned that the GPU moves memory in fixed segments, and a warp is served by as few transactions as it can manage. Fewer transactions, faster kernel. 🎉
Häufig gestellte Fragen
Ist die Lektion „Was eine Speichertransaktion ist“ kostenlos?
Ja — der vollständige Text von „Was eine Speichertransaktion ist“ ist hier im Web kostenlos zu lesen. Um sie interaktiv zu üben (integrierter Code-Editor und 24/7 KI-Tutor) und den Rest des CUDA Academy-Kurses freizuschalten, upgrade auf CoddyKit PRO. Der CUDA Academy-Kurs umfasst insgesamt 4 Lektionen.
Was lerne ich in „Was eine Speichertransaktion ist“?
Cache-Zeilen und 32-/128-Byte-Segmente. Du übst CUDA Academy mit praktischem Code, den du direkt im Browser ausführst, und ein 24/7 KI-Tutor beantwortet deine Fragen während du die Lektion bearbeitest.
Brauche ich Erfahrung, um CUDA Academy zu starten?
Keine Vorkenntnisse erforderlich. CUDA Academy auf CoddyKit ist für Anfänger bis fortgeschrittene Lernende strukturiert, sodass du hier starten oder von Anfang an beginnen und in deinem eigenen Tempo voranschreiten kannst. Dies ist Lektion 1 von 4.
Wie lange dauert die Lektion „Was eine Speichertransaktion ist“?
Die meisten CoddyKit-Lektionen dauern etwa 5–10 Minuten. Jede ist kompakt und interaktiv, sodass du stetig Fortschritte machst und genau dort weitermachst, wo du aufgehört hast – im Web und in der App.
Kann ich in dieser CUDA Academy-Lektion Code schreiben und ausführen?
Ja. Jede CUDA Academy-Lektion enthält einen integrierten Code-Editor, sodass du echten Code direkt in deinem Browser schreibst und ausführst und sofort KI-Feedback erhältst — ohne lokale Einrichtung erforderlich.
Alle Lektionen in diesem Kurs
- Was eine Speichertransaktion ist
- Koaleszierte vs. gestridete Lesezugriffe
- Structure of Arrays vs. Array of Structs
- Die effektive Bandbreite messen