0Pricing
CUDA Academy · Lektion

Thrust: Reduce, Scan und Sort

Hochrangige Primitive mit einem Aufruf

Thrust: Reduce, Scan und Sort ist eine kostenlose CUDA Academy-Lektion auf CoddyKit. Dies ist Lektion 3 von 4. Du kannst die komplette Lektion unten kostenlos lesen – dann übst du sie direkt im Browser mit einem integrierten Code-Editor und einem KI-Tutor rund um die Uhr. Sie ist Teil des CUDA Academy-Lernpfads, und dein Fortschritt wird über Web und CoddyKit-App synchronisiert. Der CUDA Academy-Kurs umfasst insgesamt 4 Lektionen.

Teile dieser Lektion wurden noch nicht übersetzt und werden auf Englisch angezeigt.

Hard Algorithms, One Line

Reductions, scans, and sorts are tricky to write fast by hand. Thrust gives you tuned versions through a single function call. 🎁

Reduce Collapses to One Value

thrust::reduce combines every element into a single result, like summing an array, all in parallel under the hood.

int total = thrust::reduce(d.begin(), d.end());

Custom Reduction Operators

Reduce defaults to addition, but you can pass an init value and a binary op to compute a product, max, or anything associative.

int m = thrust::reduce(d.begin(), d.end(),
  0, thrust::maximum<int>());

Scan Keeps the Running Total

A scan, or prefix sum, outputs the running total at each position. It is the backbone of compaction, sorting, and stream allocation.

Inclusive vs Exclusive

inclusive_scan includes the current element in its sum; exclusive_scan does not. Picking the right one avoids an off-by-one bug.

thrust::inclusive_scan(d.begin(), d.end(),
  out.begin());

Scan Is Not Obvious to Parallelize

A prefix sum looks sequential, yet Thrust runs it in parallel with a clever tree algorithm you never have to write yourself.

Sort in Place

thrust::sort orders a device_vector in place using a fast GPU radix or merge sort, far quicker than a CPU sort on big data.

thrust::sort(d.begin(), d.end());

Sort by Key

sort_by_key sorts one array and reorders a second values array to match, perfect for keeping records aligned with their keys.

thrust::sort_by_key(keys.begin(),
  keys.end(), values.begin());

Compose Primitives

Real pipelines chain these: transform then reduce, or sort then scan. Each step is one tuned call, so you focus on the logic.

Fused transform_reduce

transform_reduce maps and sums in one pass, computing things like a dot product or sum of squares without a temporary array.

float ss = thrust::transform_reduce(
  d.begin(), d.end(), sq, 0.0f, thrust::plus<float>());

Let the Library Win

These primitives are heavily optimized by NVIDIA. Reaching for them first usually beats a custom kernel and saves hours of work.

Quick Check

Recall what a prefix sum produces.

Recap

You collapsed data with reduce, built running totals with scan, ordered arrays with sort, and fused steps with transform_reduce. 🏁

Häufig gestellte Fragen

Ist die Lektion „Thrust: Reduce, Scan und Sort“ kostenlos?

Ja — der vollständige Text von „Thrust: Reduce, Scan und Sort“ ist hier im Web kostenlos zu lesen. Um sie interaktiv zu üben (integrierter Code-Editor und 24/7 KI-Tutor) und den Rest des CUDA Academy-Kurses freizuschalten, upgrade auf CoddyKit PRO. Der CUDA Academy-Kurs umfasst insgesamt 4 Lektionen.

Was lerne ich in „Thrust: Reduce, Scan und Sort“?

Hochrangige Primitive mit einem Aufruf Du übst CUDA Academy mit praktischem Code, den du direkt im Browser ausführst, und ein 24/7 KI-Tutor beantwortet deine Fragen während du die Lektion bearbeitest.

Brauche ich Erfahrung, um CUDA Academy zu starten?

Keine Vorkenntnisse erforderlich. CUDA Academy auf CoddyKit ist für Anfänger bis fortgeschrittene Lernende strukturiert, sodass du hier starten oder von Anfang an beginnen und in deinem eigenen Tempo voranschreiten kannst. Dies ist Lektion 3 von 4.

Wie lange dauert die Lektion „Thrust: Reduce, Scan und Sort“?

Die meisten CoddyKit-Lektionen dauern etwa 5–10 Minuten. Jede ist kompakt und interaktiv, sodass du stetig Fortschritte machst und genau dort weitermachst, wo du aufgehört hast – im Web und in der App.

Kann ich in dieser CUDA Academy-Lektion Code schreiben und ausführen?

Ja. Jede CUDA Academy-Lektion enthält einen integrierten Code-Editor, sodass du echten Code direkt in deinem Browser schreibst und ausführst und sofort KI-Feedback erhältst — ohne lokale Einrichtung erforderlich.

Alle Lektionen in diesem Kurs

  1. cuBLAS-GEMM richtig einsetzen
  2. Thrust-Vektoren und -Transformationen
  3. Thrust: Reduce, Scan und Sort
  4. cuDNN für Deep Learning
← Zurück zu CUDA Academy