Регистры и локальная память
Самое быстрое хранилище для каждого потока и его переполнение.
«Регистры и локальная память» — бесплатный урок CUDA Academy на CoddyKit. Это урок 1 из 4. Ты можешь прочитать весь урок бесплатно ниже — а потом практиковать его прямо в браузере с встроенным редактором кода и ИИ-репетитором 24/7. Это часть пути обучения CUDA Academy, и твой прогресс синхронизируется между веб-версией и приложением CoddyKit. Курс CUDA Academy содержит 4 уроков всего.
Части этого урока еще не переведены и отображаются на английском.
The Fastest Memory You Have
Every thread gets its own private registers, the quickest storage on the chip. They live right next to the math units, so reads cost almost nothing.
Plain Variables Become Registers
When you write a normal local variable in a kernel, the compiler usually keeps it in a register. No special syntax is needed at all. 🙂
__global__ void k() {
int x = 5; // x lives in a register
float y = x * 2.0f;
}Private to One Thread
A register value belongs to exactly one thread. Thread 7 cannot see thread 8's register, which is why per-thread work stays cleanly separated.
Registers Are a Shared Pool
Each streaming multiprocessor has a fixed register file. All resident threads split it, so using fewer registers per thread lets more threads run at once.
When You Run Out
If a kernel needs more registers than are available, the extra values get pushed elsewhere. This event is called a register spill.
Where Spills Go: Local Memory
Spilled values land in local memory, which despite the name is not on-chip. It actually lives in slow off-chip DRAM.
Local Is Still Per-Thread
Local memory is private to each thread just like registers. The difference is purely speed, since local memory is far slower to reach.
Big Local Arrays Spill
Declaring a large per-thread array often forces it into local memory, because there simply are not enough registers to hold every element.
__global__ void k() {
float buf[64]; // likely lives in local memory
}Spills Hurt Performance
Every spill turns a free register access into a slow DRAM trip. Cutting register pressure is one of the easiest optimization wins you can make.
Seeing Your Register Count
You can ask the compiler to report registers per thread. Pass --ptxas-options=-v to nvcc and it prints usage for each kernel.
nvcc --ptxas-options=-v kernel.cuCapping Registers on Purpose
The __launch_bounds__ qualifier hints the compiler to limit registers, trading a little per-thread speed for more threads running together.
__global__ void __launch_bounds__(256)
myKernel() { /* ... */ }Quick Check
You declared a big per-thread array and performance dropped. What likely happened?
Recap: Fast and Private
You learned that registers are the fastest per-thread storage, that the pool is limited, and that overflow spills into slow local memory. Keep register pressure low. 🎯
Часто задаваемые вопросы
Урок «Регистры и локальная память» бесплатный?
Да — полный текст урока «Регистры и локальная память» бесплатно доступен здесь в веб-версии. Чтобы практиковать его интерактивно (встроенный редактор кода и ИИ-репетитор 24/7) и разблокировать остальной курс CUDA Academy, подпишись на CoddyKit PRO. Курс CUDA Academy содержит 4 уроков всего.
Чему я научусь в уроке «Регистры и локальная память»?
Самое быстрое хранилище для каждого потока и его переполнение. Ты практикуешь CUDA Academy с помощью реального кода, который запускаешь прямо в браузере, и ИИ-репетитор 24/7 отвечает на твои вопросы во время урока.
Нужен ли мне опыт, чтобы начать CUDA Academy?
Предыдущий опыт не требуется. CUDA Academy на CoddyKit структурирован для всех уровней — от новичков до продвинутых, поэтому ты можешь начать отсюда или с самого начала и учиться в своем темпе. Это урок 1 из 4.
Сколько времени занимает урок «Регистры и локальная память»?
Большинство уроков CoddyKit занимают около 5–10 минут. Каждый из них компактный и интерактивный, поэтому ты постоянно делаешь прогресс и продолжаешь с того же места в веб-версии и приложении.
Можно ли писать и запускать код в этом уроке CUDA Academy?
Да. Каждый урок CUDA Academy включает встроенный редактор кода, поэтому ты пишешь и запускаешь реальный код прямо в браузере и получаешь моментальную обратную связь от AI — локальная установка не требуется.
Все уроки этого курса
- Регистры и локальная память
- Компромиссы глобальной памяти
- Константная память и её кэш
- Мысленная модель иерархии