0Pricing
CUDA Academy · 课时

内核的组成

函数签名、返回类型与 void 规则

内核的组成 是 CoddyKit 上的免费 CUDA Academy 课时。 这是第 1 节课,共 4 节。 你可以在下方免费阅读本课时的完整内容 — 然后在浏览器中使用内置代码编辑器和全天候 AI 导师进行实践。 这是 CUDA Academy 学习路径的一部分,你的进度在网页和 CoddyKit 应用中同步。 CUDA Academy 课程共包含 4 节课。

本课时的部分内容尚未翻译,以英文显示。

What a Kernel Is

A kernel is a single C++ function that runs on the GPU, executed at once by thousands of threads. You write it once, the hardware fans it out. 🚀

The __global__ Marker

You mark a kernel with the __global__ qualifier. It tells nvcc this function is called from the CPU but actually runs on the device.

__global__ void myKernel() {
    // runs on the GPU
}

Kernels Return void

Every kernel must return void. There is no return value to hand back, so results travel out through pointers to device memory instead.

__global__ void k() { /* void only */ }

Why Not Return a Value?

Thousands of threads run your kernel together, so a single return value would make no sense. Each thread writes its own slot of an output array.

Parameters Are by Value

Kernel parameters are copied to every thread, so pass small things: ints, sizes, and pointers. Never pass big objects by value.

__global__ void add(float* out, float* a, int n) {}

Pass Pointers, Not Arrays

To give a kernel an array, you pass a device pointer. The kernel reads and writes the GPU memory that pointer addresses.

__global__ void scale(float* data, float k) {}

Each Thread Picks Its Work

The same code runs in every thread, so each one uses its index to decide which element to touch. That is how one function covers a whole array.

int i = threadIdx.x; // who am I?

A Tiny Complete Kernel

Here is a full kernel: it doubles one element per thread. Notice the __global__ marker, the void return, and the pointer parameter all together.

__global__ void doubleIt(float* x) {
    int i = threadIdx.x;
    x[i] = x[i] * 2.0f;
}

Host Code Stays Separate

Your normal CPU function, often main, is host code. It sets things up and then asks the GPU to run the kernel.

int main() {
    // host side: prepare and launch
}

No Recursion or I/O Surprises

A kernel runs on hardware with tight rules, so keep it simple: avoid deep recursion and heavy standard-library calls inside it.

Naming Your Kernels

Treat a kernel like any function: give it a clear, verb-based name like vectorAdd. Good names make launches far easier to read.

__global__ void vectorAdd(float* c, float* a, float* b);

Quick Check

Let us check the kernel signature rules.

Recap: Kernel Anatomy

A kernel is a __global__ void function run by many threads. Pass pointers and small values, and let each thread use its index. Nicely done! 🎉

常见问题解答

「内核的组成」课时是免费的吗?

是的 — 「内核的组成」的完整文本可在网页上免费阅读。要进行交互式练习(内置代码编辑器和全天候 AI 导师)并解锁 CUDA Academy 课程的其余内容,请升级到 CoddyKit PRO。 CUDA Academy 课程共包含 4 节课。

「内核的组成」这节课中我会学到什么?

函数签名、返回类型与 void 规则 你通过在浏览器中直接运行的动手代码来练习 CUDA Academy,全天候 AI 导师会在你学习这节课的过程中回答你的问题。

学习 CUDA Academy 需要有经验吗?

无需任何先前经验。CoddyKit 上的 CUDA Academy 课程适合初学者到高级学习者,你可以从这里开始或从头开始,按照自己的节奏学习。 这是第 1 节课,共 4 节。

「内核的组成」课时需要多长时间?

大多数 CoddyKit 课程大约需要 5–10 分钟。每节课都很精短且互动,所以你能稳步进步,并在网页和应用中从离开的地方继续。

我能在这节 CUDA Academy 课中编写并运行代码吗?

能。每节 CUDA Academy 课都包含内置代码编辑器,你可以在浏览器中直接编写并运行真实代码,并获得即时 AI 反馈 — 无需本地设置。

此课程中的所有课时

  1. 内核的组成
  2. 三重尖括号启动语法
  3. 在内核中使用 printf
  4. 详解 cudaDeviceSynchronize
← 返回 CUDA Academy