การเรียกใช้ด้วยวงเล็บมุมสามชั้น
เรียกเคอร์เนลด้วย >>
การเรียกใช้ด้วยวงเล็บมุมสามชั้น เป็นบทเรียน CUDA Academy ฟรีบน CoddyKit นี่คือบทเรียนที่ 2 จากทั้งหมด 4 บทเรียน คุณสามารถอ่านบทเรียนทั้งหมดด้านล่างฟรี — จากนั้นลองปฏิบัติด้วยตัวคุณเองในเบราว์เซอร์พร้อมตัวแก้ไขโค้ดในตัวและติวเตอร์ AI ตลอด 24/7 บทเรียนนี้เป็นส่วนหนึ่งของเส้นทางการเรียน CUDA Academy และความก้าวหน้าของคุณจะซิงค์ข้ามเว็บและแอป CoddyKit คอร์ส CUDA Academy มีบทเรียนทั้งหมด 4 บทเรียน
บางส่วนของบทเรียนนี้ยังไม่ได้รับการแปล และแสดงเป็นภาษาอังกฤษ
Launching Is Different
You do not call a kernel like a normal function. You launch it, telling the GPU how many threads to spin up at the same time. 🚀
The <<< >>> Syntax
The launch uses CUDA's special triple angle brackets. Inside them you give a launch configuration before the normal argument list.
myKernel<<<blocks, threads>>>(args);First Number: Blocks
The first value sets how many blocks launch. A block is a group of threads that run together and share fast on-chip memory.
myKernel<<<4, threads>>>(args); // 4 blocksSecond Number: Threads per Block
The second value sets threads per block. Total threads launched equals blocks times threads per block.
myKernel<<<4, 256>>>(args); // 1024 threadsArguments Come After
After the brackets, you pass the kernel's normal arguments in parentheses, just like any C++ call.
doubleIt<<<1, 256>>>(devicePtr);Launches Are Asynchronous
A launch is asynchronous: the CPU queues the work and keeps going immediately. The GPU runs the kernel in the background.
Configs Can Be Variables
The launch numbers do not have to be literals. You often compute the block count from your data size at runtime.
int blocks = (n + 255) / 256;
add<<<blocks, 256>>>(c, a, b, n);Using dim3 for 2D
For grids and blocks in 2D or 3D, you use a dim3 value. It bundles x, y, and z sizes into the launch.
dim3 threads(16, 16);
blur<<<grid, threads>>>(img);Picking Threads per Block
A safe default for threads per block is 128 or 256. Values must be multiples of 32 and stay at or below 1024.
kernel<<<blocks, 128>>>(data);A Bad Config Fails
If you ask for too many threads per block, the launch fails silently. You must check for errors right after launching.
kernel<<<1, 2048>>>(p); // too big, failsRead It Left to Right
Read a launch as: run this kernel across this many blocks of this many threads, with these arguments. Simple once it clicks.
vectorAdd<<<grid, block>>>(c, a, b, n);Quick Check
Time to check the launch configuration.
Recap: The Launch
You launch with kernel<<
คำถามที่พบบ่อย
บทเรียน “การเรียกใช้ด้วยวงเล็บมุมสามชั้น” ฟรีหรือไม่
ใช่ — ข้อความเต็มของ “การเรียกใช้ด้วยวงเล็บมุมสามชั้น” ฟรีให้อ่านที่นี่บนเว็บ เพื่อปฏิบัติแบบโต้ตอบ (ตัวแก้ไขโค้ดในตัวและติวเตอร์ AI ตลอด 24/7) และปลดล็อคส่วนที่เหลือของคอร์ส CUDA Academy ให้อัปเกรดเป็น CoddyKit PRO คอร์ส CUDA Academy มีบทเรียนทั้งหมด 4 บทเรียน
คุณจะเรียนรู้อะไรในบทเรียน “การเรียกใช้ด้วยวงเล็บมุมสามชั้น”
เรียกเคอร์เนลด้วย >> คุณปฏิบัติ CUDA Academy ด้วยโค้ดที่ใช้งานได้จริงที่คุณเรียกใช้โดยตรงในเบราว์เซอร์ และติวเตอร์ AI ตลอด 24/7 ตอบคำถามของคุณขณะที่คุณไปผ่านบทเรียน
คุณต้องมีประสบการณ์ก่อนที่จะเริ่มเรียน CUDA Academy หรือไม่
ไม่จำเป็นต้องมีประสบการณ์มาก่อน CUDA Academy บน CoddyKit ออกแบบมาสำหรับผู้เริ่มต้นไปจนถึงผู้เรียนขั้นสูง คุณสามารถเริ่มต้นที่นี่หรือเริ่มจากตัวแรกและเรียนด้วยความเร็วของคุณเอง นี่คือบทเรียนที่ 2 จากทั้งหมด 4 บทเรียน
บทเรียน “การเรียกใช้ด้วยวงเล็บมุมสามชั้น” ใช้เวลานานแค่ไหน
บทเรียน CoddyKit ส่วนใหญ่ใช้เวลาประมาณ 5–10 นาที แต่ละบทเรียนจึงสั้นและเป็นแบบโต้ตอบ คุณสามารถก้าวหน้าอย่างต่อเนื่องและกลับมาเรียนต่อจากตรงที่เพิ่งหยุดบนเว็บและแอปได้เลย
ฉันเขียนและรันโค้ดในบทเรียน CUDA Academy นี้ได้ไหม
ได้ บทเรียน CUDA Academy ทุกบทมีตัวแก้ไขโค้ดในตัว คุณจึงเขียนและรันโค้ดจริงได้เลยในเบราว์เซอร์ และได้รับข้อเสนอแนะจาก AI ในทันที — ไม่ต้องติดตั้งในเครื่องของคุณ
บทเรียนทั้งหมดในหลักสูตรนี้
- โครงสร้างของเคอร์เนล
- การเรียกใช้ด้วยวงเล็บมุมสามชั้น
- printf ภายในเคอร์เนล
- อธิบาย cudaDeviceSynchronize