กับดักของสตรีมเริ่มต้น
ดูว่าสตรีม 0 ทำให้ทุกอย่างทำงานต่อเนื่องทีละรายการอย่างไร
กับดักของสตรีมเริ่มต้น เป็นบทเรียน CUDA Academy ฟรีบน CoddyKit นี่คือบทเรียนที่ 1 จากทั้งหมด 4 บทเรียน คุณสามารถอ่านบทเรียนทั้งหมดด้านล่างฟรี — จากนั้นลองปฏิบัติด้วยตัวคุณเองในเบราว์เซอร์พร้อมตัวแก้ไขโค้ดในตัวและติวเตอร์ AI ตลอด 24/7 บทเรียนนี้เป็นส่วนหนึ่งของเส้นทางการเรียน CUDA Academy และความก้าวหน้าของคุณจะซิงค์ข้ามเว็บและแอป CoddyKit คอร์ส CUDA Academy มีบทเรียนทั้งหมด 4 บทเรียน
บางส่วนของบทเรียนนี้ยังไม่ได้รับการแปล และแสดงเป็นภาษาอังกฤษ
What Is a Stream?
A stream is an ordered queue of GPU work. Operations in the same stream run one after another, in the order you issued them. ⏳
The Default Stream
If you never name a stream, every call lands in the default stream, also called stream 0. It is where all your work has secretly been running.
Why It Is a Trap
The default stream is synchronizing: it blocks until other streams finish, and they wait for it. Nothing overlaps, so your GPU sits idle between steps.
Everything Serializes
Issue a copy, a kernel, then another copy on stream 0 and they run strictly back to back. The GPU can never start step two before step one ends.
cudaMemcpy(d_a, h_a, n, cudaMemcpyHostToDevice);
kernel<<<grid, block>>>(d_a);
cudaMemcpy(h_a, d_a, n, cudaMemcpyDeviceToHost);Wasted Hardware
Modern GPUs have separate engines for compute and for copying. On the default stream those engines take turns instead of working together.
The Hidden Cost of Copies
Data transfers over PCIe are slow. When copies cannot overlap with compute, that transfer time is added directly onto your total runtime.
Implicit Synchronization
Many default-stream calls are blocking on the host too. cudaMemcpy returns only after the copy is done, stalling your CPU.
The Legacy Behavior
By default, work in any stream waits for stream 0, and stream 0 waits for them. This legacy rule quietly destroys concurrency you thought you had.
A Telltale Profile
In a profiler the default-stream trap shows up as a single busy lane with gaps, while copy and compute engines sit idle waiting for each other.
The Way Out
The fix is to issue independent work on separate streams. Then copies and kernels can run at the same time and fill those idle gaps.
Per-Thread Default Streams
Compiling with --default-stream per-thread gives each host thread its own default stream, easing the trap without rewriting every call.
nvcc --default-stream per-thread app.cu -o appQuick Check
Why does putting all work on the default stream hurt performance?
Recap
The default stream serializes your GPU work and blocks overlap. To go faster, you will move independent tasks onto their own streams next. 🚀
คำถามที่พบบ่อย
บทเรียน “กับดักของสตรีมเริ่มต้น” ฟรีหรือไม่
ใช่ — ข้อความเต็มของ “กับดักของสตรีมเริ่มต้น” ฟรีให้อ่านที่นี่บนเว็บ เพื่อปฏิบัติแบบโต้ตอบ (ตัวแก้ไขโค้ดในตัวและติวเตอร์ AI ตลอด 24/7) และปลดล็อคส่วนที่เหลือของคอร์ส CUDA Academy ให้อัปเกรดเป็น CoddyKit PRO คอร์ส CUDA Academy มีบทเรียนทั้งหมด 4 บทเรียน
คุณจะเรียนรู้อะไรในบทเรียน “กับดักของสตรีมเริ่มต้น”
ดูว่าสตรีม 0 ทำให้ทุกอย่างทำงานต่อเนื่องทีละรายการอย่างไร คุณปฏิบัติ CUDA Academy ด้วยโค้ดที่ใช้งานได้จริงที่คุณเรียกใช้โดยตรงในเบราว์เซอร์ และติวเตอร์ AI ตลอด 24/7 ตอบคำถามของคุณขณะที่คุณไปผ่านบทเรียน
คุณต้องมีประสบการณ์ก่อนที่จะเริ่มเรียน CUDA Academy หรือไม่
ไม่จำเป็นต้องมีประสบการณ์มาก่อน CUDA Academy บน CoddyKit ออกแบบมาสำหรับผู้เริ่มต้นไปจนถึงผู้เรียนขั้นสูง คุณสามารถเริ่มต้นที่นี่หรือเริ่มจากตัวแรกและเรียนด้วยความเร็วของคุณเอง นี่คือบทเรียนที่ 1 จากทั้งหมด 4 บทเรียน
บทเรียน “กับดักของสตรีมเริ่มต้น” ใช้เวลานานแค่ไหน
บทเรียน CoddyKit ส่วนใหญ่ใช้เวลาประมาณ 5–10 นาที แต่ละบทเรียนจึงสั้นและเป็นแบบโต้ตอบ คุณสามารถก้าวหน้าอย่างต่อเนื่องและกลับมาเรียนต่อจากตรงที่เพิ่งหยุดบนเว็บและแอปได้เลย
ฉันเขียนและรันโค้ดในบทเรียน CUDA Academy นี้ได้ไหม
ได้ บทเรียน CUDA Academy ทุกบทมีตัวแก้ไขโค้ดในตัว คุณจึงเขียนและรันโค้ดจริงได้เลยในเบราว์เซอร์ และได้รับข้อเสนอแนะจาก AI ในทันที — ไม่ต้องติดตั้งในเครื่องของคุณ
บทเรียนทั้งหมดในหลักสูตรนี้
- กับดักของสตรีมเริ่มต้น
- สร้างและใช้สตรีม
- อีเวนต์สำหรับจับเวลาและซิงค์
- ทำการคัดลอกและคำนวณซ้อนกัน