0Pricing
CUDA Academy · درس

دمج المرشحات في نواة واحدة

تقليل عمليات الإطلاق وحركة البيانات العامة

دمج المرشحات في نواة واحدة درس مجاني في CUDA Academy على CoddyKit. هذا هو الدرس 2 من أصل 4. يمكنك قراءة الدرس كاملاً أدناه مجاناً — ثم تمرن عليه مباشرة في المتصفح باستخدام محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7. هذا الدرس جزء من مسار التعلم في CUDA Academy، وتقدمك يتزامن عبر الويب وتطبيق CoddyKit. تتضمن دورة CUDA Academy 4 دروس في المجموع.

بعض أجزاء هذا الدرس لم تُترجم بعد وتظهر باللغة الإنجليزية.

Why Fuse At All

Running five kernels means five launches and five round-trips to global memory. Fusing filters into one kernel cuts both, often giving a big speedup. 🚀

Launch Overhead Adds Up

Every kernel launch costs a few microseconds. On a tiny image that fixed launch overhead can dwarf the real work, so fewer launches means more useful time.

Global Traffic Is the Enemy

Separate stages write a pixel to global memory then read it right back. Fusing keeps that value in a register, erasing the wasted round-trip entirely.

Load Once Per Thread

In a fused kernel each thread reads its pixel a single time. That one load then feeds every filter in sequence without touching global memory again.

float v = input[idx];

Chain Operations in Registers

Apply each filter to the value already in hand. The pixel flows through brightness, gamma, and contrast as plain math, all kept in fast registers.

v = v * brightness;
v = powf(v, gamma);
v = (v - 0.5f) * contrast + 0.5f;

Write Once At the End

After the whole chain runs, store the final result. One store replaces the many writes that separate kernels would have made.

output[idx] = v;

Pointwise Filters Fuse Cleanly

Filters that touch only one pixel are pointwise and fuse with zero fuss. Brightness, gamma, and color tweaks are the easiest wins to combine.

Neighborhood Filters Are Harder

A blur reads nearby pixels, so fusing it needs shared memory tiles, not just registers. Fuse pointwise stages freely and treat stencils with extra care.

Watch Register Pressure

A big fused kernel uses more registers per thread. Too many and occupancy drops or values spill, so fuse aggressively but keep an eye on the cost.

Verify After Fusing

Fusing reorders work, so always compare the fused output against the staged version. A quick diff confirms the math still matches before you celebrate.

One Kernel, Many Filters

The payoff is real: a single launch reads, transforms, and writes each pixel exactly once. That is the heart of a well-tuned fused image kernel.

Quick Check

You merged three pointwise filters into one kernel. What is the main performance win?

Recap

You learned to fuse filters: load once, chain pointwise math in registers, write once. It cuts launches and global traffic, but watch register pressure and verify the output. ✅

الأسئلة الشائعة

هل درس «دمج المرشحات في نواة واحدة» مجاني؟

نعم — نص درس «دمج المرشحات في نواة واحدة» كامل متاح مجاناً هنا على الويب. لتمرينه بشكل تفاعلي (محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7) وفتح باقي دورة CUDA Academy، انتقل إلى CoddyKit PRO. تتضمن دورة CUDA Academy 4 دروس في المجموع.

ماذا ستتعلم في «دمج المرشحات في نواة واحدة»؟

تقليل عمليات الإطلاق وحركة البيانات العامة تتمرن على CUDA Academy مع أكواد عملية تشغلها مباشرة في المتصفح، ومدرس ذكاء اصطناعي متاح 24/7 يجيب على أسئلتك أثناء عملك.

هل أحتاج إلى خبرة سابقة لأبدأ CUDA Academy؟

لا تُشترط خبرة سابقة. CUDA Academy على CoddyKit منظم للمبتدئين حتى المتقدمين، لذا يمكنك البدء من هنا أو من البداية والتقدم بسرعتك الخاصة. هذا هو الدرس 2 من أصل 4.

كم من الوقت يستغرق درس «دمج المرشحات في نواة واحدة»؟

معظم دروس CoddyKit تستغرق حوالي 5–10 دقائق. كل منها موجز وتفاعلي، لذا تحرز تقدماً مستمراً وتستأنف من حيث توقفت عبر الويب والتطبيق.

هل يمكنني كتابة وتشغيل أكواد في درس CUDA Academy هذا؟

نعم. كل درس في CUDA Academy يتضمن محرر أكواد مدمج، لذا تكتب وتشغل أكواداً حقيقية مباشرة في متصفحك وتحصل على تعليقات فورية من الذكاء الاصطناعي — بدون إعداد محلي.

جميع الدروس في هذه الدورة

  1. تصميم مسار المعالجة
  2. دمج المرشحات في نواة واحدة
  3. بث البلاطات للصور الكبيرة
  4. حلّل، حسّن، ثم أطلق
← العودة إلى CUDA Academy