Transformer Bloğunun İçinde
Tüm mimarinin nasıl bir araya geldiği
Transformer Bloğunun İçinde, CoddyKit'te ücretsiz bir NLP Academy dersidir. Bu, 4 dersinin 4. dersidir. Aşağıdan dersin tamamını ücretsiz okuyabilir, sonra tarayıcıda yerleşik kod editörü ve 7/24 yapay zeka koçu ile uygulamalı olarak pratik yapabilirsin. Bu, NLP Academy öğrenme yolunun bir parçasıdır ve ilerlemeniz web ve CoddyKit uygulaması arasında senkronize olur. NLP Academy kursu toplamda 4 dersten oluşur.
Bu dersin bazı bölümleri henüz çevrilmemiş olup İngilizce olarak gösterilmektedir.
Stacking Blocks
A Transformer is built by stacking the same block many times. Each block refines the representation a little more. 🧱
Two Main Sublayers
Every block has two parts: a multi-head attention sublayer followed by a small feed-forward network applied to each position.
The Feed-Forward Network
The feed-forward sublayer is two linear layers with a nonlinearity between them, processing each token independently for extra power.
h = relu(x @ W1 + b1) @ W2 + b2Residual Connections
Each sublayer wraps its input in a residual connection: add the input back to the output so gradients flow and nothing is lost.
out = x + sublayer(x)Layer Normalization
After adding the residual, layer normalization rescales the values. This keeps activations stable and speeds up training.
out = layer_norm(x + sublayer(x))Encoder Blocks
An encoder stack reads the input and builds rich context vectors. It uses self-attention so every token sees all the others.
Decoder Blocks
A decoder generates output one token at a time. It adds masked attention plus cross-attention back to the encoder.
Masked Attention
During generation, masking hides future tokens so the decoder cannot peek ahead and must predict the next word honestly.
Cross-Attention
Cross-attention lets the decoder query the encoder's outputs, connecting what it is writing to what it originally read.
Depth Brings Power
Stacking many blocks lets early layers catch simple patterns and later layers build abstract meaning, much like deep vision networks.
Encoder, Decoder, or Both
BERT uses only the encoder, GPT uses only the decoder, and translation models use both. The block is the shared building unit. 🚀
Quick Check
Let us check the block structure.
Recap
You assembled the Transformer block: attention plus feed-forward, wrapped in residuals and layer norm, stacked into encoders and decoders. ✨
Sıkça Sorulan Sorular
“Transformer Bloğunun İçinde” dersi ücretsiz mi?
Evet — “Transformer Bloğunun İçinde” dersin tüm metni burada web'de ücretsiz olarak okunabilir. Etkileşimli olarak pratik yapmak (yerleşik kod editörü ve 7/24 yapay zeka koçu) ve NLP Academy kursunun geri kalanını açmak için CoddyKit PRO'ya yükselt. NLP Academy kursu toplamda 4 dersten oluşur.
“Transformer Bloğunun İçinde” dersinde ne öğreneceğim?
Tüm mimarinin nasıl bir araya geldiği NLP Academy ile uygulamalı kodu tarayıcıda doğrudan çalıştırarak pratik yaparsın ve 7/24 yapay zeka koçu dersi çalışırken sorularını yanıtlar.
NLP Academy öğrenmeye başlamak için deneyim gerekli mi?
Önceden deneyim gerekmez. CoddyKit'te NLP Academy, başlangıçtan ileri seviyeye kadar yapılandırıldığı için buradan başlayabilir veya başından başlayıp kendi hızında ilerleme yapabilirsin. Bu, 4 dersinin 4. dersidir.
“Transformer Bloğunun İçinde” dersi ne kadar sürer?
Çoğu CoddyKit dersi yaklaşık 5–10 dakika sürer. Her biri kısa ve etkileşimli olduğu için sabit ilerleme yaparsın ve web ile uygulama arasında tam olarak bıraktığın yerden devam edebilirsin.
Bu NLP Academy dersinde kod yazıp çalıştırabilir miyim?
Evet. Her NLP Academy dersi yerleşik bir kod editörü içerir, bu sayede tarayıcıda gerçek kod yazıp çalıştırabilir ve anlık yapay zeka geri bildirimi alırsın — yerel kurulum gerekli değildir.
Bu kursun tüm dersleri
- Dikkat Mekanizması Fikri
- Öz-Dikkat Adım Adım
- Çok Başlı Dikkat ve Konumlar
- Transformer Bloğunun İçinde