0Pricing
NLP Academy · Lektion

Im Inneren des Transformer-Blocks

Wie die gesamte Architektur zusammenspielt

Im Inneren des Transformer-Blocks ist eine kostenlose NLP Academy-Lektion auf CoddyKit. Dies ist Lektion 4 von 4. Du kannst die komplette Lektion unten kostenlos lesen – dann übst du sie direkt im Browser mit einem integrierten Code-Editor und einem KI-Tutor rund um die Uhr. Sie ist Teil des NLP Academy-Lernpfads, und dein Fortschritt wird über Web und CoddyKit-App synchronisiert. Der NLP Academy-Kurs umfasst insgesamt 4 Lektionen.

Teile dieser Lektion wurden noch nicht übersetzt und werden auf Englisch angezeigt.

Stacking Blocks

A Transformer is built by stacking the same block many times. Each block refines the representation a little more. 🧱

Two Main Sublayers

Every block has two parts: a multi-head attention sublayer followed by a small feed-forward network applied to each position.

The Feed-Forward Network

The feed-forward sublayer is two linear layers with a nonlinearity between them, processing each token independently for extra power.

h = relu(x @ W1 + b1) @ W2 + b2

Residual Connections

Each sublayer wraps its input in a residual connection: add the input back to the output so gradients flow and nothing is lost.

out = x + sublayer(x)

Layer Normalization

After adding the residual, layer normalization rescales the values. This keeps activations stable and speeds up training.

out = layer_norm(x + sublayer(x))

Encoder Blocks

An encoder stack reads the input and builds rich context vectors. It uses self-attention so every token sees all the others.

Decoder Blocks

A decoder generates output one token at a time. It adds masked attention plus cross-attention back to the encoder.

Masked Attention

During generation, masking hides future tokens so the decoder cannot peek ahead and must predict the next word honestly.

Cross-Attention

Cross-attention lets the decoder query the encoder's outputs, connecting what it is writing to what it originally read.

Depth Brings Power

Stacking many blocks lets early layers catch simple patterns and later layers build abstract meaning, much like deep vision networks.

Encoder, Decoder, or Both

BERT uses only the encoder, GPT uses only the decoder, and translation models use both. The block is the shared building unit. 🚀

Quick Check

Let us check the block structure.

Recap

You assembled the Transformer block: attention plus feed-forward, wrapped in residuals and layer norm, stacked into encoders and decoders. ✨

Häufig gestellte Fragen

Ist die Lektion „Im Inneren des Transformer-Blocks“ kostenlos?

Ja — der vollständige Text von „Im Inneren des Transformer-Blocks“ ist hier im Web kostenlos zu lesen. Um sie interaktiv zu üben (integrierter Code-Editor und 24/7 KI-Tutor) und den Rest des NLP Academy-Kurses freizuschalten, upgrade auf CoddyKit PRO. Der NLP Academy-Kurs umfasst insgesamt 4 Lektionen.

Was lerne ich in „Im Inneren des Transformer-Blocks“?

Wie die gesamte Architektur zusammenspielt Du übst NLP Academy mit praktischem Code, den du direkt im Browser ausführst, und ein 24/7 KI-Tutor beantwortet deine Fragen während du die Lektion bearbeitest.

Brauche ich Erfahrung, um NLP Academy zu starten?

Keine Vorkenntnisse erforderlich. NLP Academy auf CoddyKit ist für Anfänger bis fortgeschrittene Lernende strukturiert, sodass du hier starten oder von Anfang an beginnen und in deinem eigenen Tempo voranschreiten kannst. Dies ist Lektion 4 von 4.

Wie lange dauert die Lektion „Im Inneren des Transformer-Blocks“?

Die meisten CoddyKit-Lektionen dauern etwa 5–10 Minuten. Jede ist kompakt und interaktiv, sodass du stetig Fortschritte machst und genau dort weitermachst, wo du aufgehört hast – im Web und in der App.

Kann ich in dieser NLP Academy-Lektion Code schreiben und ausführen?

Ja. Jede NLP Academy-Lektion enthält einen integrierten Code-Editor, sodass du echten Code direkt in deinem Browser schreibst und ausführst und sofort KI-Feedback erhältst — ohne lokale Einrichtung erforderlich.

Alle Lektionen in diesem Kurs

  1. Die Idee der Attention
  2. Self-Attention Schritt für Schritt
  3. Multi-Head-Attention und Positionen
  4. Im Inneren des Transformer-Blocks
← Zurück zu NLP Academy