Perhatian Diri: Query, Key, dan Value
Biarkan setiap token melihat token lainnya
Perhatian Diri: Query, Key, dan Value adalah pelajaran Deep Learning Academy gratis di CoddyKit. Ini adalah pelajaran 1 dari 4. Kamu bisa membaca pelajaran lengkapnya di bawah secara gratis — lalu praktikkan langsung di browser dengan editor kode bawaan dan tutor AI 24/7. Ini adalah bagian dari jalur belajar Deep Learning Academy, dan progresmu tersinkronisasi di web dan aplikasi CoddyKit. Kursus Deep Learning Academy mencakup 4 pelajaran total.
Bagian dari pelajaran ini belum diterjemahkan dan ditampilkan dalam bahasa Inggris.
Tokens That Talk
In a sequence, the meaning of one word depends on others. Self-attention lets every token look at every other token to gather the context it needs.
Three Roles per Token
Each token plays three roles: a query that asks, a key that answers, and a value that carries content. These come from the same word, used three ways.
The Query
A token's query describes what it is looking for. Think of it as the question this word is asking about the rest of the sentence.
The Key
Every token also exposes a key, a label advertising what it offers. A query is compared against all keys to find good matches.
The Value
Once a match is found, the value is the actual information that gets passed along. Keys decide how much, values decide what.
Make Q, K, V
You build queries, keys, and values by projecting the input through three learned linear layers. Same input, three different weight matrices.
q = self.W_q(x)
k = self.W_k(x)
v = self.W_v(x)Score by Similarity
To see how well a query matches a key, you take their dot product. A bigger score means the two tokens are more relevant to each other.
scores = q @ k.transpose(-2, -1)Scores to Weights
Raw scores become attention weights with softmax, so each query's weights are positive and sum to one across all keys.
weights = scores.softmax(dim=-1)Blend the Values
The output for each token is a weighted sum of all values, mixed by the attention weights. Relevant tokens contribute more.
out = weights @ vWhy It Beats RNNs
Self-attention connects any two tokens in one step, so distance does not matter. This parallel view is why transformers handle long context so well.
Learned, Not Fixed
The Q, K, V projections are trained by gradient descent. The network learns what to ask, what to advertise, and what to share, all from data.
Quick Check
Let's test how attention combines its pieces.
Recap
You learned that self-attention turns each token into a query, key, and value, scores query-key matches, and blends values by softmax weights. Nice work!
Pertanyaan yang Sering Diajukan
Apakah pelajaran “Perhatian Diri: Query, Key, dan Value” gratis?
Ya — teks lengkap “Perhatian Diri: Query, Key, dan Value” gratis dibaca di sini di web. Untuk praktiknya secara interaktif (editor kode bawaan dan tutor AI 24/7) dan buka sisa kursus Deep Learning Academy, upgrade ke CoddyKit PRO. Kursus Deep Learning Academy mencakup 4 pelajaran total.
Apa yang akan aku pelajari di “Perhatian Diri: Query, Key, dan Value”?
Biarkan setiap token melihat token lainnya Kamu berlatih Deep Learning Academy dengan kode praktik yang langsung kamu jalankan di browser, dan tutor AI 24/7 menjawab pertanyaanmu saat kamu mengerjakan pelajaran ini.
Apakah aku perlu pengalaman untuk memulai Deep Learning Academy?
Tidak diperlukan pengalaman sebelumnya. Deep Learning Academy di CoddyKit dirancang untuk pemula hingga pelajar tingkat lanjut, jadi kamu bisa memulai di sini atau dari awal dan belajar sesuai kecepatan kamu sendiri. Ini adalah pelajaran 1 dari 4.
Berapa lama pelajaran “Perhatian Diri: Query, Key, dan Value” memakan waktu?
Sebagian besar pelajaran CoddyKit memakan waktu sekitar 5–10 menit. Setiap pelajaran ringkas dan interaktif, jadi kamu membuat kemajuan stabil dan melanjutkan dari tempat kamu tinggalkan di web dan aplikasi.
Bisakah aku menulis dan menjalankan kode dalam pelajaran Deep Learning Academy ini?
Ya. Setiap pelajaran Deep Learning Academy menyertakan editor kode bawaan, jadi kamu menulis dan menjalankan kode nyata langsung di browser dan mendapatkan umpan balik AI instan — tidak diperlukan penyiapan lokal.
Semua pelajaran dalam kursus ini
- Perhatian Diri: Query, Key, dan Value
- Produk Titik Berskala dan Multi-Head
- Pengodean Posisi untuk Urutan
- Susun Blok Encoder Transformer