Self-attention paso a paso
Consultas, claves y valores
Self-attention paso a paso es una lección gratuita de NLP Academy en CoddyKit. Esta es la lección 2 de 4. Puedes leer la lección completa abajo gratuitamente — luego la practicas en el navegador con un editor de código integrado y un tutor de IA 24/7. Forma parte de la ruta de aprendizaje de NLP Academy, y tu progreso se sincroniza en la web y la app de CoddyKit. El curso de NLP Academy incluye 4 lecciones en total.
Partes de esta lección aún no han sido traducidas y se muestran en inglés.
What Is Self-Attention?
In self-attention, every word looks at every other word in the same sentence to build a richer, context-aware version of itself. 🔍
Three Roles per Word
Each word is projected into three vectors: a query, a key, and a value. These three roles drive the whole self-attention computation.
The Query
The query represents what a word is looking for. Think of it as the question this word asks about the rest of the sentence.
Keys and Values
Each word also offers a key that advertises what it contains, and a value holding the actual information to pass along if chosen.
Scoring With Dot Products
Compare a query to every key using a dot product. A larger product means the query and key align, so that word deserves more focus.
scores = query @ keys.TScale the Scores
Divide scores by the square root of the key dimension. This scaling keeps numbers stable so softmax does not become too sharp.
scores = scores / (d_k ** 0.5)Softmax for Weights
Run the scaled scores through softmax to get attention weights that are positive and sum to 1 across all words in the sentence.
weights = softmax(scores)Blend the Values
Multiply each value by its weight and add them up. The result is a new vector that blends information from the most relevant words.
output = weights @ valuesThe Full Formula
All steps combine into one tidy expression: scaled dot-product attention over queries, keys, and values, as introduced in the Transformer paper.
attn = softmax(Q @ K.T / d_k**0.5) @ VWhy Self, Not Cross
It is called self-attention because the queries, keys, and values all come from the same sequence, letting words attend to their own neighbors.
Resolving Ambiguity
In "it was tired," self-attention links it to the right noun by weighting nearby words, giving every token clearer context.
Quick Check
Let us confirm the self-attention steps.
Recap
You walked through self-attention: turn words into queries, keys, and values, score, scale, softmax, then blend the values. That is the engine. ⚙️
Preguntas frecuentes
¿La lección «Self-attention paso a paso» es gratis?
Sí — el texto completo de «Self-attention paso a paso» es gratis para leer aquí en la web. Para practicarla de forma interactiva (editor de código integrado y tutor de IA 24/7) y desbloquear el resto del curso de NLP Academy, actualiza a CoddyKit PRO. El curso de NLP Academy incluye 4 lecciones en total.
¿Qué aprenderé en «Self-attention paso a paso»?
Consultas, claves y valores Practicas NLP Academy con código real que ejecutas directamente en el navegador, y un tutor de IA 24/7 responde tus preguntas mientras trabajas en la lección.
¿Necesito experiencia previa para empezar NLP Academy?
No se requiere experiencia previa. NLP Academy en CoddyKit está estructurado para principiantes hasta estudiantes avanzados, así que puedes empezar aquí o desde el inicio y avanzar a tu ritmo. Esta es la lección 2 de 4.
¿Cuánto tiempo toma la lección «Self-attention paso a paso»?
La mayoría de las lecciones de CoddyKit toman alrededor de 5–10 minutos. Cada una es compacta e interactiva, así que avanzas constantemente y retomas exactamente por donde dejaste en la web y la app.
¿Puedo escribir y ejecutar código en esta lección de NLP Academy?
Sí. Cada lección de NLP Academy incluye un editor de código integrado, así que escribes y ejecutas código real directamente en tu navegador y obtienes retroalimentación instantánea de IA — sin configuración local necesaria.
Todas las lecciones de este curso
- La idea de la atención
- Self-attention paso a paso
- Atención multi-cabeza y posiciones
- Dentro del bloque Transformer