La idea de la atención
Permitir que el modelo se centre en lo importante
La idea de la atención es una lección gratuita de NLP Academy en CoddyKit. Esta es la lección 1 de 4. Puedes leer la lección completa abajo gratuitamente — luego la practicas en el navegador con un editor de código integrado y un tutor de IA 24/7. Forma parte de la ruta de aprendizaje de NLP Academy, y tu progreso se sincroniza en la web y la app de CoddyKit. El curso de NLP Academy incluye 4 lecciones en total.
Partes de esta lección aún no han sido traducidas y se muestran en inglés.
Why Attention?
When you read a sentence, you do not weigh every word equally. Attention gives a model that same power to focus on what matters. 🎯
The Bottleneck Problem
Older models squeezed a whole sentence into one fixed vector. That bottleneck lost detail, especially for long inputs that carry many ideas.
Look Back at Everything
Instead of one summary vector, attention lets the model look back at every input word whenever it needs to, picking what is relevant right now.
Attention as Weights
Attention assigns each word a weight between 0 and 1. Higher weight means more focus; the weights for one step always add up to 1.
A Weighted Average
The output is a weighted average of the input vectors. Words with big weights shape the result; words with tiny weights barely matter.
weights = [0.7, 0.2, 0.1]
output = sum(w * v for w, v in zip(weights, vectors))Translation Example
Translating "the cat sat" into French, the model can align each output word to the right source word instead of guessing from one blob.
Soft, Not Hard
Attention is soft: it spreads focus across all words by degree, rather than hard-picking just one. That makes it smooth and trainable.
Computing Relevance
To set the weights, the model scores how relevant each word is to the current step, then turns those scores into a probability spread.
Softmax Turns Scores Into Weights
Raw scores can be any number, so softmax squashes them into positive weights that sum to 1 and emphasize the largest score.
import numpy as np
def softmax(s):
e = np.exp(s - np.max(s))
return e / e.sum()Handling Long Inputs
Because it can reach any word directly, attention keeps long-range links intact. The first word can still influence the last with no decay.
Why It Changed NLP
Attention freed models from reading strictly in order. That single idea became the foundation of the Transformer and modern language models. 🚀
Quick Check
Let us check the core intuition behind attention.
Recap
You learned that attention weights every input word by relevance and blends them into a focused output. That focus is what powers modern NLP. ✨
Preguntas frecuentes
¿La lección «La idea de la atención» es gratis?
Sí — el texto completo de «La idea de la atención» es gratis para leer aquí en la web. Para practicarla de forma interactiva (editor de código integrado y tutor de IA 24/7) y desbloquear el resto del curso de NLP Academy, actualiza a CoddyKit PRO. El curso de NLP Academy incluye 4 lecciones en total.
¿Qué aprenderé en «La idea de la atención»?
Permitir que el modelo se centre en lo importante Practicas NLP Academy con código real que ejecutas directamente en el navegador, y un tutor de IA 24/7 responde tus preguntas mientras trabajas en la lección.
¿Necesito experiencia previa para empezar NLP Academy?
No se requiere experiencia previa. NLP Academy en CoddyKit está estructurado para principiantes hasta estudiantes avanzados, así que puedes empezar aquí o desde el inicio y avanzar a tu ritmo. Esta es la lección 1 de 4.
¿Cuánto tiempo toma la lección «La idea de la atención»?
La mayoría de las lecciones de CoddyKit toman alrededor de 5–10 minutos. Cada una es compacta e interactiva, así que avanzas constantemente y retomas exactamente por donde dejaste en la web y la app.
¿Puedo escribir y ejecutar código en esta lección de NLP Academy?
Sí. Cada lección de NLP Academy incluye un editor de código integrado, así que escribes y ejecutas código real directamente en tu navegador y obtienes retroalimentación instantánea de IA — sin configuración local necesaria.
Todas las lecciones de este curso
- La idea de la atención
- Self-attention paso a paso
- Atención multi-cabeza y posiciones
- Dentro del bloque Transformer