การเข้ารหัสตำแหน่งเพื่อบอกลำดับ
แทรกตำแหน่งในลำดับลงในโทเค็น
การเข้ารหัสตำแหน่งเพื่อบอกลำดับ เป็นบทเรียน Deep Learning Academy ฟรีบน CoddyKit นี่คือบทเรียนที่ 3 จากทั้งหมด 4 บทเรียน คุณสามารถอ่านบทเรียนทั้งหมดด้านล่างฟรี — จากนั้นลองปฏิบัติด้วยตัวคุณเองในเบราว์เซอร์พร้อมตัวแก้ไขโค้ดในตัวและติวเตอร์ AI ตลอด 24/7 บทเรียนนี้เป็นส่วนหนึ่งของเส้นทางการเรียน Deep Learning Academy และความก้าวหน้าของคุณจะซิงค์ข้ามเว็บและแอป CoddyKit คอร์ส Deep Learning Academy มีบทเรียนทั้งหมด 4 บทเรียน
บางส่วนของบทเรียนนี้ยังไม่ได้รับการแปล และแสดงเป็นภาษาอังกฤษ
Attention Ignores Order
Self-attention treats a sentence like a bag of tokens. Shuffle the words and the math barely changes, so the model has no built-in sense of order.
Order Carries Meaning
But order matters: "dog bites man" is not "man bites dog". We must hand the model some signal about each token's position.
Add, Do Not Append
The trick is to add a position vector directly to each token's embedding. Same shape, so attention sees content and place fused together.
x = token_emb + pos_embSinusoidal Encoding
The original transformer used fixed sine and cosine waves of different frequencies, one pattern per dimension, to mark each position.
Why Sine Waves
Mixing frequencies gives every position a unique fingerprint, and the smooth waves let the model generalize to lengths it never saw in training.
The Formula
Even dimensions use sine, odd dimensions use cosine, with the wavelength growing across dimensions. That is the whole encoding recipe.
pe[:, 0::2] = torch.sin(pos / div)
pe[:, 1::2] = torch.cos(pos / div)Relative Distance
A neat bonus: with sinusoids, the offset between two positions is easy to express, helping attention reason about relative distance.
Learned Positions
Modern models often skip sinusoids and use a learned embedding table, one trainable vector per position, just like word embeddings.
self.pos_emb = nn.Embedding(max_len, d_model)Fixed vs Learned
Fixed sinusoids extrapolate to new lengths for free; learned tables fit your data but are capped at the maximum length you trained on.
Rotary Encoding
Newer transformers favor rotary position encoding, which rotates query and key vectors by an angle tied to position, baking order into attention itself.
Where It Goes
Whatever scheme you pick, positional info is injected once at the input, before the first attention layer ever runs.
Quick Check
Let's check why positional encoding exists.
Recap
You learned that attention ignores order, so we add positional encodings, whether sinusoidal, learned, or rotary, to tell tokens where they sit. Well done!
คำถามที่พบบ่อย
บทเรียน “การเข้ารหัสตำแหน่งเพื่อบอกลำดับ” ฟรีหรือไม่
ใช่ — ข้อความเต็มของ “การเข้ารหัสตำแหน่งเพื่อบอกลำดับ” ฟรีให้อ่านที่นี่บนเว็บ เพื่อปฏิบัติแบบโต้ตอบ (ตัวแก้ไขโค้ดในตัวและติวเตอร์ AI ตลอด 24/7) และปลดล็อคส่วนที่เหลือของคอร์ส Deep Learning Academy ให้อัปเกรดเป็น CoddyKit PRO คอร์ส Deep Learning Academy มีบทเรียนทั้งหมด 4 บทเรียน
คุณจะเรียนรู้อะไรในบทเรียน “การเข้ารหัสตำแหน่งเพื่อบอกลำดับ”
แทรกตำแหน่งในลำดับลงในโทเค็น คุณปฏิบัติ Deep Learning Academy ด้วยโค้ดที่ใช้งานได้จริงที่คุณเรียกใช้โดยตรงในเบราว์เซอร์ และติวเตอร์ AI ตลอด 24/7 ตอบคำถามของคุณขณะที่คุณไปผ่านบทเรียน
คุณต้องมีประสบการณ์ก่อนที่จะเริ่มเรียน Deep Learning Academy หรือไม่
ไม่จำเป็นต้องมีประสบการณ์มาก่อน Deep Learning Academy บน CoddyKit ออกแบบมาสำหรับผู้เริ่มต้นไปจนถึงผู้เรียนขั้นสูง คุณสามารถเริ่มต้นที่นี่หรือเริ่มจากตัวแรกและเรียนด้วยความเร็วของคุณเอง นี่คือบทเรียนที่ 3 จากทั้งหมด 4 บทเรียน
บทเรียน “การเข้ารหัสตำแหน่งเพื่อบอกลำดับ” ใช้เวลานานแค่ไหน
บทเรียน CoddyKit ส่วนใหญ่ใช้เวลาประมาณ 5–10 นาที แต่ละบทเรียนจึงสั้นและเป็นแบบโต้ตอบ คุณสามารถก้าวหน้าอย่างต่อเนื่องและกลับมาเรียนต่อจากตรงที่เพิ่งหยุดบนเว็บและแอปได้เลย
ฉันเขียนและรันโค้ดในบทเรียน Deep Learning Academy นี้ได้ไหม
ได้ บทเรียน Deep Learning Academy ทุกบทมีตัวแก้ไขโค้ดในตัว คุณจึงเขียนและรันโค้ดจริงได้เลยในเบราว์เซอร์ และได้รับข้อเสนอแนะจาก AI ในทันที — ไม่ต้องติดตั้งในเครื่องของคุณ
บทเรียนทั้งหมดในหลักสูตรนี้
- ความสนใจในตัวเอง: คิวรี คีย์ และแวลู
- ผลคูณจุดแบบปรับสเกลและหลายหัว
- การเข้ารหัสตำแหน่งเพื่อบอกลำดับ
- วางบล็อกตัวเข้ารหัส Transformer ซ้อนกัน