อธิบาย Adam และ AdamW
อัตราการปรับตัวพร้อมการเสื่อมน้ำหนักแบบแยกส่วน
อธิบาย Adam และ AdamW เป็นบทเรียน Deep Learning Academy ฟรีบน CoddyKit นี่คือบทเรียนที่ 2 จากทั้งหมด 4 บทเรียน คุณสามารถอ่านบทเรียนทั้งหมดด้านล่างฟรี — จากนั้นลองปฏิบัติด้วยตัวคุณเองในเบราว์เซอร์พร้อมตัวแก้ไขโค้ดในตัวและติวเตอร์ AI ตลอด 24/7 บทเรียนนี้เป็นส่วนหนึ่งของเส้นทางการเรียน Deep Learning Academy และความก้าวหน้าของคุณจะซิงค์ข้ามเว็บและแอป CoddyKit คอร์ส Deep Learning Academy มีบทเรียนทั้งหมด 4 บทเรียน
บางส่วนของบทเรียนนี้ยังไม่ได้รับการแปล และแสดงเป็นภาษาอังกฤษ
One Rate Per Weight
SGD uses a single learning rate for every parameter. Adam adapts the step size for each weight on its own, based on that weight's gradient history.
Two Moving Averages
Adam tracks two running averages: the mean of gradients and the mean of their squares. Together they form the first and second moments.
First Moment Is Momentum
The first moment is basically momentum, the smoothed average of recent gradients. It decides the overall direction each weight should move.
Second Moment Scales Steps
The second moment estimates each gradient's size. Adam divides by its square root, so noisy weights take smaller steps and quiet ones take larger.
The Betas
Two decay rates, the betas, control those averages, typically 0.9 and 0.999. They balance how much recent versus older gradients matter.
Bias Correction
The averages start at zero, so early steps look too small. Adam applies a bias correction to fix this so updates are sane from step one.
Use It in PyTorch
One line gives you Adam. The default learning rate of 0.001 works well across a huge range of models, which is why it is so popular.
opt = torch.optim.Adam(model.parameters(), lr=1e-3)Adam's Weight Decay Flaw
Classic Adam mixes weight decay into the gradient, where the adaptive scaling distorts it. The regularization ends up weaker than you intended.
AdamW Fixes It
AdamW decouples weight decay from the gradient step and applies it directly to the weights. The decay now works as a clean, predictable shrink.
opt = torch.optim.AdamW(model.parameters(), lr=1e-3, weight_decay=0.01)The Modern Default
For transformers and most large models, AdamW is the standard choice. Reach for it first whenever you want fast, reliable convergence.
Adaptive, With Caveats
Adam often trains faster than SGD, yet plain SGD with momentum can generalize better on vision tasks. Try both when accuracy really counts.
Quick Check
Pin down the Adam to AdamW difference.
Recap
Adam adapts a learning rate per weight using gradient mean and variance, while AdamW fixes its weight decay. AdamW is today's go-to optimizer. ⚙️
คำถามที่พบบ่อย
บทเรียน “อธิบาย Adam และ AdamW” ฟรีหรือไม่
ใช่ — ข้อความเต็มของ “อธิบาย Adam และ AdamW” ฟรีให้อ่านที่นี่บนเว็บ เพื่อปฏิบัติแบบโต้ตอบ (ตัวแก้ไขโค้ดในตัวและติวเตอร์ AI ตลอด 24/7) และปลดล็อคส่วนที่เหลือของคอร์ส Deep Learning Academy ให้อัปเกรดเป็น CoddyKit PRO คอร์ส Deep Learning Academy มีบทเรียนทั้งหมด 4 บทเรียน
คุณจะเรียนรู้อะไรในบทเรียน “อธิบาย Adam และ AdamW”
อัตราการปรับตัวพร้อมการเสื่อมน้ำหนักแบบแยกส่วน คุณปฏิบัติ Deep Learning Academy ด้วยโค้ดที่ใช้งานได้จริงที่คุณเรียกใช้โดยตรงในเบราว์เซอร์ และติวเตอร์ AI ตลอด 24/7 ตอบคำถามของคุณขณะที่คุณไปผ่านบทเรียน
คุณต้องมีประสบการณ์ก่อนที่จะเริ่มเรียน Deep Learning Academy หรือไม่
ไม่จำเป็นต้องมีประสบการณ์มาก่อน Deep Learning Academy บน CoddyKit ออกแบบมาสำหรับผู้เริ่มต้นไปจนถึงผู้เรียนขั้นสูง คุณสามารถเริ่มต้นที่นี่หรือเริ่มจากตัวแรกและเรียนด้วยความเร็วของคุณเอง นี่คือบทเรียนที่ 2 จากทั้งหมด 4 บทเรียน
บทเรียน “อธิบาย Adam และ AdamW” ใช้เวลานานแค่ไหน
บทเรียน CoddyKit ส่วนใหญ่ใช้เวลาประมาณ 5–10 นาที แต่ละบทเรียนจึงสั้นและเป็นแบบโต้ตอบ คุณสามารถก้าวหน้าอย่างต่อเนื่องและกลับมาเรียนต่อจากตรงที่เพิ่งหยุดบนเว็บและแอปได้เลย
ฉันเขียนและรันโค้ดในบทเรียน Deep Learning Academy นี้ได้ไหม
ได้ บทเรียน Deep Learning Academy ทุกบทมีตัวแก้ไขโค้ดในตัว คุณจึงเขียนและรันโค้ดจริงได้เลยในเบราว์เซอร์ และได้รับข้อเสนอแนะจาก AI ในทันที — ไม่ต้องติดตั้งในเครื่องของคุณ
บทเรียนทั้งหมดในหลักสูตรนี้
- SGD พร้อมโมเมนตัม
- อธิบาย Adam และ AdamW
- การเสื่อมน้ำหนักกับการทำให้เป็นระเบียบแบบ L2
- ตารางอัตราการเรียนรู้และการวอร์มอัป