SSE/AVX命令セット入門
並列データ処理向けに設計された最新のSIMD命令セット(SSE、AVX)と、それらのレジスターについて概要を学びます。
「SSE/AVX命令セット入門」はCoddyKit上の無料Assembly Language & x86 Low-Level Systems Programmingレッスンです。 これはレッスン2/4です。 下記で完全なレッスンを無料で読むことができます。その後、ブラウザ内の組み込みコードエディタと24時間対応のAIチューターでハンズオン演習できます。 これはAssembly Language & x86 Low-Level Systems Programming学習パスの一部であり、ウェブとCoddyKitアプリ全体で進捗が同期されます。 Assembly Language & x86 Low-Level Systems Programmingコースには全4レッスンが含まれています。
このレッスンの一部はまだ翻訳されておらず、英語で表示されています。
Welcome to SIMD!
In this lesson, we'll explore powerful instruction sets designed for parallel processing: SSE and AVX. These extensions allow your CPU to perform the same operation on multiple pieces of data simultaneously.
This technique, called Single Instruction, Multiple Data (SIMD), is crucial for speeding up tasks like graphics rendering, scientific calculations, and video processing.
Scalar vs. Vector Processing
Imagine you need to add two lists of numbers. A traditional scalar processor adds them one pair at a time:
- 1st number + 1st number
- 2nd number + 2nd number
- ...and so on.
A vector processor (using SIMD) can add multiple pairs in a single instruction, significantly faster for large datasets.
Introducing SSE
SSE stands for Streaming SIMD Extensions. Introduced by Intel, SSE brought 128-bit wide registers and instructions to x86 processors.
This means an SSE instruction can process 128 bits of data in one go. For single-precision floating-point numbers (each 32 bits), this allows processing four numbers at once.
SSE Registers: XMM0-XMM15
SSE uses a dedicated set of 16 XMM registers, named XMM0 through XMM15. Each XMM register is 128 bits wide.
- They can hold four 32-bit single-precision floating-point numbers.
- Or two 64-bit double-precision floating-point numbers.
- Or sixteen 8-bit integers.
These registers are separate from the general-purpose registers (like EAX, EBX).
SSE Data Movement: MOVAPS
Let's look at a basic SSE instruction: MOVAPS. This instruction moves aligned packed single-precision floating-point values. It loads 128 bits of data from memory into an XMM register.
Try running this example (using NASM syntax for Linux x86-64):
section .data
; Define 4 single-precision floats (128 bits total)
my_sse_data dd 1.0, 2.0, 3.0, 4.0
section .text
global _start
_start:
; Load 128 bits from my_sse_data into XMM0
movaps xmm0, [my_sse_data]
; XMM0 now holds [1.0, 2.0, 3.0, 4.0]
; Exit the program (Linux x86/x64 syscall)
mov eax, 1 ; sys_exit
xor ebx, ebx ; exit code 0
int 0x80 ; Invoke kernel
Beyond SSE: AVX
AVX, or Advanced Vector Extensions, is a further enhancement to SIMD processing. AVX expands the SIMD registers to 256 bits, doubling the amount of data processed per instruction compared to SSE.
AVX also introduced a new instruction encoding scheme (VEX prefix) and non-destructive operations, meaning destination registers don't overwrite source registers by default.
AVX Registers: YMM0-YMM15
AVX introduced 16 new YMM registers (YMM0 through YMM15). Each YMM register is 256 bits wide.
- A YMM register can hold eight 32-bit single-precision floats.
- Or four 64-bit double-precision floats.
The lower 128 bits of each YMM register overlap with the corresponding XMM register (e.g., the lower half of YMM0 is XMM0).
AVX Data Movement: VMOVAPS
The AVX equivalent of MOVAPS is VMOVAPS. The 'V' prefix indicates a VEX-encoded AVX instruction. This instruction moves aligned packed single-precision floating-point values, but now 256 bits at a time.
Here's an example demonstrating loading 8 floats into a YMM register:
section .data
; Define 8 single-precision floats (256 bits total)
my_avx_data dd 1.0, 2.0, 3.0, 4.0, 5.0, 6.0, 7.0, 8.0
section .text
global _start
_start:
; Load 256 bits from my_avx_data into YMM0
vmovaps ymm0, [my_avx_data]
; YMM0 now holds [1.0, 2.0, 3.0, 4.0, 5.0, 6.0, 7.0, 8.0]
; Exit the program (Linux x86/x64 syscall)
mov eax, 1 ; sys_exit
xor ebx, ebx ; exit code 0
int 0x80 ; Invoke kernel
Key Differences: SSE vs. AVX
While both SSE and AVX are SIMD extensions, AVX offers significant advancements:
- Register Size: SSE uses 128-bit XMM registers; AVX uses 256-bit YMM registers.
- Data Throughput: AVX can process twice as much data per instruction as SSE.
- Non-Destructive Ops: Many AVX instructions allow a three-operand format, keeping source operands intact.
- VEX Prefix: AVX instructions use a VEX prefix, enabling more flexible encoding and future extensions.
Quick Check: SIMD Registers
Which of the following statements about SSE and AVX registers are TRUE?
Recap: Power of SIMD
We've introduced SSE and AVX, powerful SIMD instruction sets that allow your CPU to perform operations on multiple data items simultaneously. You learned about:
- The concept of SIMD and its benefits for parallel processing.
- SSE with its 128-bit XMM registers (
XMM0-XMM15). - AVX with its 256-bit YMM registers (
YMM0-YMM15). - Basic data movement instructions like
MOVAPSandVMOVAPS.
Understanding these instruction sets is key to optimizing performance for data-intensive tasks.
AI チューターと学ぶ Assembly — 無料
ブラウザでリアルコードを書いて実行し、24/7 の AI チューターから瞬時にサポートを受け、ウェブまたはアプリで続きから学習できます。
- コース
- 12
- レッスン
- 48
よくある質問
「SSE/AVX命令セット入門」レッスンは無料ですか?
はい。「SSE/AVX命令セット入門」の完全なテキストはこのウェブで無料で読めます。インタラクティブに演習し(組み込みコードエディタと24時間対応のAIチューター)、Assembly Language & x86 Low-Level Systems Programmingコースの残りをアンロックするには、CoddyKit PROにアップグレードしてください。 Assembly Language & x86 Low-Level Systems Programmingコースには全4レッスンが含まれています。
「SSE/AVX命令セット入門」で何を学びますか?
並列データ処理向けに設計された最新のSIMD命令セット(SSE、AVX)と、それらのレジスターについて概要を学びます。 ブラウザで直接実行するハンズオンコードでAssembly Language & x86 Low-Level Systems Programmingを演習し、24時間対応のAIチューターがレッスンを進める中での質問に答えます。
Assembly Language & x86 Low-Level Systems Programmingを始めるのに経験は必要ですか?
事前経験は必要ありません。CoddyKitのAssembly Language & x86 Low-Level Systems Programmingは初級者から上級者向けに構成されているため、ここから始めるか最初から始めて、自分のペースで進むことができます。 これはレッスン2/4です。
「SSE/AVX命令セット入門」レッスンにはどのくらい時間がかかりますか?
ほとんどのCoddyKitレッスンは約5~10分かかります。各レッスンはコンパクトでインタラクティブなので、着実に進歩し、ウェブとアプリ全体で正確に前回の場所から再開できます。
このAssembly Language & x86 Low-Level Systems Programmingレッスンでコードを書いて実行できますか?
はい。すべてのAssembly Language & x86 Low-Level Systems Programmingレッスンに組み込みコードエディタが含まれているため、ブラウザでリアルコードを書いて実行し、即座のAIフィードバックを取得できます。ローカル設定は不要です。
このコースのすべてのレッスン
- x87 FPUプログラミングの基礎
- SSE/AVX命令セット入門
- SIMDによるコードのベクトル化
- 浮動小数点の精度、丸め、例外