Assembly Language & x86 Low-Level Systems Programming · 课时

SSE/AVX 指令集简介

概览现代 SIMD 指令集(SSE、AVX)及其寄存器,了解它们如何用于并行数据处理。

第 2 / 4 课11 个步骤

SSE/AVX 指令集简介 是 CoddyKit 上的免费 Assembly Language & x86 Low-Level Systems Programming 课时。 这是第 2 节课,共 4 节。 你可以在下方免费阅读本课时的完整内容 — 然后在浏览器中使用内置代码编辑器和全天候 AI 导师进行实践。 这是 Assembly Language & x86 Low-Level Systems Programming 学习路径的一部分,你的进度在网页和 CoddyKit 应用中同步。 Assembly Language & x86 Low-Level Systems Programming 课程共包含 4 节课。

本课时的部分内容尚未翻译,以英文显示。

Welcome to SIMD!

In this lesson, we'll explore powerful instruction sets designed for parallel processing: SSE and AVX. These extensions allow your CPU to perform the same operation on multiple pieces of data simultaneously.

This technique, called Single Instruction, Multiple Data (SIMD), is crucial for speeding up tasks like graphics rendering, scientific calculations, and video processing.

Scalar vs. Vector Processing

Imagine you need to add two lists of numbers. A traditional scalar processor adds them one pair at a time:

  • 1st number + 1st number
  • 2nd number + 2nd number
  • ...and so on.

A vector processor (using SIMD) can add multiple pairs in a single instruction, significantly faster for large datasets.

Introducing SSE

SSE stands for Streaming SIMD Extensions. Introduced by Intel, SSE brought 128-bit wide registers and instructions to x86 processors.

This means an SSE instruction can process 128 bits of data in one go. For single-precision floating-point numbers (each 32 bits), this allows processing four numbers at once.

SSE Registers: XMM0-XMM15

SSE uses a dedicated set of 16 XMM registers, named XMM0 through XMM15. Each XMM register is 128 bits wide.

  • They can hold four 32-bit single-precision floating-point numbers.
  • Or two 64-bit double-precision floating-point numbers.
  • Or sixteen 8-bit integers.

These registers are separate from the general-purpose registers (like EAX, EBX).

SSE Data Movement: MOVAPS

Let's look at a basic SSE instruction: MOVAPS. This instruction moves aligned packed single-precision floating-point values. It loads 128 bits of data from memory into an XMM register.

Try running this example (using NASM syntax for Linux x86-64):

section .data
  ; Define 4 single-precision floats (128 bits total)
  my_sse_data dd 1.0, 2.0, 3.0, 4.0

section .text
  global _start

_start:
  ; Load 128 bits from my_sse_data into XMM0
  movaps xmm0, [my_sse_data]

  ; XMM0 now holds [1.0, 2.0, 3.0, 4.0]

  ; Exit the program (Linux x86/x64 syscall)
  mov eax, 1    ; sys_exit
  xor ebx, ebx  ; exit code 0
  int 0x80      ; Invoke kernel

Beyond SSE: AVX

AVX, or Advanced Vector Extensions, is a further enhancement to SIMD processing. AVX expands the SIMD registers to 256 bits, doubling the amount of data processed per instruction compared to SSE.

AVX also introduced a new instruction encoding scheme (VEX prefix) and non-destructive operations, meaning destination registers don't overwrite source registers by default.

AVX Registers: YMM0-YMM15

AVX introduced 16 new YMM registers (YMM0 through YMM15). Each YMM register is 256 bits wide.

  • A YMM register can hold eight 32-bit single-precision floats.
  • Or four 64-bit double-precision floats.

The lower 128 bits of each YMM register overlap with the corresponding XMM register (e.g., the lower half of YMM0 is XMM0).

AVX Data Movement: VMOVAPS

The AVX equivalent of MOVAPS is VMOVAPS. The 'V' prefix indicates a VEX-encoded AVX instruction. This instruction moves aligned packed single-precision floating-point values, but now 256 bits at a time.

Here's an example demonstrating loading 8 floats into a YMM register:

section .data
  ; Define 8 single-precision floats (256 bits total)
  my_avx_data dd 1.0, 2.0, 3.0, 4.0, 5.0, 6.0, 7.0, 8.0

section .text
  global _start

_start:
  ; Load 256 bits from my_avx_data into YMM0
  vmovaps ymm0, [my_avx_data]

  ; YMM0 now holds [1.0, 2.0, 3.0, 4.0, 5.0, 6.0, 7.0, 8.0]

  ; Exit the program (Linux x86/x64 syscall)
  mov eax, 1    ; sys_exit
  xor ebx, ebx  ; exit code 0
  int 0x80      ; Invoke kernel

Key Differences: SSE vs. AVX

While both SSE and AVX are SIMD extensions, AVX offers significant advancements:

  • Register Size: SSE uses 128-bit XMM registers; AVX uses 256-bit YMM registers.
  • Data Throughput: AVX can process twice as much data per instruction as SSE.
  • Non-Destructive Ops: Many AVX instructions allow a three-operand format, keeping source operands intact.
  • VEX Prefix: AVX instructions use a VEX prefix, enabling more flexible encoding and future extensions.

Quick Check: SIMD Registers

Which of the following statements about SSE and AVX registers are TRUE?

Recap: Power of SIMD

We've introduced SSE and AVX, powerful SIMD instruction sets that allow your CPU to perform operations on multiple data items simultaneously. You learned about:

  • The concept of SIMD and its benefits for parallel processing.
  • SSE with its 128-bit XMM registers (XMM0-XMM15).
  • AVX with its 256-bit YMM registers (YMM0-YMM15).
  • Basic data movement instructions like MOVAPS and VMOVAPS.

Understanding these instruction sets is key to optimizing performance for data-intensive tasks.

免费开始

用 AI 导师学习 Assembly — 免费

在浏览器中编写并运行真实代码,获得全天候 AI 导师的即时帮助,并在网页或应用中继续学习。

课程
12
课程
48

常见问题解答

「SSE/AVX 指令集简介」课时是免费的吗?

是的 — 「SSE/AVX 指令集简介」的完整文本可在网页上免费阅读。要进行交互式练习(内置代码编辑器和全天候 AI 导师)并解锁 Assembly Language & x86 Low-Level Systems Programming 课程的其余内容,请升级到 CoddyKit PRO。 Assembly Language & x86 Low-Level Systems Programming 课程共包含 4 节课。

「SSE/AVX 指令集简介」这节课中我会学到什么?

概览现代 SIMD 指令集(SSE、AVX)及其寄存器,了解它们如何用于并行数据处理。 你通过在浏览器中直接运行的动手代码来练习 Assembly Language & x86 Low-Level Systems Programming,全天候 AI 导师会在你学习这节课的过程中回答你的问题。

学习 Assembly Language & x86 Low-Level Systems Programming 需要有经验吗?

无需任何先前经验。CoddyKit 上的 Assembly Language & x86 Low-Level Systems Programming 课程适合初学者到高级学习者,你可以从这里开始或从头开始,按照自己的节奏学习。 这是第 2 节课,共 4 节。

「SSE/AVX 指令集简介」课时需要多长时间?

大多数 CoddyKit 课程大约需要 5–10 分钟。每节课都很精短且互动,所以你能稳步进步,并在网页和应用中从离开的地方继续。

我能在这节 Assembly Language & x86 Low-Level Systems Programming 课中编写并运行代码吗?

能。每节 Assembly Language & x86 Low-Level Systems Programming 课都包含内置代码编辑器,你可以在浏览器中直接编写并运行真实代码,并获得即时 AI 反馈 — 无需本地设置。

此课程中的所有课时

  1. x87 FPU 编程基础
  2. SSE/AVX 指令集简介
  3. 使用 SIMD 将代码向量化
  4. 浮点精度、舍入与异常
← 返回 Assembly Language & x86 Low-Level Systems Programming