Assembly Language & x86 Low-Level Systems Programming · 강의

SSE/AVX 명령어 집합 소개

병렬 데이터 처리를 위해 설계된 최신 SIMD 명령어 집합(SSE, AVX)과 해당 레지스터를 개괄적으로 살펴봅니다.

레슨 2/411개 단계

SSE/AVX 명령어 집합 소개은(는) CoddyKit의 무료 Assembly Language & x86 Low-Level Systems Programming 강의입니다. 이것은 4개 중 2번째 강의입니다. 아래에서 전체 강의를 무료로 읽을 수 있으며, 내장 코드 에디터와 24/7 AI 튜터와 함께 브라우저에서 직접 실습할 수 있습니다. 이 강의는 Assembly Language & x86 Low-Level Systems Programming 학습 경로의 일부이며, 진행 상황이 웹과 CoddyKit 앱에 동기화됩니다. Assembly Language & x86 Low-Level Systems Programming 강의에는 총 4개의 강의가 포함되어 있습니다.

이 강의의 일부는 아직 번역되지 않았으며 영어로 표시됩니다.

Welcome to SIMD!

In this lesson, we'll explore powerful instruction sets designed for parallel processing: SSE and AVX. These extensions allow your CPU to perform the same operation on multiple pieces of data simultaneously.

This technique, called Single Instruction, Multiple Data (SIMD), is crucial for speeding up tasks like graphics rendering, scientific calculations, and video processing.

Scalar vs. Vector Processing

Imagine you need to add two lists of numbers. A traditional scalar processor adds them one pair at a time:

  • 1st number + 1st number
  • 2nd number + 2nd number
  • ...and so on.

A vector processor (using SIMD) can add multiple pairs in a single instruction, significantly faster for large datasets.

Introducing SSE

SSE stands for Streaming SIMD Extensions. Introduced by Intel, SSE brought 128-bit wide registers and instructions to x86 processors.

This means an SSE instruction can process 128 bits of data in one go. For single-precision floating-point numbers (each 32 bits), this allows processing four numbers at once.

SSE Registers: XMM0-XMM15

SSE uses a dedicated set of 16 XMM registers, named XMM0 through XMM15. Each XMM register is 128 bits wide.

  • They can hold four 32-bit single-precision floating-point numbers.
  • Or two 64-bit double-precision floating-point numbers.
  • Or sixteen 8-bit integers.

These registers are separate from the general-purpose registers (like EAX, EBX).

SSE Data Movement: MOVAPS

Let's look at a basic SSE instruction: MOVAPS. This instruction moves aligned packed single-precision floating-point values. It loads 128 bits of data from memory into an XMM register.

Try running this example (using NASM syntax for Linux x86-64):

section .data
  ; Define 4 single-precision floats (128 bits total)
  my_sse_data dd 1.0, 2.0, 3.0, 4.0

section .text
  global _start

_start:
  ; Load 128 bits from my_sse_data into XMM0
  movaps xmm0, [my_sse_data]

  ; XMM0 now holds [1.0, 2.0, 3.0, 4.0]

  ; Exit the program (Linux x86/x64 syscall)
  mov eax, 1    ; sys_exit
  xor ebx, ebx  ; exit code 0
  int 0x80      ; Invoke kernel

Beyond SSE: AVX

AVX, or Advanced Vector Extensions, is a further enhancement to SIMD processing. AVX expands the SIMD registers to 256 bits, doubling the amount of data processed per instruction compared to SSE.

AVX also introduced a new instruction encoding scheme (VEX prefix) and non-destructive operations, meaning destination registers don't overwrite source registers by default.

AVX Registers: YMM0-YMM15

AVX introduced 16 new YMM registers (YMM0 through YMM15). Each YMM register is 256 bits wide.

  • A YMM register can hold eight 32-bit single-precision floats.
  • Or four 64-bit double-precision floats.

The lower 128 bits of each YMM register overlap with the corresponding XMM register (e.g., the lower half of YMM0 is XMM0).

AVX Data Movement: VMOVAPS

The AVX equivalent of MOVAPS is VMOVAPS. The 'V' prefix indicates a VEX-encoded AVX instruction. This instruction moves aligned packed single-precision floating-point values, but now 256 bits at a time.

Here's an example demonstrating loading 8 floats into a YMM register:

section .data
  ; Define 8 single-precision floats (256 bits total)
  my_avx_data dd 1.0, 2.0, 3.0, 4.0, 5.0, 6.0, 7.0, 8.0

section .text
  global _start

_start:
  ; Load 256 bits from my_avx_data into YMM0
  vmovaps ymm0, [my_avx_data]

  ; YMM0 now holds [1.0, 2.0, 3.0, 4.0, 5.0, 6.0, 7.0, 8.0]

  ; Exit the program (Linux x86/x64 syscall)
  mov eax, 1    ; sys_exit
  xor ebx, ebx  ; exit code 0
  int 0x80      ; Invoke kernel

Key Differences: SSE vs. AVX

While both SSE and AVX are SIMD extensions, AVX offers significant advancements:

  • Register Size: SSE uses 128-bit XMM registers; AVX uses 256-bit YMM registers.
  • Data Throughput: AVX can process twice as much data per instruction as SSE.
  • Non-Destructive Ops: Many AVX instructions allow a three-operand format, keeping source operands intact.
  • VEX Prefix: AVX instructions use a VEX prefix, enabling more flexible encoding and future extensions.

Quick Check: SIMD Registers

Which of the following statements about SSE and AVX registers are TRUE?

Recap: Power of SIMD

We've introduced SSE and AVX, powerful SIMD instruction sets that allow your CPU to perform operations on multiple data items simultaneously. You learned about:

  • The concept of SIMD and its benefits for parallel processing.
  • SSE with its 128-bit XMM registers (XMM0-XMM15).
  • AVX with its 256-bit YMM registers (YMM0-YMM15).
  • Basic data movement instructions like MOVAPS and VMOVAPS.

Understanding these instruction sets is key to optimizing performance for data-intensive tasks.

무료로 시작

AI 튜터와 함께 Assembly을(를) 배우세요 — 무료

브라우저에서 실제 코드를 작성하고 실행하며, 24/7 AI 튜터로부터 즉각적인 도움을 받고, 웹이나 앱에서 중단한 부분부터 계속 학습하세요.

코스
12
레슨
48

자주 묻는 질문

“SSE/AVX 명령어 집합 소개” 강의는 무료인가요?

네 — “SSE/AVX 명령어 집합 소개” 전체 내용을 이 웹사이트에서 무료로 읽을 수 있습니다. 인터랙티브하게 실습하려면(내장 코드 에디터와 24/7 AI 튜터), CoddyKit PRO로 업그레이드하면 Assembly Language & x86 Low-Level Systems Programming 강의 전체를 잠금 해제할 수 있습니다. Assembly Language & x86 Low-Level Systems Programming 강의에는 총 4개의 강의가 포함되어 있습니다.

“SSE/AVX 명령어 집합 소개”에서 뭘 배우나요?

병렬 데이터 처리를 위해 설계된 최신 SIMD 명령어 집합(SSE, AVX)과 해당 레지스터를 개괄적으로 살펴봅니다. 브라우저에서 직접 실행하는 실습 코드로 Assembly Language & x86 Low-Level Systems Programming을(를) 배우며, 24/7 AI 튜터가 강의를 진행하면서 질문에 답변해줍니다.

Assembly Language & x86 Low-Level Systems Programming을(를) 시작하는 데 경험이 필요한가요?

사전 경험은 필요하지 않습니다. CoddyKit의 Assembly Language & x86 Low-Level Systems Programming은(는) 초급자부터 고급 학습자까지를 위해 구성되어 있으므로, 여기서 시작하거나 처음부터 시작할 수 있으며 자신의 속도대로 진행할 수 있습니다. 이것은 4개 중 2번째 강의입니다.

“SSE/AVX 명령어 집합 소개” 강의는 얼마나 걸리나요?

대부분의 CoddyKit 강의는 약 5~10분이 소요됩니다. 각 강의는 간결하고 인터랙티브하여 꾸준한 진행이 가능하며, 웹과 앱에서 중단한 부분부터 바로 시작할 수 있습니다.

이 Assembly Language & x86 Low-Level Systems Programming 강의에서 코드를 작성하고 실행할 수 있나요?

네. 모든 Assembly Language & x86 Low-Level Systems Programming 강의에는 내장 코드 에디터가 포함되어 있으므로, 브라우저에서 바로 실제 코드를 작성하고 실행한 후 즉시 AI 피드백을 받을 수 있습니다 — 로컬 설정이 필요 없습니다.

이 강의의 모든 강의

  1. x87 FPU 프로그래밍 기초
  2. SSE/AVX 명령어 집합 소개
  3. SIMD를 사용한 코드 벡터화
  4. 부동소수점 정밀도, 반올림 및 예외
← Assembly Language & x86 Low-Level Systems Programming(으)로 돌아가기