0Pricing
Assembly Language & x86 Low-Level Systems Programming · Lesson

x87 FPU Programming Basics

Learn how to use the x87 Floating-Point Unit (FPU) for high-precision floating-point arithmetic in assembly language.

x87 FPU Programming Basics is a free Assembly Language & x86 Low-Level Systems Programming lesson on CoddyKit — lesson 1 of 4. You can read the complete lesson below for free — then practise it hands-on in the browser with a built-in code editor and a 24/7 AI tutor. It is part of the Assembly Language & x86 Low-Level Systems Programming learning path, one of 4 lessons in the course, and your progress syncs across the web and the CoddyKit app.

Meet the x87 FPU

The x87 Floating-Point Unit (FPU) is a specialized part of the CPU designed to handle mathematical operations on real numbers, also known as floating-point numbers.

Unlike integer arithmetic, floating-point math requires different internal representations and calculations, which the FPU excels at with high precision.

The FPU Register Stack

The x87 FPU uses a unique register stack, not general-purpose registers. This stack consists of eight 80-bit registers, denoted as ST(0) through ST(7).

  • ST(0) is always the top of the stack.
  • Operations push new values onto the stack or pop values from it, shifting existing values.
  • It behaves like a Last-In, First-Out (LIFO) stack.

Loading Data with FLD

To work with floating-point numbers, you first need to load them onto the FPU stack. The FLD instruction pushes a floating-point value from memory onto the top of the FPU stack, making it ST(0).

In assembly, we often define floating-point constants in the data section. Common types are single-precision (32-bit, DD) and double-precision (64-bit, DQ).

Try loading a value onto the stack:

section .data
  float_val dq 3.1415926535

section .text
  global _start

_start:
  finit                 ; Initialize FPU
  fld qword [float_val] ; Load float_val onto FPU stack (ST(0))

  ; At this point, ST(0) contains 3.1415926535
  ; We just exit as printing floats is complex in basic assembly.

  mov rax, 60           ; syscall number for exit
  xor rdi, rdi          ; exit code 0
  syscall

Storing FPU Results

After calculations, you'll want to store the result from the FPU stack back into memory. The FST and FSTP instructions are used for this.

  • FST: Copies the value from ST(0) to a memory location or another FPU register, leaving ST(0) unchanged.
  • FSTP: Copies the value from ST(0) to memory/register, then pops it from the stack, decreasing the stack pointer. The previous ST(1) becomes ST(0).

Let's store a value:

section .data
  float_val dq 123.45
  result_val dq 0.0     ; Will store result here

section .text
  global _start

_start:
  finit                 ; Initialize FPU
  fld qword [float_val] ; ST(0) = 123.45

  fstp qword [result_val] ; Store ST(0) to result_val, then pop.
                          ; FPU stack is now empty.

  ; result_val now holds 123.45 in memory.

  mov rax, 60           ; syscall number for exit
  xor rdi, rdi          ; exit code 0
  syscall

FPU Arithmetic Operations

The FPU provides instructions for common arithmetic operations. These typically operate on ST(0) and another operand (either another stack register or a memory operand).

  • FADD: Add (e.g., FADD ST(1), ST(0) adds ST(0) to ST(1)).
  • FMUL: Multiply
  • FSUB: Subtract
  • FDIV: Divide

Using FADDP ST(1), ST(0) adds ST(0) to ST(1), stores in ST(1), and pops ST(0). This leaves the sum on top of the stack.

Here's an addition example:

section .data
  val1 dq 10.5
  val2 dq 2.0
  sum_result dq 0.0

section .text
  global _start

_start:
  finit                 ; Initialize FPU
  fld qword [val1]      ; ST(0) = 10.5
  fld qword [val2]      ; ST(0) = 2.0, ST(1) = 10.5

  faddp st(1), st(0)    ; ST(1) = ST(1) + ST(0) (10.5 + 2.0 = 12.5).
                        ; Pop ST(0). Now ST(0) = 12.5.

  fstp qword [sum_result] ; Store 12.5 to sum_result and pop.

  mov rax, 60           ; syscall number for exit
  xor rdi, rdi          ; exit code 0
  syscall

FPU Built-in Constants

The FPU can load commonly used constants directly onto its stack, saving you from defining them in memory. This improves efficiency and precision.

  • FLD1: Pushes 1.0 onto the stack.
  • FLDZ: Pushes 0.0 onto the stack.
  • FLDPI: Pushes the value of Pi (π) onto the stack.

Let's load Pi:

section .data
  pi_val dq 0.0 ; To store PI

section .text
  global _start

_start:
  finit           ; Initialize FPU
  fldpi           ; ST(0) = PI (approx 3.14159...)

  fstp qword [pi_val] ; Store PI to pi_val and pop.

  mov rax, 60     ; syscall number for exit
  xor rdi, rdi    ; exit code 0
  syscall

Converting Integers & Floats

Sometimes you need to convert between integer and floating-point types. The FPU provides instructions for this:

  • FILD (Float Integer Load): Loads a signed integer from memory, converts it to a floating-point format, and pushes it onto the FPU stack.
  • FISTP (Float Integer Store and Pop): Stores ST(0) as an integer to memory and then pops it from the stack. The value is truncated towards zero during conversion.

Let's convert an integer to a float, add, then convert back:

section .data
  int_val dd 5
  float_add dq 2.5
  int_result dd 0

section .text
  global _start

_start:
  finit               ; Initialize FPU
  fild dword [int_val] ; ST(0) = 5.0 (from 5)
  fld qword [float_add] ; ST(0) = 2.5, ST(1) = 5.0

  faddp st(1), st(0)  ; ST(1) = 5.0 + 2.5 = 7.5. Pop ST(0).
                      ; Now ST(0) = 7.5.

  fistp dword [int_result] ; Store 7.5 as integer (7) to int_result and pop.

  mov rax, 60         ; syscall number for exit
  xor rdi, rdi        ; exit code 0
  syscall

Comparing Floating-Point Values

Comparing floating-point numbers requires special FPU instructions. You can't directly use integer comparison instructions like CMP.

  • FCOM: Compares ST(0) with an operand (another FPU register or memory) and sets FPU status flags.
  • FCOMP: Same as FCOM, but pops ST(0) after comparison.

To use these flags for conditional jumps (like JE, JB), you must transfer them from the FPU status word to the CPU's EFLAGS register:

  1. FSTSW AX: Stores the FPU Status Word into the AX register.
  2. SAHF: Transfers the AH register (which now contains the relevant FPU flags) into the CPU's EFLAGS register, specifically the ZF, PF, and CF flags.

Putting it Together: (A + B) * C

Let's combine what we've learned to perform a simple calculation: (A + B) * C. We'll load three values, add two, multiply by the third, and store the final integer result.

This example demonstrates stack manipulation and arithmetic operations.

section .data
  val_A dq 3.0
  val_B dq 1.5
  val_C dq 2.0
  final_int_result dd 0

section .text
  global _start

_start:
  finit           ; Initialize FPU

  fld qword [val_A] ; ST(0) = 3.0
  fld qword [val_B] ; ST(0) = 1.5, ST(1) = 3.0

  faddp st(1), st(0) ; Add ST(0) (1.5) to ST(1) (3.0), store in ST(1).
                     ; Pop ST(0). Now ST(0) = 4.5 (sum of A+B)

  fld qword [val_C] ; ST(0) = 2.0, ST(1) = 4.5 (A+B)

  fmulp st(1), st(0) ; Multiply ST(0) (2.0) by ST(1) (4.5), store in ST(1).
                     ; Pop ST(0). Now ST(0) = 9.0 ((A+B)*C)

  fistp dword [final_int_result] ; Store 9.0 as integer (9) to final_int_result and pop.

  mov rax, 60     ; syscall number for exit
  xor rdi, rdi    ; exit code 0
  syscall

FPU Stack Challenge

Consider the following x87 FPU assembly code snippet. What will be the value of ST(0) after its execution?

section .data
  val_X dq 10.0
  val_Y dq 3.0

section .text
  finit
  fld qword [val_X] ; ST(0) = 10.0
  fld qword [val_Y] ; ST(0) = 3.0, ST(1) = 10.0
  faddp st(1), st(0) ; ST(0) = 13.0
  fld1              ; ST(0) = 1.0, ST(1) = 13.0
  fsub              ; ST(0) = ST(0) - ST(1) (1.0 - 13.0 = -12.0)

x87 FPU Summary

You've taken your first steps into x87 FPU programming! We covered:

  • The FPU's 8-register stack (ST(0) to ST(7)).
  • Loading values with FLD and storing with FST/FSTP.
  • Basic arithmetic: FADD, FSUB, FMUL, FDIV.
  • Using built-in constants like FLD1, FLDZ, FLDPI.
  • Converting between integers and floats with FILD and FISTP.
  • How to prepare FPU comparison results for conditional jumps.

The x87 FPU is powerful for precise calculations, though modern systems often use SIMD extensions like SSE/AVX for speed, which you'll explore next!

Frequently asked questions

Is the “x87 FPU Programming Basics” lesson free?

Yes — the full text of “x87 FPU Programming Basics” is free to read here on the web, and the Assembly Language & x86 Low-Level Systems Programming course includes 4 lessons in total. To practise it interactively (a built-in code editor and a 24/7 AI tutor) and unlock the rest of the Assembly Language & x86 Low-Level Systems Programming course, upgrade to CoddyKit PRO.

What will I learn in “x87 FPU Programming Basics”?

Learn how to use the x87 Floating-Point Unit (FPU) for high-precision floating-point arithmetic in assembly language. You practise Assembly Language & x86 Low-Level Systems Programming with hands-on code you run directly in the browser, and a 24/7 AI tutor answers your questions as you work through the lesson.

Do I need any experience to start Assembly Language & x86 Low-Level Systems Programming?

No prior experience is required. Assembly Language & x86 Low-Level Systems Programming on CoddyKit is structured for beginners through advanced learners; this is — lesson 1 of 4, so you can start here or from the beginning and move at your own pace.

How long does the “x87 FPU Programming Basics” lesson take?

Most CoddyKit lessons take about 5–10 minutes. Each one is bite-sized and interactive, so you make steady progress and pick up exactly where you left off across the web and the app.

Can I write and run code in this Assembly Language & x86 Low-Level Systems Programming lesson?

Yes. Every Assembly Language & x86 Low-Level Systems Programming lesson includes a built-in code editor, so you write and run real code right in your browser and get instant AI feedback — no local setup required.

All lessons in this course

  1. x87 FPU Programming Basics
  2. SSE/AVX Instruction Sets Introduction
  3. Vectorizing Code with SIMD
  4. Floating-Point Precision, Rounding, and Exceptions
← Back to Assembly Language & x86 Low-Level Systems Programming