Data Representation and Types
Learn how integers, characters, and other data types are stored in memory and manipulated using assembly instructions.
Data Representation and Types is a free Assembly Language & x86 Low-Level Systems Programming lesson on CoddyKit — lesson 3 of 4. You can read the complete lesson below for free — then practise it hands-on in the browser with a built-in code editor and a 24/7 AI tutor. It is part of the Assembly Language & x86 Low-Level Systems Programming learning path, one of 4 lessons in the course, and your progress syncs across the web and the CoddyKit app.
What are Data Types?
In assembly language, we work directly with raw bits and bytes. But how do we know if a sequence of bytes represents a number, a character, or something else?
This is where data types come in! They give meaning to the raw data, helping both you and the CPU understand how to interpret and manipulate information.
Common Data Sizes
x86 assembly defines standard sizes for data. These directly correspond to how much memory space a piece of data occupies:
- BYTE: 8 bits
- WORD: 16 bits (2 bytes)
- DWORD: Double Word, 32 bits (4 bytes)
- QWORD: Quad Word, 64 bits (8 bytes)
These sizes are fundamental for declaring variables and working with registers.
Defining Data: DB, DW, DD, DQ
To store data in memory, we use data definition directives. These tell the assembler to reserve space and optionally initialize it with a value.
DB: Define Byte (8-bit)DW: Define Word (16-bit)DD: Define Doubleword (32-bit)DQ: Define Quadword (64-bit)
You'll see these often when creating variables in your programs.
Unsigned Integers
An unsigned integer is a number that is always positive or zero. All of its bits are used to represent the magnitude of the number.
For example, an 8-bit unsigned byte can hold values from 0 to 255. A 16-bit unsigned word can hold values from 0 to 65,535.
When you don't need negative numbers, unsigned types are perfect and give you a larger positive range.
Signed Integers (Two's Complement)
Signed integers can represent both positive and negative values. One bit, usually the Most Significant Bit (MSB), is used to indicate the sign (0 for positive, 1 for negative).
Negative numbers are typically represented using Two's Complement. This system makes arithmetic operations work seamlessly for both positive and negative values.
An 8-bit signed byte ranges from -128 to +127.
Character Data: ASCII
Characters like 'A', 'b', or '7' are also stored as numbers! The most common standard for this is ASCII (American Standard Code for Information Interchange).
Each character is assigned a unique 8-bit (1-byte) numerical value. For example, the character 'A' is represented by the decimal value 65 (or hexadecimal 0x41).
You can define single characters or entire strings using the DB directive.
Code: Defining & Accessing Data
This example shows how to define different data types and then load their values into CPU registers. This demonstrates how assembly treats these named memory locations.
section .data
; Define various data types
myByte db 10 ; An 8-bit unsigned integer
myWord dw 256 ; A 16-bit unsigned integer
myDword dd 65536 ; A 32-bit unsigned integer
myChar db 'X' ; An 8-bit character (ASCII value 88)
myString db "Hello", 0 ; A string (null-terminated)
section .text
global _start
_start:
; Move byte into AL register
mov al, [myByte]
; Move word into BX register
mov bx, [myWord]
; Move dword into ECX register
mov ecx, [myDword]
; Move char into DL register
mov dl, [myChar]
; Exit gracefully (Linux syscall)
mov eax, 1 ; sys_exit syscall number
xor ebx, ebx ; Exit code 0
int 0x80 ; Invoke kernel
Data Alignment Benefits
Data alignment means placing data in memory at an address that is a multiple of its size. For example, a DWORD (4 bytes) might be aligned to an address ending in 0, 4, 8, or C (hex).
While not strictly required by all CPUs, proper alignment can significantly improve performance. The CPU can fetch aligned data more efficiently, often in a single memory access, avoiding extra work.
Assemblers sometimes provide directives like ALIGN to help ensure proper alignment.
Why Data Types Matter
Understanding data types is crucial because it dictates:
- Memory Usage: How much space your data consumes.
- Instruction Choice: Which assembly instructions (e.g.,
ADD,MOV) are appropriate for the data size. - Interpretation: Whether the CPU treats
0xFFas255(unsigned) or-1(signed).
Careful type selection prevents errors and ensures your programs behave as expected at the lowest level.
Quick Check: Data Sizes
You've learned about common data sizes and how they're defined. Let's test your knowledge!
Recap: Data Representation
Great job! You've explored the fundamental concepts of data representation in x86 assembly.
- We define data using directives like
DB,DW,DD, andDQfor various sizes. - Integers can be signed (positive/negative) or unsigned (positive only).
- Characters are stored using the ASCII standard, where each character has a numerical value.
- Understanding data alignment can help optimize performance.
Next, we'll continue building on this knowledge to perform more complex operations!
Frequently asked questions
Is the “Data Representation and Types” lesson free?
Yes — the full text of “Data Representation and Types” is free to read here on the web, and the Assembly Language & x86 Low-Level Systems Programming course includes 4 lessons in total. To practise it interactively (a built-in code editor and a 24/7 AI tutor) and unlock the rest of the Assembly Language & x86 Low-Level Systems Programming course, upgrade to CoddyKit PRO.
What will I learn in “Data Representation and Types”?
Learn how integers, characters, and other data types are stored in memory and manipulated using assembly instructions. You practise Assembly Language & x86 Low-Level Systems Programming with hands-on code you run directly in the browser, and a 24/7 AI tutor answers your questions as you work through the lesson.
Do I need any experience to start Assembly Language & x86 Low-Level Systems Programming?
No prior experience is required. Assembly Language & x86 Low-Level Systems Programming on CoddyKit is structured for beginners through advanced learners; this is — lesson 3 of 4, so you can start here or from the beginning and move at your own pace.
How long does the “Data Representation and Types” lesson take?
Most CoddyKit lessons take about 5–10 minutes. Each one is bite-sized and interactive, so you make steady progress and pick up exactly where you left off across the web and the app.
Can I write and run code in this Assembly Language & x86 Low-Level Systems Programming lesson?
Yes. Every Assembly Language & x86 Low-Level Systems Programming lesson includes a built-in code editor, so you write and run real code right in your browser and get instant AI feedback — no local setup required.
All lessons in this course
- x86 Registers Demystified
- Memory Addressing Modes
- Data Representation and Types
- The FLAGS Register and Status Bits