Representação de Dados em Binários
Aprenda como tipos de dados (inteiros, números de ponto flutuante e strings) são armazenados na memória e em arquivos, incluindo conceitos como endianidade.
Representação de Dados em Binários é uma aula grátis de Reverse Engineering & Binary Analysis Basics no CoddyKit. Esta é a aula 2 de 4. Você pode ler a aula completa abaixo gratuitamente — depois pratica ao vivo no navegador com um editor de código integrado e um tutor de IA 24/7. Faz parte do caminho de aprendizado de Reverse Engineering & Binary Analysis Basics, e seu progresso é sincronizado entre a web e o app CoddyKit. O curso de Reverse Engineering & Binary Analysis Basics inclui 4 aulas no total.
Partes desta aula ainda não foram traduzidas e aparecem em inglês.
What is Binary Data?
When you reverse engineer, you're looking at a program's raw binary form. This means understanding how all kinds of data—numbers, text, and more—are stored as sequences of bits and bytes.
A bit is the smallest unit, either 0 or 1. Eight bits make a byte. Everything a computer does, from calculations to displaying text, relies on these fundamental units.
Numbers as Bits: Integers
Integers are whole numbers. They can be signed (positive or negative) or unsigned (only non-negative). The number of bytes used determines the range of values an integer can hold.
- Byte (8-bit): 0 to 255 (unsigned) or -128 to 127 (signed).
- Word (16-bit): Up to 65,535 (unsigned).
- DWord (32-bit): Up to ~4 billion (unsigned).
- QWord (64-bit): Much larger numbers!
Integer Size Matters
Let's see how different integer sizes affect the maximum value. This Python code shows the max value for an unsigned 8-bit integer (a byte) and a signed 8-bit integer.
# Max unsigned 8-bit integer
max_8_bit_unsigned = 2**8 - 1
print(f"Max 8-bit unsigned: {max_8_bit_unsigned}")
# Max signed 8-bit integer
max_8_bit_signed = 2**7 - 1
min_8_bit_signed = -2**7
print(f"Max 8-bit signed: {max_8_bit_signed}")
print(f"Min 8-bit signed: {min_8_bit_signed}")Floating-Point Numbers (Floats)
Numbers with decimal points, like 3.14 or -0.5, are called floating-point numbers. They are stored differently from integers to handle their fractional parts.
Most systems use the IEEE 754 standard for floats. This standard defines how a number's sign, exponent, and fractional part are represented in bits. Common sizes are 32-bit (single-precision) and 64-bit (double-precision).
Text: Characters & Strings
Text characters are also stored as numbers. The most common mapping for English characters is ASCII, where each character (like 'A' or '!') corresponds to a specific 8-bit number.
For a wider range of characters (emojis, foreign languages), Unicode is used. UTF-8 is a popular Unicode encoding that uses 1 to 4 bytes per character, making it flexible and backward-compatible with ASCII.
A string is simply a sequence of these characters, often ending with a special null byte (0x00) to mark its end.
How Strings Become Bytes
Here's how a simple string is represented as bytes using UTF-8. Notice how each character gets a numerical value.
message = "Hello"
bytes_message = message.encode('utf-8')
print(f"String: '{message}'")
print(f"Bytes (UTF-8): {bytes_message}")
# Example with a non-ASCII character
smiley = "😊"
bytes_smiley = smiley.encode('utf-8')
print(f"String: '{smiley}'")
print(f"Bytes (UTF-8): {bytes_smiley}")Endianness: Byte Order
When a piece of data, like a 32-bit integer, takes up more than one byte, there's a choice to be made: which byte comes first in memory? This order is called endianness.
- Big-endian: The most significant byte (MSB) comes first. Think of reading numbers left-to-right, like "123" where '1' is the most significant digit.
- Little-endian: The least significant byte (LSB) comes first. This is like writing "321" if '1' were the most significant.
Visualizing Endianness
Let's take the 32-bit hexadecimal number 0x12345678. This number has four bytes: 12, 34, 56, 78.
- Big-endian: Stores bytes in memory as
12 34 56 78(MSB first). - Little-endian: Stores bytes in memory as
78 56 34 12(LSB first).
Most modern Intel/AMD CPUs (x86/x64) are little-endian. Network protocols often use big-endian.
Endianness & Reverse Engineering
Understanding endianness is crucial when you're working with raw binary data, especially across different systems or file formats.
- If you read a 32-bit integer from a big-endian file on a little-endian system without conversion, the value will be incorrect.
- Network packets often use big-endian, so analyzing network traffic requires awareness.
- Many embedded systems (like ARM processors) can be configured for either, adding complexity.
Endianness Check
Imagine a 32-bit integer with the hexadecimal value 0xAABBCCDD is stored in memory. If the system is little-endian, what would be the order of bytes in memory, starting from the lowest address?
Data Representation Recap
Great job! In this lesson, we explored how data is represented in binaries:
- Integers: Stored as signed or unsigned numbers, with size determining range.
- Floating-points: Use standards like IEEE 754 for decimals.
- Characters & Strings: Mapped to numbers (ASCII, UTF-8) and often null-terminated.
- Endianness: The byte order (big-endian or little-endian) for multi-byte data, critical for correct interpretation.
Understanding these fundamentals is key to interpreting any binary file or memory dump. Next, we'll look at common binary file formats!
Perguntas Frequentes
A aula “Representação de Dados em Binários” é grátis?
Sim — o texto completo de “Representação de Dados em Binários” é grátis para ler aqui na web. Para praticá-la interativamente (um editor de código integrado e um tutor de IA 24/7) e desbloquear o restante do curso de Reverse Engineering & Binary Analysis Basics, atualize para CoddyKit PRO. O curso de Reverse Engineering & Binary Analysis Basics inclui 4 aulas no total.
O que vou aprender em “Representação de Dados em Binários”?
Aprenda como tipos de dados (inteiros, números de ponto flutuante e strings) são armazenados na memória e em arquivos, incluindo conceitos como endianidade. Você pratica Reverse Engineering & Binary Analysis Basics com código prático que executa diretamente no navegador, e um tutor de IA 24/7 responde suas dúvidas enquanto trabalha na aula.
Preciso ter experiência prévia para começar Reverse Engineering & Binary Analysis Basics?
Nenhuma experiência prévia é necessária. Reverse Engineering & Binary Analysis Basics no CoddyKit é estruturado para alunos iniciantes até avançados, então você pode começar aqui ou desde o início e aprender no seu ritmo. Esta é a aula 2 de 4.
Quanto tempo leva a aula “Representação de Dados em Binários”?
A maioria das aulas CoddyKit leva cerca de 5–10 minutos. Cada uma é compacta e interativa, então você faz progresso constante e retoma exatamente de onde parou entre web e app.
Posso escrever e executar código nesta aula de Reverse Engineering & Binary Analysis Basics?
Sim. Cada aula de Reverse Engineering & Binary Analysis Basics inclui um editor de código integrado, então você escreve e executa código real direto no navegador e recebe feedback de IA instantaneamente — nenhuma configuração local necessária.
Todas as aulas deste curso
- Visão Geral das Arquiteturas de CPU
- Representação de Dados em Binários
- Formatos Comuns de Arquivos Binários
- Endianness e Ordenação de Bytes