二进制文件中的数据表示
学习整数、浮点数和字符串等数据类型如何存储在内存和文件中,并了解字节序等概念。
二进制文件中的数据表示 是 CoddyKit 上的免费 Reverse Engineering & Binary Analysis Basics 课时。 这是第 2 节课,共 4 节。 你可以在下方免费阅读本课时的完整内容 — 然后在浏览器中使用内置代码编辑器和全天候 AI 导师进行实践。 这是 Reverse Engineering & Binary Analysis Basics 学习路径的一部分,你的进度在网页和 CoddyKit 应用中同步。 Reverse Engineering & Binary Analysis Basics 课程共包含 4 节课。
本课时的部分内容尚未翻译,以英文显示。
What is Binary Data?
When you reverse engineer, you're looking at a program's raw binary form. This means understanding how all kinds of data—numbers, text, and more—are stored as sequences of bits and bytes.
A bit is the smallest unit, either 0 or 1. Eight bits make a byte. Everything a computer does, from calculations to displaying text, relies on these fundamental units.
Numbers as Bits: Integers
Integers are whole numbers. They can be signed (positive or negative) or unsigned (only non-negative). The number of bytes used determines the range of values an integer can hold.
- Byte (8-bit): 0 to 255 (unsigned) or -128 to 127 (signed).
- Word (16-bit): Up to 65,535 (unsigned).
- DWord (32-bit): Up to ~4 billion (unsigned).
- QWord (64-bit): Much larger numbers!
Integer Size Matters
Let's see how different integer sizes affect the maximum value. This Python code shows the max value for an unsigned 8-bit integer (a byte) and a signed 8-bit integer.
# Max unsigned 8-bit integer
max_8_bit_unsigned = 2**8 - 1
print(f"Max 8-bit unsigned: {max_8_bit_unsigned}")
# Max signed 8-bit integer
max_8_bit_signed = 2**7 - 1
min_8_bit_signed = -2**7
print(f"Max 8-bit signed: {max_8_bit_signed}")
print(f"Min 8-bit signed: {min_8_bit_signed}")Floating-Point Numbers (Floats)
Numbers with decimal points, like 3.14 or -0.5, are called floating-point numbers. They are stored differently from integers to handle their fractional parts.
Most systems use the IEEE 754 standard for floats. This standard defines how a number's sign, exponent, and fractional part are represented in bits. Common sizes are 32-bit (single-precision) and 64-bit (double-precision).
Text: Characters & Strings
Text characters are also stored as numbers. The most common mapping for English characters is ASCII, where each character (like 'A' or '!') corresponds to a specific 8-bit number.
For a wider range of characters (emojis, foreign languages), Unicode is used. UTF-8 is a popular Unicode encoding that uses 1 to 4 bytes per character, making it flexible and backward-compatible with ASCII.
A string is simply a sequence of these characters, often ending with a special null byte (0x00) to mark its end.
How Strings Become Bytes
Here's how a simple string is represented as bytes using UTF-8. Notice how each character gets a numerical value.
message = "Hello"
bytes_message = message.encode('utf-8')
print(f"String: '{message}'")
print(f"Bytes (UTF-8): {bytes_message}")
# Example with a non-ASCII character
smiley = "😊"
bytes_smiley = smiley.encode('utf-8')
print(f"String: '{smiley}'")
print(f"Bytes (UTF-8): {bytes_smiley}")Endianness: Byte Order
When a piece of data, like a 32-bit integer, takes up more than one byte, there's a choice to be made: which byte comes first in memory? This order is called endianness.
- Big-endian: The most significant byte (MSB) comes first. Think of reading numbers left-to-right, like "123" where '1' is the most significant digit.
- Little-endian: The least significant byte (LSB) comes first. This is like writing "321" if '1' were the most significant.
Visualizing Endianness
Let's take the 32-bit hexadecimal number 0x12345678. This number has four bytes: 12, 34, 56, 78.
- Big-endian: Stores bytes in memory as
12 34 56 78(MSB first). - Little-endian: Stores bytes in memory as
78 56 34 12(LSB first).
Most modern Intel/AMD CPUs (x86/x64) are little-endian. Network protocols often use big-endian.
Endianness & Reverse Engineering
Understanding endianness is crucial when you're working with raw binary data, especially across different systems or file formats.
- If you read a 32-bit integer from a big-endian file on a little-endian system without conversion, the value will be incorrect.
- Network packets often use big-endian, so analyzing network traffic requires awareness.
- Many embedded systems (like ARM processors) can be configured for either, adding complexity.
Endianness Check
Imagine a 32-bit integer with the hexadecimal value 0xAABBCCDD is stored in memory. If the system is little-endian, what would be the order of bytes in memory, starting from the lowest address?
Data Representation Recap
Great job! In this lesson, we explored how data is represented in binaries:
- Integers: Stored as signed or unsigned numbers, with size determining range.
- Floating-points: Use standards like IEEE 754 for decimals.
- Characters & Strings: Mapped to numbers (ASCII, UTF-8) and often null-terminated.
- Endianness: The byte order (big-endian or little-endian) for multi-byte data, critical for correct interpretation.
Understanding these fundamentals is key to interpreting any binary file or memory dump. Next, we'll look at common binary file formats!
常见问题解答
「二进制文件中的数据表示」课时是免费的吗?
是的 — 「二进制文件中的数据表示」的完整文本可在网页上免费阅读。要进行交互式练习(内置代码编辑器和全天候 AI 导师)并解锁 Reverse Engineering & Binary Analysis Basics 课程的其余内容,请升级到 CoddyKit PRO。 Reverse Engineering & Binary Analysis Basics 课程共包含 4 节课。
「二进制文件中的数据表示」这节课中我会学到什么?
学习整数、浮点数和字符串等数据类型如何存储在内存和文件中,并了解字节序等概念。 你通过在浏览器中直接运行的动手代码来练习 Reverse Engineering & Binary Analysis Basics,全天候 AI 导师会在你学习这节课的过程中回答你的问题。
学习 Reverse Engineering & Binary Analysis Basics 需要有经验吗?
无需任何先前经验。CoddyKit 上的 Reverse Engineering & Binary Analysis Basics 课程适合初学者到高级学习者,你可以从这里开始或从头开始,按照自己的节奏学习。 这是第 2 节课,共 4 节。
「二进制文件中的数据表示」课时需要多长时间?
大多数 CoddyKit 课程大约需要 5–10 分钟。每节课都很精短且互动,所以你能稳步进步,并在网页和应用中从离开的地方继续。
我能在这节 Reverse Engineering & Binary Analysis Basics 课中编写并运行代码吗?
能。每节 Reverse Engineering & Binary Analysis Basics 课都包含内置代码编辑器,你可以在浏览器中直接编写并运行真实代码,并获得即时 AI 反馈 — 无需本地设置。