Identifying Functions and Data
Learn techniques to locate significant functions, strings, and other data within disassembled binaries.
Identifying Functions and Data is a free Reverse Engineering & Binary Analysis Basics lesson on CoddyKit — lesson 2 of 4. You can read the complete lesson below for free — then practise it hands-on in the browser with a built-in code editor and a 24/7 AI tutor. It is part of the Reverse Engineering & Binary Analysis Basics learning path, one of 4 lessons in the course, and your progress syncs across the web and the CoddyKit app.
Spotting Key Parts of a Binary
Welcome! In reverse engineering, our goal is to understand how a program works without its source code. A critical first step is to identify its core components: functions and data.
These elements are like the building blocks and raw materials of any software. Learning to spot them quickly will significantly speed up your analysis.
Strings: Your First Clues
Strings are often the easiest and most valuable clues in a binary. They can reveal a program's purpose, error messages, user prompts, file paths, network addresses, or API calls.
- Error messages:
"Error: File not found" - URLs/Paths:
"https://malicious.com/update","C:\Windows\System32\config.dat" - User Prompts:
"Enter password:"
Finding them is usually the first step for any analyst.
Locating Strings in Disassemblers
Most disassemblers, like Ghidra or IDA Pro, have a dedicated feature to list all identified strings within a binary. This saves you from manually scanning through raw bytes.
When you find an interesting string, you can usually cross-reference it to see where in the code it's being used. This immediately points you to relevant functions.
Functions: Program's Building Blocks
A function (or subroutine) is a self-contained block of code designed to perform a specific task. Programs are built from many functions calling each other.
Identifying functions helps you break down a complex program into smaller, manageable pieces, making it easier to understand its overall logic and flow.
Recognizing Function Entry Points
Functions often start with a specific sequence of instructions called a prologue. This setup typically prepares the stack for local variables and saves the previous stack frame.
A common x86 prologue looks like this:
push ebpmov ebp, esp
This sequence pushes the old base pointer onto the stack and sets the current stack pointer as the new base pointer.
Function Exits: Epilogues
Just as functions have entry points, they also have exit points, marked by an epilogue. The epilogue restores the stack to its state before the function call and returns control to the caller.
A typical x86 epilogue might be:
mov esp, ebppop ebpret
This restores the stack pointer, pops the old base pointer, and returns from the function.
Spotting Common Library Functions
Most programs use functions from system libraries (e.g., for printing to screen, file I/O, network communication). Disassemblers are often smart enough to identify these for you.
They do this by looking at imported symbols (like the Import Address Table in Windows PE files or Procedure Linkage Table in Linux ELF files) or by matching known function signatures.
Where Data Resides: Data Sections
Beyond code, binaries contain various data sections. Understanding these helps you locate global variables, constants, and other program-wide information:
.data: Initialized global and static variables..bss: Uninitialized global and static variables (zeroed out at runtime)..rdata: Read-only data, such as strings and constants.
These sections are usually clearly labeled in disassemblers.
Global vs. Local Variables
Distinguishing between global and local variables is key. Global variables are accessible throughout the program and are usually stored in .data or .bss sections.
Local variables, on the other hand, are created on the stack when a function is called and are only accessible within that function. They are typically referenced relative to the stack frame pointer (e.g., [ebp-0x4]).
Quick Check: Data Clues
You are analyzing a binary and see a reference to an address within the .rdata section. What kind of data is most likely stored at this address?
Key Takeaways
You've learned fundamental techniques for static analysis!
- Strings offer immediate insights into program functionality.
- Function prologues and epilogues help define code boundaries.
- Recognizing library functions speeds up analysis.
- Understanding data sections (
.data,.bss,.rdata) helps locate global variables and constants.
These skills are essential for navigating and understanding disassembled binaries.
Frequently asked questions
Is the “Identifying Functions and Data” lesson free?
Yes — the full text of “Identifying Functions and Data” is free to read here on the web, and the Reverse Engineering & Binary Analysis Basics course includes 4 lessons in total. To practise it interactively (a built-in code editor and a 24/7 AI tutor) and unlock the rest of the Reverse Engineering & Binary Analysis Basics course, upgrade to CoddyKit PRO.
What will I learn in “Identifying Functions and Data”?
Learn techniques to locate significant functions, strings, and other data within disassembled binaries. You practise Reverse Engineering & Binary Analysis Basics with hands-on code you run directly in the browser, and a 24/7 AI tutor answers your questions as you work through the lesson.
Do I need any experience to start Reverse Engineering & Binary Analysis Basics?
No prior experience is required. Reverse Engineering & Binary Analysis Basics on CoddyKit is structured for beginners through advanced learners; this is — lesson 2 of 4, so you can start here or from the beginning and move at your own pace.
How long does the “Identifying Functions and Data” lesson take?
Most CoddyKit lessons take about 5–10 minutes. Each one is bite-sized and interactive, so you make steady progress and pick up exactly where you left off across the web and the app.
Can I write and run code in this Reverse Engineering & Binary Analysis Basics lesson?
Yes. Every Reverse Engineering & Binary Analysis Basics lesson includes a built-in code editor, so you write and run real code right in your browser and get instant AI feedback — no local setup required.
All lessons in this course
- Introduction to Disassemblers
- Identifying Functions and Data
- Control Flow Graph Analysis
- String & Cross-Reference Analysis