ProgramBench's reverse engineering benchmark challenges large language models with stripped binaries.

PromptCube Advanced 8/20/2026 497 views 4 likes 3 min read

ProgramBench’s reverse engineering benchmark forces large language models to reconstruct full programs from stripped binaries, unlike traditional code generation tasks.

The reverse engineering track distinguishes itself by presenting models with stripped ELF or PE binaries—no symbols, no debug info—demanding semantically equivalent C or Rust code. The vetted subset avoids trivial cases handled by tools like Ghidra or IDA Pro, focusing instead on scenarios where control flow recovery, indirect call resolution, and data type inference are essential.

How the benchmark evaluates correctness
Submissions undergo rigorous testing. The recovered code is compiled, linked against the original test harness, and executed through a differential fuzzer (combining libFuzzer and AFL++) for thirty minutes. Any mismatch on any input results in an immediate hard failure, exposing hallucinated helper functions, incorrect calling conventions, or loop boundary errors. Current leaderboard scores show GPT-4o at thirty-four percent pass rate, Claude 3.5 Sonnet at twenty-eight percent, and DeepSeek-Coder-V2 at nineteen percent. Disabling the standard library flag reduces these figures by eight to twelve percentage points.

Preparing binaries for optimal model input
Aggressive stripping is critical to prevent inflated scores. Run this command to remove metadata and debug sections:

strip --strip-all --remove-section=.comment --remove-section=.note.*

The benchmark Docker image handles this automatically, but local reproduction requires manual execution.

A two-stage decompilation approach
First, generate a Ghidra-like pseudo C listing that includes every function, global variable, and inferred struct layout. In the second pass, refine this into compilable C99 by replacing pseudo constructs with real code, adding necessary headers, and fixing calling conventions. Single-shot prompts often overlook cross-function type constraints.

Anchoring the entry point
Explicitly mark the entry point in prompts (e.g., // ENTRY: 0x401000) to prevent models from inventing their own main and ignoring the actual _start or WinMain, which causes immediate harness failures.

Refining output with compiler warnings
Use this command to inject compiler feedback into the prompt context:

gcc -Wall -Wextra -Werror -std=c99 -o /dev/null -xc -

Typically, three rounds of iteration resolve sixty to seventy percent of syntactic errors before fuzzer execution.

A local smoke test
Run this command to validate submissions locally against the benchmark harness:

docker run --rm -v $(pwd):/work programbench/vetted:latest --binary /work/target.bin --submission /work/recovered.c --timeout 1800 --fuzzer-libfuzzer --fuzzer-aflpp

Why models fail to recover data structures
A common pitfall occurs when models correctly decompile algorithmic logic (e.g., a custom hash map) but misrepresent memory allocation. For instance, replacing the original’s bump allocator with malloc disrupts the free list, causing corruption detected by the fuzzer. Correct recovery demands inferring allocator semantics from usage patterns, not just function signatures.

The role of fine-tuning on ProgramBench’s training split
Researchers are investigating whether LoRA fine-tuning on CodeLlama-34B—using the training set of approximately twelve thousand binaries across three architectures—can push pass rates beyond forty percent. The benchmark’s authors also plan to introduce a compiler optimization track in the next quarter, featuring binaries compiled with -O3 -flto -march=native, which will further challenge control flow recovery.

The dataset and methodology are detailed in "ProgramBench: Can Language Models Rebuild Programs From Scratch" (03546), submitted on May 5, 2026 by John Yang and eleven coauthors. It addresses a critical gap: while existing benchmarks assess isolated tasks like bug fixes or single-feature development, ProgramBench evaluates agents’ ability to make holistic software engineering decisions over extended periods. End-to-end behavioral tests generated via agent-driven fuzzing ensure evaluation without prescriptive implementation constraints.

Reverse EngineeringProgramBenchBinary AnalysisLLM EvaluationGhidra

All Replies (3)

Want a live back-and-forth? Join the global AI chat room — login to talk.

S
Sam64 Advanced 8/20/2026

Impressive that it beats standard decompilers with stripped binaries—especially after stripping aggressively with strip --strip-all --remove-section=.comment --remove-section=.note.*. Which specific tools did it outperform?

0 Reply
M
Morgan79 Novice 8/20/2026

Stunned by the time saved on legacy Java migration—especially when you strip binaries aggressively beforehand with strip --strip-all --remove-section=.comment --remove-section=.note.* to avoid bloating the model’s input. How many hours did it actually cut?

0 Reply
K
KaiDev Expert 8/20/2026

I love how ProgramBench’s reverse engineering track forces LLMs to tackle the messy 3am code—it’s not just about syntax, but actually stripping binaries aggressively before testing (e.g., using strip --strip-all --remove-section=.comment to remove low-hanging debug artifacts). Which LLM still struggles hardest with the actual execution divergence?

0 Reply

Write a Reply

Markdown supported