CheckerBench: Can Long-Horizon Agents Synthesize Static-Analysis Checkers?
Abstract
Static-analysis checker synthesis requires agents to interpret a defect specification, inspect a repository, implement analyzer-specific logic, and refine the checker through repeated compilation and analysis feedback. Existing coding-agent benchmarks focus on tasks such as patch generation or vulnerability detection and rarely assess whether an agent can develop a working checker in a repository from start to finish. We introduce CheckerBench, an executable benchmark of 300 tasks derived from 297 CVEs across 167 repositories, 85 CWEs, and five language ecosystems. Each task includes vulnerable and fixed revisions, a pinned analysis environment, and a checker scaffold. We further introduce CheckerLab, a common evaluation framework that independently rebuilds submitted checkers and measures vulnerable-fixed diagnostic contrast, patch localization, false positives, and tool use. Across 21 model-harness configurations and three independent repeats per configuration, mean Pass@1 is 32.30%, while the best reaches 45.33%. These results show that reliable, reusable checker development remains challenging for current coding agents.
Community
We introduce CheckerBench, an executable benchmark for evaluating whether coding agents can turn real-world vulnerabilities into reusable static-analysis checkers.
CheckerBench contains 300 tasks derived from 297 CVEs across 167 repositories, 85 CWEs, and five language ecosystems. Agents must inspect repositories, implement analyzer-specific logic, and iteratively refine their checkers through compilation and analysis feedback. Each task includes vulnerable and fixed revisions, a pinned analysis environment, and a checker scaffold.
We also introduce CheckerLab, a unified evaluation framework that independently rebuilds submitted checkers and measures vulnerable–fixed diagnostic contrast, patch localization, false positives, and tool use.
Across 21 model–harness configurations and three independent repeats per configuration, mean Pass@1 is 32.30%, while the best configuration reaches 45.33%. These results reveal substantial room for progress in long-horizon agentic checker development.
The dataset and runtime assets are publicly available to support reproducible research on coding agents, automated static analysis, and software security.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Correct Code, Broken Contributions? SWE-CC: Benchmarking Repository Policy Compliance for Coding Agents (2026)
- OdinEval: A Reproducible Benchmark for LLM-Based Program Repair in the Odin Programming Language (2026)
- Evaluating Coding Agents on Kernel Exploit Generation (2026)
- WitnessGym: Benchmarking Coding Agents on the Construction of Bug Witnesses (2026)
- Test-Driven Approaches to Software Engineering with Large Language Models: A Survey of Phases, Tasks, and Agent Skills (2026)
- SWE-Bench Pro Verified: A Reliable Benchmark for Software Engineering Agents (2026)
- From Dead Code and Static Requirements to Working Engines: Software Revival with Coding Agents (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2610.07557 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 1
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper