EarlyEval: Cheaper Agent Evaluation via Early Outcome Prediction Paper • 2609.02783 • Published 5 days ago • 113
CyberFactory: Scaling Cyber Security Capabilities with Instances from the Wild Paper • 2608.23181 • Published 14 days ago • 34
Are Performance-Optimization Benchmarks Reliably Measuring Coding Agents? Paper • 2607.01211 • Published Jul 1 • 14
Second Thought: Reasoning in Parallel as LLM Agents Act and Observe Paper • 2608.13667 • Published 24 days ago • 17
How Do Practitioners Build SE Agents? Insights from a Mixed-Methods Study Paper • 2607.10856 • Published Jul 12 • 7
ResearchStudio-Reel: Automate the Last Mile of Research from Paper to Poster, Video, and Blog Paper • 2607.04438 • Published Jul 5 • 64
SecureAgentBench: Benchmarking Secure Code Generation under Realistic Vulnerability Scenarios Paper • 2509.22097 • Published Sep 26, 2025 • 3
Reasoning Runtime Behavior of a Program with LLM: How Far Are We? Paper • 2403.16437 • Published Mar 25, 2024 • 2
Running Agents 1.52k Big Code Models Leaderboard 📈 1.52k Explore and compare code model performance on a leaderboard