CyberFactory: Scaling Cyber Security Capabilities with Instances from the Wild Paper • 2608.23181 • Published 5 days ago • 32
ParaTempo: Efficient Parallel Reasoning via Temporal Confidence Paper • 2608.16425 • Published 12 days ago • 38
SkillForge: Self-Distilling Agents for Project-Specific Issue Resolution Paper • 2608.18933 • Published 10 days ago • 11
Second Thought: Reasoning in Parallel as LLM Agents Act and Observe Paper • 2608.13667 • Published 15 days ago • 16
SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring Paper • 2608.09802 • Published 19 days ago • 133
Are Performance-Optimization Benchmarks Reliably Measuring Coding Agents? Paper • 2607.01211 • Published Jul 1 • 14
How Do Practitioners Build SE Agents? Insights from a Mixed-Methods Study Paper • 2607.10856 • Published Jul 12 • 7
Dockerless: Environment-Free Program Verifier for Coding Agents Paper • 2606.28436 • Published Jun 26 • 116
FastContext: Training Efficient Repository Explorer for Coding Agents Paper • 2606.14066 • Published Jun 12 • 96
SWE-Explore: Benchmarking How Coding Agents Explore Repositories Paper • 2606.07297 • Published Jun 5 • 123
Rethinking the Value of Agent-Generated Tests for LLM-Based Software Engineering Agents Paper • 2602.07900 • Published Feb 8 • 4
DLLM-Searcher: Adapting Diffusion Large Language Model for Search Agents Paper • 2602.07035 • Published Feb 3 • 31
CodeOCR: On the Effectiveness of Vision Language Models in Code Understanding Paper • 2602.01785 • Published Feb 2 • 97
PaperBanana: Automating Academic Illustration for AI Scientists Paper • 2601.23265 • Published Jan 30 • 230
SWE-Pruner: Self-Adaptive Context Pruning for Coding Agents Paper • 2601.16746 • Published Jan 23 • 94