Locked at the Entrance, Open Inside: Where RLVR Narrows the Solution Space Paper • 2608.29188 • Published 9 days ago • 10 • 3
DRACO: Fine-Grained Credit Assignment with Dynamic Rubrics for Long-Horizon Agent Training Paper • 2609.04094 • Published 4 days ago • 25 • 5
Compile by Training: Turning Natural-Language Specifications into Local Neural Functions Paper • 2609.04199 • Published 4 days ago • 356 • 4
Ignorance or Incompetence? Constructing Knowledge-Gated, Verifiable Tasks for LLM Agents Paper • 2608.30322 • Published 7 days ago • 4 • 4
From Production Traffic to Post-Training: Building a Self-Hosted LLM That Covers the Corporate Request Mix Paper • 2609.01572 • Published 6 days ago • 33 • 3
Super Library Agent: Joint Generation and Maintenance of Multiple Applications Beyond the Single Codebase Paper • 2608.29310 • Published 9 days ago • 27 • 4
LoopArena: Benchmarking Models as Runtime Controllers for Loop Engineering Paper • 2608.28281 • Published 10 days ago • 101 • 5
What Does an Evaluation License? A Commit-Bound Census of Claim Replay in Inspect Evals Paper • 2608.19269 • Published 9 days ago • 5 • 3
Agentic Game Development as a Verifiable Trajectory Data Engine for Scaling World Models Paper • 2608.25518 • Published 12 days ago • 196 • 5
PILOT in the Loop: Live Self-Improvement for Long-Horizon Agents Paper • 2608.26530 • Published 11 days ago • 34 • 4
TorchMorph: CUDA-accelerated Morphological Transforms Paper • 2608.24738 • Published 13 days ago • 4 • 3
Quantization-Aware Healing: A Practical Recipe for Recovering Compressed, 4-Bit LLMs Paper • 2608.20953 • Published 17 days ago • 11 • 6
Daedalus-150M: A Convolution-Attention Hybrid Designed for CPU Inference Paper • 2608.20210 • Published 18 days ago • 8 • 4
QuoteBench: How Matched Scores Can Hide Command-Path Failures Paper • 2608.13547 • Published 25 days ago • 8 • 4
The Embedder's Dilemma: LLMs Are Better, but at What Cost? Paper • 2608.12875 • Published 25 days ago • 15 • 5
SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science? Paper • 2608.19799 • Published 18 days ago • 65 • 4
SkillGate: Training In-Policy Skill Selection in Long-Horizon Agents Paper • 2608.18852 • Published 19 days ago • 8 • 3
How Do Agents Fail on AutoResearch: End-to-End Diagnostic Evaluation on 100 Real-World Frontier Research Tasks Paper • 2608.14905 • Published 24 days ago • 31 • 4
Second Thought: Reasoning in Parallel as LLM Agents Act and Observe Paper • 2608.13667 • Published 25 days ago • 17 • 4
Specification-first convergence with an AI coding agent: a case study of dismantling a core architectural invariant across 189 files in a 717k-line codebase with no test oracle and no human code review Paper • 2608.12440 • Published 26 days ago • 10 • 5