DRACO: Fine-Grained Credit Assignment with Dynamic Rubrics for Long-Horizon Agent Training Paper • 2609.04094 • Published 3 days ago • 23 • 5
Compile by Training: Turning Natural-Language Specifications into Local Neural Functions Paper • 2609.04199 • Published 3 days ago • 314 • 3
Ignorance or Incompetence? Constructing Knowledge-Gated, Verifiable Tasks for LLM Agents Paper • 2608.30322 • Published 6 days ago • 4 • 4
From Production Traffic to Post-Training: Building a Self-Hosted LLM That Covers the Corporate Request Mix Paper • 2609.01572 • Published 5 days ago • 33 • 3
Super Library Agent: Joint Generation and Maintenance of Multiple Applications Beyond the Single Codebase Paper • 2608.29310 • Published 8 days ago • 27 • 4
LoopArena: Benchmarking Models as Runtime Controllers for Loop Engineering Paper • 2608.28281 • Published 9 days ago • 101 • 5
What Does an Evaluation License? A Commit-Bound Census of Claim Replay in Inspect Evals Paper • 2608.19269 • Published 8 days ago • 5 • 3
Agentic Game Development as a Verifiable Trajectory Data Engine for Scaling World Models Paper • 2608.25518 • Published 11 days ago • 196 • 5
PILOT in the Loop: Live Self-Improvement for Long-Horizon Agents Paper • 2608.26530 • Published 10 days ago • 33 • 4
TorchMorph: CUDA-accelerated Morphological Transforms Paper • 2608.24738 • Published 12 days ago • 4 • 3
Quantization-Aware Healing: A Practical Recipe for Recovering Compressed, 4-Bit LLMs Paper • 2608.20953 • Published 16 days ago • 11 • 4
Daedalus-150M: A Convolution-Attention Hybrid Designed for CPU Inference Paper • 2608.20210 • Published 17 days ago • 8 • 4
QuoteBench: How Matched Scores Can Hide Command-Path Failures Paper • 2608.13547 • Published 24 days ago • 8 • 4
The Embedder's Dilemma: LLMs Are Better, but at What Cost? Paper • 2608.12875 • Published 24 days ago • 15 • 5
SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science? Paper • 2608.19799 • Published 17 days ago • 65 • 4
SkillGate: Training In-Policy Skill Selection in Long-Horizon Agents Paper • 2608.18852 • Published 18 days ago • 8 • 3
How Do Agents Fail on AutoResearch: End-to-End Diagnostic Evaluation on 100 Real-World Frontier Research Tasks Paper • 2608.14905 • Published 23 days ago • 31 • 4
Second Thought: Reasoning in Parallel as LLM Agents Act and Observe Paper • 2608.13667 • Published 24 days ago • 17 • 4
Specification-first convergence with an AI coding agent: a case study of dismantling a core architectural invariant across 189 files in a 717k-line codebase with no test oracle and no human code review Paper • 2608.12440 • Published 25 days ago • 10 • 5
Ready Cohorts: Bounding GPU Opportunity and Avoiding Host Round Trips in LLM-Agent Control Paper • 2608.12123 • Published 25 days ago • 2 • 3