TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Paper • 2607.17423 • Published 3 days ago • 146
LATO.2: Factorized 3D Mesh Generation with Vertex and Topology Flow Paper • 2607.10623 • Published 10 days ago • 14
AutoRubric-T2I: Robust Rule-Based Reward Model for Text-to-Image Alignment Paper • 2605.17602 • Published May 20 • 19
Mean Mode Screaming: Mean--Variance Split Residuals for 1000-Layer Diffusion Transformers Paper • 2605.06169 • Published May 7 • 238
Maestro: Reinforcement Learning to Orchestrate Hierarchical Model-Skill Ensembles Paper • 2605.22177 • Published May 21 • 21
Video2GUI: Synthesizing Large-Scale Interaction Trajectories for Generalized GUI Agent Pretraining Paper • 2605.14747 • Published May 14 • 147
DiffRetriever: Parallel Representative Tokens for Retrieval with Diffusion Language Models Paper • 2605.07210 • Published May 8 • 4