ZimaBlue: Evolving Generalizable World Action Models through Scalable Video Pre-training Paper • 2609.00188 • Published 18 days ago • 53
Beyond Data Scaling: Representation-Centric Continued Pre-training for Vision-Language-Action Models Paper • 2608.27550 • Published 22 days ago • 95
VLAct Collection Beyond Data Scaling: Representation-Centric Continued Pre-training for Vision-Language-Action Models • 12 items • Updated 18 days ago • 3
StreamOPD: A Post-Training Recipe with Spatio-Temporal Cue Gating for Streaming Video Understanding Paper • 2608.16320 • Published Aug 17 • 10
Mage Collection A family of lightweight multimodal models, including understanding and generation. • 8 items • Updated Jul 26 • 30
StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding Paper • 2608.05703 • Published Aug 6 • 17
Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model Paper • 2607.24904 • Published Jul 27 • 38
SciForma: Structure-Faithful Generation of Scientific Diagrams Paper • 2607.18091 • Published Jul 20 • 25
Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing Paper • 2607.19064 • Published Jul 21 • 78
ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning Paper • 2605.20342 • Published May 19 • 34
LongLive-2.0: An NVFP4 Parallel Infrastructure for Long Video Generation Paper • 2605.18739 • Published May 18 • 117
FlowAnchor: Stabilizing the Editing Signal for Inversion-Free Video Editing Paper • 2604.22586 • Published Apr 24 • 16
Contexts are Never Long Enough: Structured Reasoning for Scalable Question Answering over Long Document Sets Paper • 2604.22294 • Published Apr 24 • 18
LLM Safety From Within: Detecting Harmful Content with Internal Representations Paper • 2604.18519 • Published Apr 20 • 26
DiffNR: Diffusion-Enhanced Neural Representation Optimization for Sparse-View 3D Tomographic Reconstruction Paper • 2604.21518 • Published Apr 23 • 28
Agentic World Modeling: Foundations, Capabilities, Laws, and Beyond Paper • 2604.22748 • Published Apr 24 • 234
VP-VLA: Visual Prompting as an Interface for Vision-Language-Action Models Paper • 2603.22003 • Published Mar 23 • 12
Attend Before Attention: Efficient and Scalable Video Understanding via Autoregressive Gazing Paper • 2603.12254 • Published Mar 12 • 23
Rethinking Token-Level Policy Optimization for Multimodal Chain-of-Thought Paper • 2603.22847 • Published Mar 24 • 26