PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models Paper • 2607.24957 • Published 9 days ago • 18
Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents Paper • 2607.28227 • Published 6 days ago • 299
Streaming Multi-Agent Autoregressive Diffusion Model with World State Registers Paper • 2607.21594 • Published 13 days ago • 16
Multimodal Continuous Reasoning via Asymmetric Mutual Variational Learning Paper • 2607.00461 • Published Jul 1 • 28
DataClaw0: Agentic Tailoring Multimodal Data from Raw Streams Paper • 2606.21337 • Published Jun 19 • 75
LoSoNA: A Benchmark for Local Social Norm Adaptation in Group Conversations Paper • 2606.14600 • Published Jun 12 • 7
Robust-U1: Can MLLMs Self-Recover Corrupted Visual Content for Robust Understanding? Paper • 2606.08063 • Published Jun 6 • 82