ServImage: An Image Generation and Editing Benchmark from Real-world Commercial Imaging Services Paper • 2604.24023 • Published Apr 27
The Cylindrical Representation Hypothesis for Language Model Steering Paper • 2605.01844 • Published May 3 • 3
M3MAD-Bench: Are Multi-Agent Debates Really Effective Across Domains and Modalities? Paper • 2601.02854 • Published Jan 6
Beyond Survival: Evaluating LLMs in Social Deduction Games with Human-Aligned Strategies Paper • 2510.11389 • Published Oct 13, 2025
Do LLMs "Feel"? Emotion Circuits Discovery and Control Paper • 2510.11328 • Published Oct 13, 2025 • 8
Audio Jailbreak: An Open Comprehensive Benchmark for Jailbreaking Large Audio-Language Models Paper • 2505.15406 • Published May 21, 2025 • 5
Evaluate Bias without Manual Test Sets: A Concept Representation Perspective for LLMs Paper • 2505.15524 • Published May 21, 2025 • 8
Word Form Matters: LLMs' Semantic Reconstruction under Typoglycemia Paper • 2503.01714 • Published Mar 3, 2025 • 5
MedTrinity-25M: A Large-scale Multimodal Dataset with Multigranular Annotations for Medicine Paper • 2408.02900 • Published Aug 6, 2024 • 31
The Cylindrical Representation Hypothesis for Language Model Steering Paper • 2605.01844 • Published May 3 • 3
PAN: A World Model for General, Interactable, and Long-Horizon World Simulation Paper • 2511.09057 • Published Nov 12, 2025 • 82