Keep-or-Drop? Adaptive Tokenizer for Compact Video Representation Paper • 2608.24293 • Published 8 days ago • 9
FactorJEPA: Factorizing Monolithic Futures into Layout-Agent-Interaction Channels for Crowded and Chaotic Global South Urban Worlds Paper • 2608.01049 • Published about 1 month ago • 13
LeapTalk: Breaking the Latency-Quality Trade-off in Talking Head Generation Paper • 2608.00079 • Published Jul 29 • 18
P-MTP: Efficient Document Parsing via Multi-Token Prediction with Progressive Depth Scaling Paper • 2606.24447 • Published Jun 23 • 1
AgenticSTS: A Bounded-Memory Testbed for Long-Horizon LLM Agents Paper • 2607.02255 • Published Jul 2 • 69
Stable-Layers: Fine-Tuning Image Layer Decomposition Models with VLM-Scored Reinforcement Learning Paper • 2605.30257 • Published May 28 • 7
αDepth: Learning Single-Pass Soft Boundary Decomposition for Stereo Conversion Paper • 2606.00386 • Published May 29 • 8
PaddleOCR-VL-1.6: Expanding the Frontier of Document Parsing with Under-Optimized Region Refinement and Progressive Post-Training Paper • 2606.03264 • Published Jun 2 • 28
GrepSeek: Training Search Agents for Direct Corpus Interaction Paper • 2605.29307 • Published May 28 • 118
Multi-view Consistent 3D Gaussian Head Avatars 'without' Multi-view Generation Paper • 2605.25220 • Published May 24 • 9
LocateAnything: Fast and High-Quality Vision-Language Grounding with Parallel Box Decoding Paper • 2605.27365 • Published May 26 • 147