Ex-Omni-2D: Expressive Omni-Modal Dialogue Models with Native Visual Presence Paper • 2608.10720 • Published Aug 11 • 16
Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents Paper • 2607.28227 • Published Jul 30 • 312
CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization Paper • 2607.25659 • Published Jul 28 • 85
TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM Paper • 2607.27205 • Published Jul 29 • 141
Self-Supervised Learning of Structured Dynamics from Videos Paper • 2607.21576 • Published Jul 23 • 21
ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU Paper • 2607.19191 • Published Jul 21 • 314
ReViV: Reconstructing the Viewer and the View in 4D from Monocular Egocentric Video Paper • 2607.17790 • Published Jul 20 • 6
HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enchancement Paper • 2607.18217 • Published Jul 20 • 63
ABot-N1: Toward a General Visual Language Navigation Foundation Model Paper • 2607.10383 • Published Jul 14 • 103
LATO.2: Factorized 3D Mesh Generation with Vertex and Topology Flow Paper • 2607.10623 • Published Jul 12 • 17
MemSlides: A Hierarchical Memory Driven Agent Framework for Personalized Slide Generation with Multi-turn Local Revision Paper • 2606.17162 • Published Jun 15 • 179
Skill-MAS: Evolving Meta-Skill for Automatic Multi-Agent Systems Paper • 2606.18837 • Published Jun 17 • 59
Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models Paper • 2606.11324 • Published Jun 9 • 173
Learning from the Self-future: On-policy Self-distillation for dLLMs Paper • 2606.18195 • Published Jun 16 • 77
WeaveBench: A Long-Horizon, Real-World Benchmark for Computer-Use Agents with Hybrid Interfaces Paper • 2606.09426 • Published Jun 8 • 108
From Correctness to Utility: Gain-Based Prefix Evaluation for LLM Reasoning Paper • 2606.07190 • Published Jun 5 • 35
Toward Generalist Autonomous Research via Hypothesis-Tree Refinement Paper • 2606.11926 • Published Jun 10 • 131
Small RL Controller, Large Language Model: RL-Guided Adaptive Sampling for Test-Time Scaling Paper • 2606.03102 • Published Jun 2 • 15