JarvisHub: An Open Harness for Canvas-Native Multimodal Creative Agents Paper • 2607.23588 • Published 6 days ago • 120
From Player to Master: Enhancing Test-Time Learning of LLM Agents via Reinforcement Learning over Memory Paper • 2606.08656 • Published Jun 7 • 3
VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation Paper • 2605.16079 • Published May 15 • 29
Towards Efficient Large Language Reasoning Models via Extreme-Ratio Chain-of-Thought Compression Paper • 2602.08324 • Published May 15
VimRAG: Navigating Massive Visual Context in Retrieval-Augmented Generation via Multimodal Memory Graph Paper • 2602.12735 • Published Feb 13 • 8
GroupGPT: A Token-efficient and Privacy-preserving Agentic Framework for Multi-User Chat Assistant Paper • 2603.01059 • Published Mar 1 • 1
Unify-Agent: A Unified Multimodal Agent for World-Grounded Image Synthesis Paper • 2603.29620 • Published Mar 31 • 49
SkillFlow:Benchmarking Lifelong Skill Discovery and Evolution for Autonomous Agents Paper • 2604.17308 • Published Apr 19 • 23
OpenSearch-VL: An Open Recipe for Frontier Multimodal Search Agents Paper • 2605.05185 • Published May 6 • 106
SCOPE: Structured Decomposition and Conditional Skill Orchestration for Complex Image Generation Paper • 2605.08043 • Published May 8 • 10
Vision-DeepResearch: Incentivizing DeepResearch Capability in Multimodal Large Language Models Paper • 2601.22060 • Published Jan 29 • 155
Vision-DeepResearch Benchmark: Rethinking Visual and Textual Search for Multimodal Large Language Models Paper • 2602.02185 • Published Feb 2 • 118
UniCorn: Towards Self-Improving Unified Multimodal Models through Self-Generated Supervision Paper • 2601.03193 • Published Jan 6 • 51
Actial: Activate Spatial Reasoning Ability of Multimodal Large Language Models Paper • 2511.01618 • Published Nov 3, 2025 • 11
CompBench: Benchmarking Complex Instruction-guided Image Editing Paper • 2505.12200 • Published May 18, 2025
Learning Only with Images: Visual Reinforcement Learning with Reasoning, Rendering, and Visual Feedback Paper • 2507.20766 • Published Jul 28, 2025 • 1
IWR-Bench: Can LVLMs reconstruct interactive webpage from a user interaction video? Paper • 2509.24709 • Published Sep 29, 2025 • 7
Agentic Jigsaw Interaction Learning for Enhancing Visual Perception and Reasoning in Vision-Language Models Paper • 2510.01304 • Published Oct 1, 2025 • 11