Robust-U1: Can MLLMs Self-Recover Corrupted Visual Content for Robust Understanding?

Robust-U1 is a unified Multimodal Large Language Model (MLLM) that can self-recover corrupted visual content and perform multimodal reasoning over it. This enables robust visual understanding even under real-world image degradations.

🏰 Pretrained Checkpoints

Checkpoint Link Note
BAGEL-7B-MoT ByteDance-Seed/BAGEL-7B-MoT Used as initial weights for training.
Robust-U1 Jiaqi-hkust/Robust-U1 Final model for visual self-recovery and multimodal reasoning.
Robust-U1-RL Jiaqi-hkust/Robust-U1-RL Fine-tuned with reinforcement learning.
Robust-U1-SFT Jiaqi-hkust/Robust-U1-SFT Fine-tuned with supervised learning.

⭐️ Citation

If you find this repository useful, please cite the paper:

@inproceedings{tang2026robustu1,
      title={Robust-U1: Can MLLMs Self-Recover Corrupted Visual Content for Robust Understanding?},
      author={Tang, Jiaqi and Chen, Jianmin and Zhai, Youyang and Wei, Wei and Liu, Runtao and Zhao, Mengjie and Wu, Xiangyu and Xiao, Qingfa and Chen, Qifeng},
      booktitle={Proceedings of the 43rd International Conference on Machine Learning (ICML)},
      year={2026},
}
Downloads last month
15
Safetensors
Model size
15B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Space using Jiaqi-hkust/Robust-U1 1

Collection including Jiaqi-hkust/Robust-U1

Paper for Jiaqi-hkust/Robust-U1