Vision-Language Models Qwen/Qwen2.5-VL-7B-Instruct Image-Text-to-Text • 8B • Updated Apr 6, 2025 • 7.98M • • 1.7k microsoft/Florence-2-large Image-Text-to-Text • 0.8B • Updated Aug 4, 2025 • 639k • 1.85k google/paligemma2-3b-pt-224 Image-Text-to-Text • 3B • Updated Dec 5, 2024 • 15.2k • 180
Aerial & Drone Object Detection PekingU/rtdetr_r50vd Object Detection • 43M • Updated Feb 6, 2025 • 191k • 36 silveroupti/VisDrone Viewer • Updated Apr 10 • 18.8k • 314 PekingU/rtdetr_r50vd_coco_o365 Object Detection • 43M • Updated Jul 1, 2024 • 187k • 17 IDEA-Research/grounding-dino-base Zero-Shot Object Detection • 0.2B • Updated May 12, 2024 • 1.65M • 206
IDEA-Research/grounding-dino-base Zero-Shot Object Detection • 0.2B • Updated May 12, 2024 • 1.65M • 206
Applied Machine Learning for Computer Vision and VLMs Qwen/Qwen3.5-9B Image-Text-to-Text • 10B • Updated Mar 2 • 12.1M • • 1.9k dx8152/Qwen-Edit-2509-Multiple-angles Image-to-Image • Updated Apr 21 • 160k • • 976 KangLiao/Puffin Text-to-3D • Updated Mar 6 • 24
OCR & Document AI nvidia/nemotron-ocr-v2 Image-to-Text • Updated May 22 • 1.32k • 251 deepseek-ai/DeepSeek-OCR Image-Text-to-Text • 3B • Updated Nov 4, 2025 • 2.46M • • 3.35k zai-org/GLM-OCR Image-Text-to-Text • 1B • Updated May 19 • 2.05M • • 2.02k
Theory and Representation learning V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning Paper • 2603.14482 • Published Mar 15 • 38 Real-Time Object Detection Meets DINOv3 Paper • 2509.20787 • Published Sep 25, 2025 • 11
V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning Paper • 2603.14482 • Published Mar 15 • 38
Vision-Language Models Qwen/Qwen2.5-VL-7B-Instruct Image-Text-to-Text • 8B • Updated Apr 6, 2025 • 7.98M • • 1.7k microsoft/Florence-2-large Image-Text-to-Text • 0.8B • Updated Aug 4, 2025 • 639k • 1.85k google/paligemma2-3b-pt-224 Image-Text-to-Text • 3B • Updated Dec 5, 2024 • 15.2k • 180
OCR & Document AI nvidia/nemotron-ocr-v2 Image-to-Text • Updated May 22 • 1.32k • 251 deepseek-ai/DeepSeek-OCR Image-Text-to-Text • 3B • Updated Nov 4, 2025 • 2.46M • • 3.35k zai-org/GLM-OCR Image-Text-to-Text • 1B • Updated May 19 • 2.05M • • 2.02k
Aerial & Drone Object Detection PekingU/rtdetr_r50vd Object Detection • 43M • Updated Feb 6, 2025 • 191k • 36 silveroupti/VisDrone Viewer • Updated Apr 10 • 18.8k • 314 PekingU/rtdetr_r50vd_coco_o365 Object Detection • 43M • Updated Jul 1, 2024 • 187k • 17 IDEA-Research/grounding-dino-base Zero-Shot Object Detection • 0.2B • Updated May 12, 2024 • 1.65M • 206
IDEA-Research/grounding-dino-base Zero-Shot Object Detection • 0.2B • Updated May 12, 2024 • 1.65M • 206
Theory and Representation learning V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning Paper • 2603.14482 • Published Mar 15 • 38 Real-Time Object Detection Meets DINOv3 Paper • 2509.20787 • Published Sep 25, 2025 • 11
V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning Paper • 2603.14482 • Published Mar 15 • 38
Applied Machine Learning for Computer Vision and VLMs Qwen/Qwen3.5-9B Image-Text-to-Text • 10B • Updated Mar 2 • 12.1M • • 1.9k dx8152/Qwen-Edit-2509-Multiple-angles Image-to-Image • Updated Apr 21 • 160k • • 976 KangLiao/Puffin Text-to-3D • Updated Mar 6 • 24