Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
silveroupti
's Collections
Vision-Language Models
OCR & Document AI
Aerial & Drone Object Detection
Theory and Representation learning
Applied Machine Learning for Computer Vision and VLMs
Vision-Language Models
updated
Apr 10
Upvote
-
Sort: Collection
Qwen/Qwen2.5-VL-7B-Instruct
Image-Text-to-Text
•
8B
•
Updated
Apr 6, 2025
•
9.35M
•
•
1.65k
microsoft/Florence-2-large
Image-Text-to-Text
•
0.8B
•
Updated
Aug 4, 2025
•
786k
•
1.84k
google/paligemma2-3b-pt-224
Image-Text-to-Text
•
3B
•
Updated
Dec 5, 2024
•
12.6k
•
175
Upvote
-
Sort: Collection
Share collection
View history
Collection guide
Browse collections