Visual Document Retrieval
Transformers
Safetensors
sentence-transformers
ColPali
multilingual
colqwen3
feature-extraction
multi-vector
text
image
video
multimodal-embedding
vidore
multilingual-embedding
custom_code
Instructions to use TomoroAI/tomoro-colqwen3-embed-4b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use TomoroAI/tomoro-colqwen3-embed-4b with Transformers:
# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("TomoroAI/tomoro-colqwen3-embed-4b", trust_remote_code=True, device_map="auto") - sentence-transformers
How to use TomoroAI/tomoro-colqwen3-embed-4b with sentence-transformers:
from sentence_transformers import SentenceTransformer model = SentenceTransformer("TomoroAI/tomoro-colqwen3-embed-4b", trust_remote_code=True) sentences = [ "The weather is lovely today.", "It's so sunny outside!", "He drove to the stadium." ] embeddings = model.encode(sentences) similarities = model.similarity(embeddings, embeddings) print(similarities.shape) # [3, 3] - ColPali
How to use TomoroAI/tomoro-colqwen3-embed-4b with ColPali:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
Complete Transformers 5.x multimodal and save support
#5
by hxssgaa - opened
Follow-up to #4 and the matching 8B loading fix.
This completes the remaining Transformers 5.x compatibility work:
- Request multimodal token type IDs by default and delegate construction to
ProcessorMixin, giving image tokens type 1 and video tokens type 2. - Preserve Transformers 5's tied-weight dictionary when nesting the Qwen3-VL model, fixing
save_pretrained()andMultiVectorEncoder.save().
Validated from main commit be3da8989cc4fe7038ec7fffb3293e6d701c5bf9 with Transformers 5.15.0. The 4B and 8B model/processor Python files are byte-identical, so this applies the same tested patch. Validation used an executable downsized checkpoint with the same remote code:
process_images(): modality IDs[0, 1]- Video processing: modality IDs
[0, 2] - Image and video forward passes
model.save_pretrained()MultiVectorEncoder.save()with Sentence Transformers 6.0.0.dev0 at626eb602b088878a8173bcab69225f68411b99e1
hxssgaa changed pull request status to merged