Complete Transformers 5.x multimodal and save support

#5
Tomoro AI Ltd org

Follow-up to #4 and the matching 8B loading fix.

This completes the remaining Transformers 5.x compatibility work:

  • Request multimodal token type IDs by default and delegate construction to ProcessorMixin, giving image tokens type 1 and video tokens type 2.
  • Preserve Transformers 5's tied-weight dictionary when nesting the Qwen3-VL model, fixing save_pretrained() and MultiVectorEncoder.save().

Validated from main commit be3da8989cc4fe7038ec7fffb3293e6d701c5bf9 with Transformers 5.15.0. The 4B and 8B model/processor Python files are byte-identical, so this applies the same tested patch. Validation used an executable downsized checkpoint with the same remote code:

  • process_images(): modality IDs [0, 1]
  • Video processing: modality IDs [0, 2]
  • Image and video forward passes
  • model.save_pretrained()
  • MultiVectorEncoder.save() with Sentence Transformers 6.0.0.dev0 at 626eb602b088878a8173bcab69225f68411b99e1
hxssgaa changed pull request status to merged

Sign up or log in to comment