MiniMax-M2.5

No longer available on HF due to storage restrictions - archived here

See MiniMax-M2.5 in action: demonstration video

Tested on a M3 Ultra 512GB RAM using Inferencer app v1.10

  • Single inference ~36.5 tokens/s @ 1000 tokens
  • Batched inference ~44 total tokens/s across two inferences
  • Memory usage: ~239 GiB
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for inferencerlabs/MiniMax-M2.5-MLX-Q9

Finetuned
(27)
this model