3-bit VLM base + 4-bit MTP drafter. Pair both for accelerated instruct decode. reasoning_effort baked to low.
Jorge Leon
leonsarmiento
AI & ML interests
AI for Environment
Recent Activity
updated a model about 20 hours ago
leonsarmiento/Qwen3.8-27B-3bit-mtp-mlx updated a collection about 20 hours ago
Qwen3.8-27B MLX Quantizations published a model about 20 hours ago
leonsarmiento/Qwen3.8-27B-3bit-mtp-mlx