Instructions to use nitinpanj/qwen38-flash-next-v3 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use nitinpanj/qwen38-flash-next-v3 with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf nitinpanj/qwen38-flash-next-v3:Q4_K_M # Run inference directly in the terminal: llama cli -hf nitinpanj/qwen38-flash-next-v3:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf nitinpanj/qwen38-flash-next-v3:Q4_K_M # Run inference directly in the terminal: llama cli -hf nitinpanj/qwen38-flash-next-v3:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf nitinpanj/qwen38-flash-next-v3:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf nitinpanj/qwen38-flash-next-v3:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf nitinpanj/qwen38-flash-next-v3:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf nitinpanj/qwen38-flash-next-v3:Q4_K_M
Use Docker
docker model run hf.co/nitinpanj/qwen38-flash-next-v3:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use nitinpanj/qwen38-flash-next-v3 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "nitinpanj/qwen38-flash-next-v3" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "nitinpanj/qwen38-flash-next-v3", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/nitinpanj/qwen38-flash-next-v3:Q4_K_M
- Ollama
How to use nitinpanj/qwen38-flash-next-v3 with Ollama:
ollama run hf.co/nitinpanj/qwen38-flash-next-v3:Q4_K_M
- Unsloth Desktop
- Pi
How to use nitinpanj/qwen38-flash-next-v3 with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf nitinpanj/qwen38-flash-next-v3:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "nitinpanj/qwen38-flash-next-v3:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use nitinpanj/qwen38-flash-next-v3 with Docker Model Runner:
docker model run hf.co/nitinpanj/qwen38-flash-next-v3:Q4_K_M
- Lemonade
How to use nitinpanj/qwen38-flash-next-v3 with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull nitinpanj/qwen38-flash-next-v3:Q4_K_M
Run and chat with the model
lemonade run user.qwen38-flash-next-v3-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use nitinpanj/qwen38-flash-next-v3 with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf nitinpanj/qwen38-flash-next-v3:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default nitinpanj/qwen38-flash-next-v3:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use nitinpanj/qwen38-flash-next-v3 with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf nitinpanj/qwen38-flash-next-v3:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "nitinpanj/qwen38-flash-next-v3:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Qwen3.8-Flash-Next โ Q4_0-Q8out-v3
Run a 95.5 GiB model on a 64 GB Mac, at ~27 tokens/sec.
This checkpoint is tuned for one specific job: running Qwen3.8-Flash-Next on Apple Silicon with 64 GB of unified memory, by keeping the rarely-used expert weights on SSD and reading them only when a token actually needs them.
โ ๏ธ Read this first
This model does not run on stock llama.cpp. Not "runs slower" โ it will try to load all 95.5 GiB into memory and fail.
You need this fork, which adds --moe-stream:
๐ https://github.com/npanj/llama.cpp
Upstream llama.cpp does support the qwen4exp architecture. What it does not have is expert
streaming, which is the entire reason this checkpoint fits on a 64 GB machine.
What you need
| Machine | Apple Silicon Mac, 64 GB unified memory |
| Disk | ~100 GB free on the internal SSD |
| Software | npanj/llama.cpp, built with Metal |
The disk matters more than you'd expect โ expert weights are read from it continuously while generating. An external USB drive will be much slower.
Quick start
Full instructions, troubleshooting and flag-by-flag explanation: docs/qwen38-flash-next-v3.md
1. Build the fork
git clone https://github.com/npanj/llama.cpp
cd llama.cpp
cmake -B build -DGGML_METAL=ON -DCMAKE_BUILD_TYPE=Release
cmake --build build -j8 --config Release
2. Download the checkpoint (95.5 GiB, 3 shards)
D=~/models/qwen38-flash-next-v3 && mkdir -p $D
for i in 1 2 3; do
curl -fL --retry 5 -C - -o $D/Qwen3.8-Flash-Next-Q4_0-Q8out-v3-0000$i-of-00003.gguf \
https://huggingface.co/nitinpanj/qwen38-flash-next-v3/resolve/main/Qwen3.8-Flash-Next-Q4_0-Q8out-v3-0000$i-of-00003.gguf
done
3. Download the draft head (1.9 GiB โ worth about +50% generation speed)
D=~/models/qwen38-flash-next-mtp && mkdir -p $D
curl -fL --retry 5 -C - -o $D/mtp-shared-Q4_K_M.gguf \
https://huggingface.co/nitinpanj/qwen38-flash-next-v3/resolve/main/MTP/mtp-shared-Q4_K_M.gguf
4. Let the GPU wire enough memory โ don't skip this, and it resets on reboot
sudo sysctl iogpu.wired_limit_mb=59392
5. Run
export LLAMA_MOE_STREAM_LOOKAHEAD=1 LLAMA_MOE_STREAM_WAVE_CAP=200 \
LLAMA_MOE_STREAM_PARTITION=1 LLAMA_QWEN4EXP_SPARSE_FA=1
./build/bin/llama-server \
-m ~/models/qwen38-flash-next-v3/Qwen3.8-Flash-Next-Q4_0-Q8out-v3-00001-of-00003.gguf \
-md ~/models/qwen38-flash-next-mtp/mtp-shared-Q4_K_M.gguf \
-ngl 99 \
--moe-stream --moe-stream-cache 36 --moe-stream-io-threads 8 --moe-stream-direct \
-c 98304 -b 4096 -ub 4096 -cms 512 -np 1 -fa on \
--cache-reuse 0 --cache-ram 512 \
--jinja --reasoning-format deepseek \
--spec-type draft-mtp --spec-draft-n-max 3 --spec-draft-p-min 0.3 \
--spec-draft-ngl 99 --spec-max-prompt 0 \
--host 127.0.0.1 --port 8080
Then open http://127.0.0.1:8080, or point any OpenAI-compatible client at
http://127.0.0.1:8080/v1.
First load takes a few minutes โ it is reading 95.5 GiB off disk.
Files
| File | Size | What it is |
|---|---|---|
Qwen3.8-Flash-Next-Q4_0-Q8out-v3-0000{1,2,3}-of-00003.gguf |
95.5 GiB | The model. Point -m at shard 1; it finds the rest. |
MTP/mtp-shared-Q4_K_M.gguf |
1.9 GiB | Draft head for speculative decoding. |
About the draft head. It has no token embeddings of its own โ it borrows the main model's, which is how it stays at 1.9 GiB instead of 2.6. So it only works when passed with
-mdalongside this checkpoint, on this fork. It cannot be loaded on its own.Its tensor names are fork-specific too. Upstream PR #28243 adds shared-MTP support for this architecture using unsloth's original naming, not the names this fork renames to. If that lands, unsloth's sidecar will load on upstream llama.cpp directly and this file still will not.
Prefer to build your own from unsloth's published sidecar? The scripts and an explanation of the conversion are at
scripts/mtp/.
Measured performance
Apple M5 Pro, 64 GB, with the configuration above:
| Prompt reading (4k) | ~367 tokens/sec |
| Generation, draft head on | ~27.6 tokens/sec |
| Generation, draft head off | 18.0-18.6 tokens/sec |
| Generation at 29k context, real chat traffic | ~20.6 tokens/sec |
Generation slows as context fills. That is expected โ at depth the cost is dominated by reading the attention cache, not by the streamed expert weights.
What "v3" is
bartowski's Q4_0 with a Q8_0 output.weight spliced in, plus unsloth's UD-IQ4_XS spliced into five
tensor groups that stay resident in memory (attn, hc, token_embd, ssm_out, shexp).
Measured against the unspliced checkpoint, paired 40-chunk perplexity:
| Perplexity | 5.2777 โ 4.3148 (โ17.8%), better on 40 of 40 chunks |
| Draft acceptance | 0.751 โ 0.817 |
| Decode speed | โ2.7% |
| Disk | +1.69 GiB |
The five resident groups are not arbitrary. Trimming to attn,hc,token_embd measures the same
perplexity but drops decode 14% โ ssm_out and shexp do nothing for perplexity and are worth
about 11 points of decode.
If your Mac isn't 64 GB
- Less memory: lower
--moe-stream-cacheand the matching wired cap. It still loads; expect it to be slower. Untested below 64 GB. - More memory: raise the cache โ but in small steps. On a 64 GB machine, going from 36 to 38 GiB of cache collapsed generation from ~24 to ~3.7 tokens/sec as macOS began swapping. More is not monotonically better.
- Not a Mac: untested. Expert streaming is not Metal-specific in principle, but nothing here has been tuned or measured on CUDA or CPU.
Credit
- Base model: Qwen3.8-Flash-Next by Qwen.
- Quantization sources: bartowski (Q4_0) and unsloth (UD-IQ4_XS, and the MTP sidecar this draft head is converted from).
- Expert streaming (
--moe-stream), the feature that makes this run at all, comes from mihailescu2m/llama.cpp.
This repository is the spliced checkpoint and the converted draft head. Model weights remain under the original Qwen license.
Problems? Open an issue at github.com/npanj/llama.cpp/issues. Please say which Mac and how much memory you have, and include the first ~30 lines the server prints at startup.
- Downloads last month
- 34,499
4-bit