Instructions to use Doctor-Shotgun/MS3.2-24B-Magnum-Diamond with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Doctor-Shotgun/MS3.2-24B-Magnum-Diamond with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="Doctor-Shotgun/MS3.2-24B-Magnum-Diamond") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("Doctor-Shotgun/MS3.2-24B-Magnum-Diamond") model = AutoModelForCausalLM.from_pretrained("Doctor-Shotgun/MS3.2-24B-Magnum-Diamond", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Inference
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Doctor-Shotgun/MS3.2-24B-Magnum-Diamond with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Doctor-Shotgun/MS3.2-24B-Magnum-Diamond" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Doctor-Shotgun/MS3.2-24B-Magnum-Diamond", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Doctor-Shotgun/MS3.2-24B-Magnum-Diamond
- SGLang
How to use Doctor-Shotgun/MS3.2-24B-Magnum-Diamond with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Doctor-Shotgun/MS3.2-24B-Magnum-Diamond" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Doctor-Shotgun/MS3.2-24B-Magnum-Diamond", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Doctor-Shotgun/MS3.2-24B-Magnum-Diamond" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Doctor-Shotgun/MS3.2-24B-Magnum-Diamond", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use Doctor-Shotgun/MS3.2-24B-Magnum-Diamond with Docker Model Runner:
docker model run hf.co/Doctor-Shotgun/MS3.2-24B-Magnum-Diamond
Instability, some issues, could you please help?
Hi. You did an excellent fine-tune. Just to be sure, I'll mention the model name: Magnum Mistral Small 3.2 24B. But there's one problem - during roleplay, it doesn't wrap actions, descriptions, and emotions in asterisks like this. Or it might suddenly write something off-topic from the roleplay, which is rare but still affects stability. Sometimes it wraps dialogue in asterisks, which again messes up formatting "like this". Is there any way to fix this instability?
I can show the problem in Discord with screenshots and other things.
Write me please if you can solve this problem. Discord swatman8331
System prompt does not help to solve the problem.
That's odd because my experience with Magnum has been quite the opposite - it has a bias in favor of adding asterisk wrapping even when the character card's first message is not written with asterisk wrapping.
It's hard to comment on what's happening on your end without more information:
- What frontend are you using? Chat completions or raw text completions?
- What quant are you using?
- What sampler settings are you using?
- What character cards/prompts are you using?
I would ask you to write to me in discord, because there is a faster and more specific connection than in the discussions here. I can just write you a link, you will see the result yourself.
Please visit the website, select any character for conversation, choose your model from the list (it's top 2). There will be an issue during roleplay.
By the way, I was asked to ask you - are you interested in becoming a pre-trainer?
Problem - It’s the template
Issue
Hmm so it's through a chat service - it's a bit hard to say since I'm not sure what is being done exactly behind the scenes. My suspicion would be a prompt templating issue or sampler issue.
I suggest you contact the site staff, they have been trying to solve this problem for several days now. Or we can contact you. Please provide your discord profile contacts or email.
Sent you a requets on Discord, can discuss there if it's easier for you.
I wrote to you on discord.