Instructions to use Alsebay/L3-8B-SMaid-v0.1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Alsebay/L3-8B-SMaid-v0.1 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="Alsebay/L3-8B-SMaid-v0.1") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("Alsebay/L3-8B-SMaid-v0.1") model = AutoModelForCausalLM.from_pretrained("Alsebay/L3-8B-SMaid-v0.1", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Alsebay/L3-8B-SMaid-v0.1 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Alsebay/L3-8B-SMaid-v0.1" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Alsebay/L3-8B-SMaid-v0.1", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Alsebay/L3-8B-SMaid-v0.1
- SGLang
How to use Alsebay/L3-8B-SMaid-v0.1 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Alsebay/L3-8B-SMaid-v0.1" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Alsebay/L3-8B-SMaid-v0.1", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Alsebay/L3-8B-SMaid-v0.1" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Alsebay/L3-8B-SMaid-v0.1", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use Alsebay/L3-8B-SMaid-v0.1 with Docker Model Runner:
docker model run hf.co/Alsebay/L3-8B-SMaid-v0.1
Thank @mradermacher so much for help me find out that LumiMaid use 'smaug-bpe' pre-tokenizer. So that mean all its quant is unable to use. That mean you can only use Transformer to load this model for now (maybe they will fix or add feature in future)
Update: Both version have different presents (settings) to work well
Overall:
Sao10K Stheno > SMaid V0.3 > SMaid V0.1 in Chai Benchmark
SMaid V0.1 = Sao10K Stheno > SMaid V0.3 in my custom EQ bench (Sadness and deep thought and Depression test)
Disclaimed: same seed, same character card, same scenario. 4 times try for each models.
Best of L3-8B merge series for me. I choose 2 best variants to publish.
SMaid-V0.1: More smart, understand well content, more novelwriting. I like this version.
SMaid-V0.3: Upgrade from v0.1. More talkative, active, energetic (wrong setting, lol).
No V0.2 because I deleted it, it's a worst model of series.
I think Stheno and Lumumaid can be like a 'ying-yang', so I combine them, lol. Have test on Chaiverse, both of them got > 1995 elo score from begining. (Thanks Sao10K let me know about ChaiVerse :) )
SMaid = Stheno (it's very good) + LumiMaid (not too good, but the writing style is good)
Recommend present (You can feedback if any setting is better)
Temperature - 1.1-1.25
Min-P - 0.075
Top-K - 50
Top_P - 0.5
Repetition Penalty - 1.1
Below is the auto-generate by Mergekit
This is a merge of pre-trained language models created using mergekit.
Merge Details
Merge Method
This model was merged using the DARE TIES merge method using NeverSleep/Llama-3-Lumimaid-8B-v0.1-OAS as a base.
Models Merged
The following models were included in the merge:
Configuration
The following YAML configuration was used to produce this model:
slices:
- sources:
- layer_range: [0, 16]
model: NeverSleep/Llama-3-Lumimaid-8B-v0.1-OAS
parameters:
density: 0.5
weight: 1.0
- layer_range: [0, 16]
model: Sao10K/L3-8B-Stheno-v3.2
parameters:
density: 0.5
weight: 0.9
- sources:
- layer_range: [16, 24]
model: Sao10K/L3-8B-Stheno-v3.2
parameters:
density: 0.75
weight: 0.5
- layer_range: [16, 24]
model: NeverSleep/Llama-3-Lumimaid-8B-v0.1-OAS
parameters:
density: 0.25
weight: 0.5
- sources:
- layer_range: [24, 32]
model: NeverSleep/Llama-3-Lumimaid-8B-v0.1-OAS
parameters:
density: 0.5
weight: 0.5
- layer_range: [24, 32]
model: Sao10K/L3-8B-Stheno-v3.2
parameters:
density: 0.5
weight: 1.0
merge_method: dare_ties
base_model: NeverSleep/Llama-3-Lumimaid-8B-v0.1-OAS
parameters:
int8_mask: true
dtype: bfloat16
- Downloads last month
- 21