Instructions to use BricksDisplay/jevling-e2b-v0.1-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use BricksDisplay/jevling-e2b-v0.1-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf BricksDisplay/jevling-e2b-v0.1-GGUF:BF16 # Run inference directly in the terminal: llama cli -hf BricksDisplay/jevling-e2b-v0.1-GGUF:BF16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf BricksDisplay/jevling-e2b-v0.1-GGUF:BF16 # Run inference directly in the terminal: llama cli -hf BricksDisplay/jevling-e2b-v0.1-GGUF:BF16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf BricksDisplay/jevling-e2b-v0.1-GGUF:BF16 # Run inference directly in the terminal: ./llama-cli -hf BricksDisplay/jevling-e2b-v0.1-GGUF:BF16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf BricksDisplay/jevling-e2b-v0.1-GGUF:BF16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf BricksDisplay/jevling-e2b-v0.1-GGUF:BF16
Use Docker
docker model run hf.co/BricksDisplay/jevling-e2b-v0.1-GGUF:BF16
- LM Studio
- Jan
- Ollama
How to use BricksDisplay/jevling-e2b-v0.1-GGUF with Ollama:
ollama run hf.co/BricksDisplay/jevling-e2b-v0.1-GGUF:BF16
- Unsloth Desktop
- Docker Model Runner
How to use BricksDisplay/jevling-e2b-v0.1-GGUF with Docker Model Runner:
docker model run hf.co/BricksDisplay/jevling-e2b-v0.1-GGUF:BF16
- Lemonade
How to use BricksDisplay/jevling-e2b-v0.1-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull BricksDisplay/jevling-e2b-v0.1-GGUF:BF16
Run and chat with the model
lemonade run user.jevling-e2b-v0.1-GGUF-BF16
List all available models
lemonade list
- Atomic Chat
Jevling-E2B-v0.1 β GGUF
Quantised build of BricksDisplay/jevling-e2b-v0.1 for on-device use (16 GB RAM). This is a System One decision model: one forward pass, no generated text, several typed questions per call, each answered as a calibrated probability. It does not work with stock llama.cpp chat/completion endpoints β they can load the weights but cannot ask a typed question or read the answer slot.
Use the maintained implementation
tools/system-one in mybigday/system-one-llama.cpp, branch feat/system-one:
git clone -b feat/system-one https://github.com/mybigday/system-one-llama.cpp
cd system-one-llama.cpp && cmake -B build && cmake --build build --target llama-system-one llama-system-one-batch -j
One call, three typed questions:
build/bin/llama-system-one -m jevling-e2b-v0.1-q8_0-embf16.gguf \
--state "Customer: I was charged twice for the same subscription this month. Please refund the duplicate charge." \
--choice "queue:Which team should handle this ticket?:billing=payments and refunds,technical=a product fault,account=login or profile settings" \
--noul "refund:Is the customer asking for a refund?" \
--score "urgency:How urgent is this?:routine,soon,urgent,critical" --json
Answers come back as probabilities per option (--json is the same shape the server returns: noul P(true), choice with a distribution, score as the expected level). Pass option descriptions β the model was trained with them. Use --system-one-request request.json for anything with colons or many questions; llama-system-one-batch for throughput. The prompt template and readout live inside the GGUF (tokenizer.chat_template.system_one, system_one.readout); do not supply a chat template of your own.
Measured: β1.7 s per 5-question request on 16 CPU threads, β2.9 GB RSS; run with --swa-full. Keep flash-attention off on CPU and threads = physical cores.
Evaluation
All numbers are accuracy on datasets the models were not trained on, asked in the System One format (all questions of an item in one request; option descriptions given where the dataset has them). Items per dataset: 120β150 unless noted. Both models of the series are shown; this card's model in bold.
| benchmark | task | Jevling-0.8B-v0.1 | Jevling-E2B-v0.1 |
|---|---|---|---|
| MASSIVE (en-US) | scenario classification, 18-way | .675 | .733 |
| BBC News | topic, 5-way | .933 | .958 |
| TREC | question type, 6-way | .858 | .850 |
| PAWS | paraphrase yes/no | .508 | .625 |
| CommitmentBank | NLI, 3-way | .893 | .875 |
| StrategyQA | yes/no reasoning | .483 | .567 |
| PubMedQA | yes/no/maybe | .758 | .667 |
| SciQ | 4-way science QA | .942 | .975 |
| Social IQa | 3-way | .575 | .725 |
| TruthfulQA (MC) | multiple choice | .450 | .633 |
| XStoryCloze (en) | 2-way | .933 | .958 |
| QuALITY | long-document 4-way QA | .417 | .500 |
| RewardBench | pairwise preference | .600 | .817 |
| Hermes function-calling | tool choice | .996 | .988 |
| Financial PhraseBank | sentiment, 3-way | .608 | .658 |
| JevBench easy / original / hard (231 items) | typed decisions | 1.000 / .833 / .441 | 1.000 / .903 / .441 |
| zh-TW kiosk set (ours, synthetic-derived, 255 states) | intent acc / completeness AUROC / is-order / noise / size | .969 / .971 / .996 / 1.000 / 1.000 | .973 / .989 / .995 / 1.000 / 1.000 |
JevBench hard (.44) is where every open model we know of sits (TypeSafe's Jev API scores .73 on it); on the zh-TW kiosk set the same harness scores the two models at .971 / .973 and the Jev API at .931 β the kiosk set is our own synthetic-derived data, so read that comparison as "fit for the distribution it was built for", not as a general claim.
Limitations
- Long-document, multi-hop, probability and date/time arithmetic questions are weak (JevBench hard .44; QuALITY .4β.5); compute arithmetic in code and put the result in the state.
- Chinese coverage comes from synthetic kiosk-style data; validate on your own transcripts before relying on it.
- When a question is undecidable the model picks the majority-style option; add an explicit "none of the above" option if you need abstention.
- Not a chat model: it does not generate text.
Licence and release status
v0.1 is a research / non-commercial release. Some of the public datasets in the training mix carry non-commercial or research-only terms, so these weights are released under CC-BY-NC-4.0 on top of the Gemma Terms of Use (gemma-4 base).
- Downloads last month
- 59
16-bit
Model tree for BricksDisplay/jevling-e2b-v0.1-GGUF
Base model
google/gemma-4-E2B