Jevling-E2B-v0.1 β€” GGUF

Quantised build of BricksDisplay/jevling-e2b-v0.1 for on-device use (16 GB RAM). This is a System One decision model: one forward pass, no generated text, several typed questions per call, each answered as a calibrated probability. It does not work with stock llama.cpp chat/completion endpoints β€” they can load the weights but cannot ask a typed question or read the answer slot.

Use the maintained implementation

tools/system-one in mybigday/system-one-llama.cpp, branch feat/system-one:

git clone -b feat/system-one https://github.com/mybigday/system-one-llama.cpp
cd system-one-llama.cpp && cmake -B build && cmake --build build --target llama-system-one llama-system-one-batch -j

One call, three typed questions:

build/bin/llama-system-one -m jevling-e2b-v0.1-q8_0-embf16.gguf \
    --state "Customer: I was charged twice for the same subscription this month. Please refund the duplicate charge." \
    --choice "queue:Which team should handle this ticket?:billing=payments and refunds,technical=a product fault,account=login or profile settings" \
    --noul   "refund:Is the customer asking for a refund?" \
    --score  "urgency:How urgent is this?:routine,soon,urgent,critical" --json

Answers come back as probabilities per option (--json is the same shape the server returns: noul P(true), choice with a distribution, score as the expected level). Pass option descriptions β€” the model was trained with them. Use --system-one-request request.json for anything with colons or many questions; llama-system-one-batch for throughput. The prompt template and readout live inside the GGUF (tokenizer.chat_template.system_one, system_one.readout); do not supply a chat template of your own.

Measured: β‰ˆ1.7 s per 5-question request on 16 CPU threads, β‰ˆ2.9 GB RSS; run with --swa-full. Keep flash-attention off on CPU and threads = physical cores.

Evaluation

All numbers are accuracy on datasets the models were not trained on, asked in the System One format (all questions of an item in one request; option descriptions given where the dataset has them). Items per dataset: 120–150 unless noted. Both models of the series are shown; this card's model in bold.

benchmark task Jevling-0.8B-v0.1 Jevling-E2B-v0.1
MASSIVE (en-US) scenario classification, 18-way .675 .733
BBC News topic, 5-way .933 .958
TREC question type, 6-way .858 .850
PAWS paraphrase yes/no .508 .625
CommitmentBank NLI, 3-way .893 .875
StrategyQA yes/no reasoning .483 .567
PubMedQA yes/no/maybe .758 .667
SciQ 4-way science QA .942 .975
Social IQa 3-way .575 .725
TruthfulQA (MC) multiple choice .450 .633
XStoryCloze (en) 2-way .933 .958
QuALITY long-document 4-way QA .417 .500
RewardBench pairwise preference .600 .817
Hermes function-calling tool choice .996 .988
Financial PhraseBank sentiment, 3-way .608 .658
JevBench easy / original / hard (231 items) typed decisions 1.000 / .833 / .441 1.000 / .903 / .441
zh-TW kiosk set (ours, synthetic-derived, 255 states) intent acc / completeness AUROC / is-order / noise / size .969 / .971 / .996 / 1.000 / 1.000 .973 / .989 / .995 / 1.000 / 1.000

JevBench hard (.44) is where every open model we know of sits (TypeSafe's Jev API scores .73 on it); on the zh-TW kiosk set the same harness scores the two models at .971 / .973 and the Jev API at .931 β€” the kiosk set is our own synthetic-derived data, so read that comparison as "fit for the distribution it was built for", not as a general claim.

Limitations

  • Long-document, multi-hop, probability and date/time arithmetic questions are weak (JevBench hard .44; QuALITY .4–.5); compute arithmetic in code and put the result in the state.
  • Chinese coverage comes from synthetic kiosk-style data; validate on your own transcripts before relying on it.
  • When a question is undecidable the model picks the majority-style option; add an explicit "none of the above" option if you need abstention.
  • Not a chat model: it does not generate text.

Licence and release status

v0.1 is a research / non-commercial release. Some of the public datasets in the training mix carry non-commercial or research-only terms, so these weights are released under CC-BY-NC-4.0 on top of the Gemma Terms of Use (gemma-4 base).

Downloads last month
59
GGUF
Model size
5B params
Architecture
gemma4
Hardware compatibility
Log In to add your hardware

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for BricksDisplay/jevling-e2b-v0.1-GGUF

Quantized
(1)
this model

Collection including BricksDisplay/jevling-e2b-v0.1-GGUF