Kodiak-v0.2-1B, accuracy mode

The most accurate and most trustworthy Kodiak. Three independently trained Kodiak-v0.2-1B models answer every question, and their calibrated answers are averaged. Each model is overconfident in different places, so averaging cancels much of it. On decisions it was never trained for, it beats Qwen3-8B (0.706 vs 0.688), and when it says "can't tell" it's right 90.5% of the time. It costs about 3× the compute of the single model.

Kodiak is an open decision model by Cortex Agent LLC. You give it a state (text, a list of texts, or JSON) and typed questions. It returns calibrated choice, score or "can't tell" answers. Use it to automate routine "read this and decide" work: routing, triage, guardrails and checks. Send the cases it isn't sure about to a person or an LLM. Each member is built on Ettin-encoder-1B (Johns Hopkins, MIT license), with 1.04B parameters. Code, docs and the full public build log: https://github.com/grizzlypeaksoftware/kodiak

# pip install "kodiak-s1[infer] @ git+https://github.com/grizzlypeaksoftware/kodiak"
from kodiak_s1.hub import Kodiak

kodiak = Kodiak.from_pretrained("cortex-agent-llc/kodiak-v0.2-1b-accuracy")   # loads all three members
kodiak.decide(
    "Hi, I ordered the walnut desk two weeks ago. Tracking has said 'label created' for 10 days.",
    [{"type": "choice", "id": "intent", "text": "What does the customer want?",
      "labels": ["delivery status or expedite", "cancel and refund", "product question"]},
     {"type": "choice", "id": "carrier", "text": "Which carrier is shipping it?", "labels": ["UPS", "FedEx", "USPS"]}],
)

For the fastest answers, use the single model: Kodiak-v0.2-1B.

How it compares (frozen eval set v0.2, choice questions)

Kodiak XL v2 preview (1B) Kodiak-v0.2-1B v0.2 accuracy mode Qwen3-8B (LLM)
Never-seen tasks, forced accuracy 0.659 ± 0.013 0.689 ± 0.008 0.706 0.688
Familiar tasks 0.881 0.877 0.889 0.710
Ranks its own mistakes last (never-seen; share of the gap to a perfect order) – 55.6% 56.6% 14.7%
Calibration error (never-seen; lower is better)¹ 0.113 0.085 0.062 0.293
When it says "can't tell", it's right 0.87 0.88 0.905 –
Latency (GPU, one request) 38 ms 38 ms ~3× ~1,500 ms

Accuracy mode averages the three training runs behind the Kodiak-v0.2-1B ± figures. Its abstain threshold (0.6) was tuned on validation data only, the highest-accuracy threshold whose validation abstain precision is at least 0.90. "Never-seen" means tasks and label sets the model was never trained on (zero-shot). The eval set and the LLM baseline setup are in the repository.

Ranking is what a confidence threshold relies on: sort the answers by confidence, and see how much of the gap between a random order and the perfect order (all mistakes last) the model closes. It doesn't change under any recalibration. Kodiak-v0.2-1B: 54.8% / 56.1% / 55.9% over three runs. ¹ The calibration comparison is raw. Qwen3-8B is mostly overconfident by a constant amount, so a fitted recalibration (isotonic, fit on the other never-seen tasks) brings it from 0.293 to 0.178; the same treatment gives Kodiak-v0.2-1B 0.044-0.058. Thanks to dipankarsarkar for pointing this out. Full numbers: reports/v02-release-vs-llm-8b.md in the repository.

What's new in v0.2

  • Four new kinds of decision. These are checking whether an answer is grounded in its source (and which sentence isn't), choosing the next step for an assistant from full API specs (call a tool, ask for missing information, or answer directly), verifying a claim against a text, and judging product relevance for a search (exact match, substitute, complement or irrelevant). On our held-out skills test the score rose from 0.52 to 0.98. On public real-world data, the gains carry over:

    Benchmark XL v2 preview Kodiak-v0.2-1B
    Amazon ESCI (product relevance) 0.04 0.23
    HoVer (claim verification) 0.12 0.23
    BFCL (function calling) 0.20 0.28
    ANLI (adversarial inference) 0.01 0.10
    CLINC150 (intent) 0.84 0.86

    These are Decision Index scores, chance-corrected so that 0 means random guessing. Each is a three-run mean on a fixed sample of up to 1,000 questions per benchmark.

  • Reads options by meaning, not wording. Training also shows the answer options in different wordings: a short label, a sentence, or a paraphrase. This raised never-seen accuracy by 2.7 points and cut calibration error by about 20%. The biggest gain was on options with ambiguous words, such as poem sentiment with "mixed" (0.54 to 0.70).

  • All the training data is synthetic or permissively licensed. The synthetic data was written and checked by open-weight models (gpt-oss-120b and DeepSeek-V3.2). No closed-model outputs and no evaluation data are in training.

Known limits (read before using)

  • Option wording still matters. The same question with reworded options gets the same answer about 68% of the time on never-seen tasks (up from 62%). Long, sentence-style options can tilt it toward one answer: on one phishing benchmark (PhishNChips) it labels most emails "phishing" when the two options are long descriptions, and it scores near chance there. Keep options short and distinct, and test a few wordings on your own data. This is the main goal for v0.3.
  • Wording traps. A message that repeats one option's words inside a condition can pull the answer toward that option. For example: "if it can't arrive by Monday, cancel and refund me".
  • Not fixed in v0.2. Response-level hallucination detection on long answers (RAGTruth) and some API-call benchmarks (API-Bank, When2Call) are still near chance. Entity-level financial sentiment (FinEntity) dropped from 0.36 to 0.16 compared with XL v2.
  • "Can't tell" precision is 0.905 at the default threshold. Raise null_threshold per request if false abstentions cost you.
  • Long inputs. The position limit is about 8,000 tokens (the state, every question and every option together). Longer requests are refused rather than truncated.
  • Ratings ("how urgent is this?") are rough. Prefer choice questions.
  • Speed. About 3× the single model: roughly 0.1 s per request on a GPU, and around a second on a CPU. Use the single model when speed matters.
  • Validate it on your own data. Don't use it for decisions about people without human review. Teach it your own decisions with the fine-tuning kit in the repository (a CSV in, a before/after report out).

License and citation

Apache-2.0 (weights and code). Base model: Ettin-encoder-1B (MIT). Training-data licenses are listed in data/LICENSES.md in the repository. Cortex Agent LLC / Grizzly Peak Software, 2026.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for cortex-agent-llc/kodiak-v0.2-1b-accuracy

Finetuned
(18)
this model