AI & ML interests

None defined yet.

Recent Activity

comgen42 
posted an update 2 days ago
view post
Post
89
Kodiak v0.4 is out: cortex-agent-llc/kodiak-v0.4-1b (plus an accuracy mode that averages three runs).

This release came out of public feedback. After v0.3, someone showed that two of its skills were answering from keywords instead of reading the case. So we built tests a keyword shortcut can't pass: the same case twice, with one detail changed so the right answer flips.

Refund eligibility is now fixed: it gets both versions right 79% of the time, up from 22%. The push-to-main-with-failing-tests case that started this now gets "ask the user first". Agent step safety improved but isn't fixed, and the model card says plainly not to use it as a safety control. Sarcasm still shows no signal on real tweets, and accuracy on brand-new kinds of task hasn't moved. That's the big problem for the next version.

Now a pause. I'm shutting the training box down for two weeks while I travel in Turkey. I'll be walking through ancient ruins: places where people built things that lasted thousands of years without any of our tools. I'm hoping that does what travel usually does for me, shakes loose some ideas. I'll be thinking about how to teach Kodiak to handle tasks it has never seen, how to run my publishing company better, and where agentic automation actually earns its keep.

No training while I'm gone. When I'm back, Kodiak gets the ideas.

Every number, including the misses, is in the public build log: github.com/grizzlypeaksoftware/kodiak
comgen42 
posted an update 3 days ago
view post
Post
101
Kodiak v0.2 1B got its official Decision Index listing this week: 11.8, rank 96 of 115.

I'll be honest, that one stung. I've put a lot into this model. Our own run of the public benchmarks said about 19, and the index confirmed that part (19.1). But most of the full score comes from private tests, and on tasks from new domains Kodiak scored close to zero. It's good at what it was trained for: intent routing at 91.6, common sense, ranking. It falls apart on kinds of decisions it has never seen. Sarcasm on real tweets came out worse than random.

There's a particular kind of tired that comes from working hard on something and having a scoreboard tell you you're near the bottom.

Then I remember what I've learned from history. The Wright brothers came home from Kitty Hawk in 1901 convinced the published lift tables were wrong. Instead of quitting, they built their own wind tunnel and tested hundreds of wing shapes. James Dyson went through more than 5,000 prototypes before one worked. Failing wasn't the exception for the people who built things that mattered. It was the job. What separated them was that they kept going and kept measuring honestly.

So that's the plan. The index just told us exactly where the weakness is: generalizing to new kinds of tasks. Meanwhile a Hugging Face user found two shortcuts in v0.3, we fixed one and are testing the fix for the other right now, and every number goes into the public log, good or bad.

Rank 96 is a data point, not a verdict. Back to work.

github.com/grizzlypeaksoftware/kodiak
  • 2 replies
·
comgen42 
posted an update 5 days ago
view post
Post
3254
🐻 Kodiak-v0.3-1B: seven new kinds of decision, measured on real data.

Kodiak is an open 1B encoder that answers typed questions with calibrated confidence, or says "can't tell". New in v0.3:
• pairwise judge (which answer is better?)
• long-answer hallucination checks
• stance and sarcasm
• policy violation and refund-eligibility checks against your written rules
• agent step safety (run it, ask first, or never)

We stopped trusting our own synthetic tests and judged v0.3 on real labelled data it never trained on (RAGBench, MT-Bench human judgments, SemEval stance): 0.21 → 0.36, averaged over 3 training runs. Nothing else got worse, and when it says "can't tell" it's right 93% of the time (up from 88%).

Known limits are in the model card.

📦 cortex-agent-llc/kodiak-v0.3-1b
🎯 Accuracy mode: cortex-agent-llc/kodiak-v0.3-1b-accuracy
🕹️ Demo: comgen42/kodiak-demo
📝 What we built, and the kill that changed our process: https://cortexagent.com/blog/kodiak-v0-3-seven-new-kinds-of-decision-measured-on-real-data
  • 4 replies
·
comgen42 
posted an update 9 days ago
view post
Post
117
🐻 Kodiak-v0.2-1B is out: an open 1B encoder that makes decisions instead of writing text.

Give it a state and typed questions. You get back calibrated answers in one forward pass, and it says "can't tell" when it doesn't know.

On tasks it was never trained for (our frozen eval set v0.2):
• 0.689 accuracy (3-run mean) vs 0.688 for Qwen3-8B, at ~40× the speed
• calibration error 0.085 vs 0.293
• accuracy mode (3 models averaged): 0.706

New in v0.2: grounding checks, picking an assistant's next tool step from API specs, claim verification, and product relevance.

Known limits are in the model card.

📦 cortex-agent-llc/kodiak-v0.2-1b
🎯 Accuracy mode: cortex-agent-llc/kodiak-v0.2-1b-accuracy
🕹️ Demo: comgen42/kodiak-demo
📝 What we built, and what didn't work: https://cortexagent.com/blog/kodiak-v0-2-an-open-1b-decision-model-you-can-download-today
  • 4 replies
·