Looking forward to the Qwen 3.8 model drop, and congratulations on joining the trillion parameters club.
But my eyes are on the promised 27B model. Small (<50B) models decline on OpenRouter, because they are being run on local devices. If people see the family trees here on HF they will understand.
- evolutionary strategies - behavior steering experiments - expanded dataset - bringing back ORPO - more orthogonal evals to keep overfitting minimum - most probably will take abliterations as base, either mine or somebody else's - random entropy addition from huggingface fine tunes (take what is popular on hf and randomly introduce into the lineage) - bring more vibe coding: turns out LLMs know how to fine tune