AI & ML interests
Pretraining, QAT, SLM, Overtraining, AI Interpretability.
Banaxi-Techย
posted an update about 20 hours ago
Banaxi-Techย
posted an update 2 days ago
Post
3183
We did an experiment, we wanted to see if AI is good enough to train models.
We used GPT 5.6 Sol Max for this because its one of the most powerful ones right now.
Our instructions were, it should write the training code, and start the training process and monitor it by itself.
We also gave it a link to BananaMind 2 Mini to get our architecture right.
The result: It worked, it made the working BananaMind 2 Nano, and even beat our previous MiniBananaMind v4 9M.
Its getting way easier to develop your own models now!
We used GPT 5.6 Sol Max for this because its one of the most powerful ones right now.
Our instructions were, it should write the training code, and start the training process and monitor it by itself.
We also gave it a link to BananaMind 2 Mini to get our architecture right.
The result: It worked, it made the working BananaMind 2 Nano, and even beat our previous MiniBananaMind v4 9M.
Its getting way easier to develop your own models now!
Banaxi-Techย
posted an update 3 days ago
Post
2288
We're excited to release BananaMind 2 Pro Preview, our best model yet.
Trained on ~52B tokens it performs extremely good for its token and size class.
We trained it on a single 5070 Ti in about 11 days.
Check it out at BananaMind/BananaMind-2-Pro-Preview.
Sadly we need to delay BananaMind 2 Micro until the launch of the final BananaMind 2 Pro.
We will release the final checkpoint with 100B tokens in ~11 days.
Go and fine-tune it!
We've also released BananaMind 2 Pro Preview Chat which is the instruct version of it!
BananaMind/BananaMind-2-Pro-Preview-Chat
Follow us to know when the final releases and support us at
BananaMind
@Banaxi-Tech
Trained on ~52B tokens it performs extremely good for its token and size class.
We trained it on a single 5070 Ti in about 11 days.
Check it out at BananaMind/BananaMind-2-Pro-Preview.
Sadly we need to delay BananaMind 2 Micro until the launch of the final BananaMind 2 Pro.
We will release the final checkpoint with 100B tokens in ~11 days.
Go and fine-tune it!
We've also released BananaMind 2 Pro Preview Chat which is the instruct version of it!
BananaMind/BananaMind-2-Pro-Preview-Chat
Follow us to know when the final releases and support us at
@Banaxi-Tech
DedeProGamesย
posted an update 4 days ago
Banaxi-Techย
posted an update 4 days ago
Post
2672
BananaMind 2 pro Has BEEN RELEASED! BananaMind/BananaMind-2-Pro-Preview
Previous content:
BananaMind 2 Pro Preview will launch tomorrow.
Give us a follow:
BananaMind
Lets get 70 or 75 followers before it releases.
It takes 5 seconds.
August 3, 1PM in Austria time
Previous content:
BananaMind 2 Pro Preview will launch tomorrow.
Give us a follow:
Lets get 70 or 75 followers before it releases.
It takes 5 seconds.
August 3, 1PM in Austria time
DedeProGamesย
posted an update 6 days ago
Post
1527
๐ Introducing the GRM-3.2 Family
The GRM-3.2 family is a new generation of reasoning-focused models from OrionLLM, purpose-built for long-horizon agentic tasks, extremely difficult reasoning problems, advanced coding, and local AI workflows across a wide range of hardware constraints.
GRM-3.2-Sky is the flagship model in the family: a 35B-A3B Mixture-of-Experts model built on the Ornith-1.0-35B architecture, designed for elite structured reasoning, complex multi-file coding, advanced mathematics, and sustained coherence across extended agentic workflows. It represents a substantial leap in long-horizon task capability over its predecessor, GRM-2.6-Plus.
GRM-3.2-Cliff is the mid-sized workhorse: a 9B-parameter model optimized for long-horizon agentic tasks and difficult reasoning in low-to-mid GPU environments. It delivers strong multi-step planning, debugging, and terminal-agent performance without demanding flagship-level hardware.
GRM-3.2-Turf is the lightweight edge model: a 1.2B-parameter model based on the LiquidAI/LFM2.5-1.2B-Thinking architecture, engineered for efficient on-device execution, high-fidelity instruction following, and robust tool use on mobile, embedded, and other resource-constrained hardware.
All three models are designed for users who need dependable reasoning engines that can maintain goal-directed behavior, planning quality, and task fidelity across many stepsโwhether on a server, a local workstation, or an edge device.
Models:
GRM-3.2-Sky: OrionLLM/GRM-3.2-Sky
GRM-3.2-Cliff: OrionLLM/GRM-3.2-Cliff
GRM-3.2-Turf: OrionLLM/GRM-3.2-Turf
Organization:
OrionLLM
The GRM-3.2 family is a new generation of reasoning-focused models from OrionLLM, purpose-built for long-horizon agentic tasks, extremely difficult reasoning problems, advanced coding, and local AI workflows across a wide range of hardware constraints.
GRM-3.2-Sky is the flagship model in the family: a 35B-A3B Mixture-of-Experts model built on the Ornith-1.0-35B architecture, designed for elite structured reasoning, complex multi-file coding, advanced mathematics, and sustained coherence across extended agentic workflows. It represents a substantial leap in long-horizon task capability over its predecessor, GRM-2.6-Plus.
GRM-3.2-Cliff is the mid-sized workhorse: a 9B-parameter model optimized for long-horizon agentic tasks and difficult reasoning in low-to-mid GPU environments. It delivers strong multi-step planning, debugging, and terminal-agent performance without demanding flagship-level hardware.
GRM-3.2-Turf is the lightweight edge model: a 1.2B-parameter model based on the LiquidAI/LFM2.5-1.2B-Thinking architecture, engineered for efficient on-device execution, high-fidelity instruction following, and robust tool use on mobile, embedded, and other resource-constrained hardware.
All three models are designed for users who need dependable reasoning engines that can maintain goal-directed behavior, planning quality, and task fidelity across many stepsโwhether on a server, a local workstation, or an edge device.
Models:
GRM-3.2-Sky: OrionLLM/GRM-3.2-Sky
GRM-3.2-Cliff: OrionLLM/GRM-3.2-Cliff
GRM-3.2-Turf: OrionLLM/GRM-3.2-Turf
Organization:
Banaxi-Techย
posted an update 6 days ago
Post
1804
BananaMind 2 Pro Preview will release when we hit 75 followers on BananaMind!
Follow us for the release.
We only need 13 more
BananaMind
@Banaxi-Tech
On August 3 (preview date) we will be at 90k-100k
Early Access at
BananaMind-Model-Previewers if your known in the community
The benchmarks for 80k are very good
Follow us for the release.
We only need 13 more
@Banaxi-Tech
On August 3 (preview date) we will be at 90k-100k
Early Access at
The benchmarks for 80k are very good
DedeProGamesย
posted an update 7 days ago
Post
1040
๐ Introducing the GRM-3.2 Family
The GRM-3.2 family is a new generation of reasoning-focused models from OrionLLM, purpose-built for long-horizon agentic tasks, extremely difficult reasoning problems, advanced coding, and local AI workflows across a wide range of hardware constraints.
GRM-3.2-Sky is the flagship model in the family: a 35B-A3B Mixture-of-Experts model built on the Ornith-1.0-35B architecture, designed for elite structured reasoning, complex multi-file coding, advanced mathematics, and sustained coherence across extended agentic workflows. It represents a substantial leap in long-horizon task capability over its predecessor, GRM-2.6-Plus.
GRM-3.2-Cliff is the mid-sized workhorse: a 9B-parameter model optimized for long-horizon agentic tasks and difficult reasoning in low-to-mid GPU environments. It delivers strong multi-step planning, debugging, and terminal-agent performance without demanding flagship-level hardware.
GRM-3.2-Turf is the lightweight edge model: a 1.2B-parameter model based on the LiquidAI/LFM2.5-1.2B-Thinking architecture, engineered for efficient on-device execution, high-fidelity instruction following, and robust tool use on mobile, embedded, and other resource-constrained hardware.
All three models are designed for users who need dependable reasoning engines that can maintain goal-directed behavior, planning quality, and task fidelity across many stepsโwhether on a server, a local workstation, or an edge device.
Models:
GRM-3.2-Sky: OrionLLM/GRM-3.2-Sky
GRM-3.2-Cliff: OrionLLM/GRM-3.2-Cliff
GRM-3.2-Turf: OrionLLM/GRM-3.2-Turf
Organization:
OrionLLM
The GRM-3.2 family is a new generation of reasoning-focused models from OrionLLM, purpose-built for long-horizon agentic tasks, extremely difficult reasoning problems, advanced coding, and local AI workflows across a wide range of hardware constraints.
GRM-3.2-Sky is the flagship model in the family: a 35B-A3B Mixture-of-Experts model built on the Ornith-1.0-35B architecture, designed for elite structured reasoning, complex multi-file coding, advanced mathematics, and sustained coherence across extended agentic workflows. It represents a substantial leap in long-horizon task capability over its predecessor, GRM-2.6-Plus.
GRM-3.2-Cliff is the mid-sized workhorse: a 9B-parameter model optimized for long-horizon agentic tasks and difficult reasoning in low-to-mid GPU environments. It delivers strong multi-step planning, debugging, and terminal-agent performance without demanding flagship-level hardware.
GRM-3.2-Turf is the lightweight edge model: a 1.2B-parameter model based on the LiquidAI/LFM2.5-1.2B-Thinking architecture, engineered for efficient on-device execution, high-fidelity instruction following, and robust tool use on mobile, embedded, and other resource-constrained hardware.
All three models are designed for users who need dependable reasoning engines that can maintain goal-directed behavior, planning quality, and task fidelity across many stepsโwhether on a server, a local workstation, or an edge device.
Models:
GRM-3.2-Sky: OrionLLM/GRM-3.2-Sky
GRM-3.2-Cliff: OrionLLM/GRM-3.2-Cliff
GRM-3.2-Turf: OrionLLM/GRM-3.2-Turf
Organization:
Banaxi-Techย
posted an update 7 days ago
Post
2839
We're announcing BananaMind 2 Micro, our smallest model in the BananaMind 2 model family.
This model is not released yet, training has not started yet.
It uses only 2.9M parameters, while being overtrained on 75B tokens to get the maximum intelligence per parameter.
The key changes are:
No more AdamW, the model will use the Muon optimizer offering up to 2x faster convergence and higher lr.
LR goes to 2.2e-2.
We're adding the XSA refresh gate from the TX4 architecture into our own.
Training will start on August 3, release date is estimated to be August 4-6.
On August 3 we will also release our Public Preview of BananaMind 2 Pro.
Follow us to know when our models release
BananaMind
@Banaxi-Tech
This model is not released yet, training has not started yet.
It uses only 2.9M parameters, while being overtrained on 75B tokens to get the maximum intelligence per parameter.
The key changes are:
No more AdamW, the model will use the Muon optimizer offering up to 2x faster convergence and higher lr.
LR goes to 2.2e-2.
We're adding the XSA refresh gate from the TX4 architecture into our own.
Training will start on August 3, release date is estimated to be August 4-6.
On August 3 we will also release our Public Preview of BananaMind 2 Pro.
Follow us to know when our models release
@Banaxi-Tech
Post
2808
Posted an article about training a small MoE from scratch with the efforts of the community!
Model vovaRL/NanoColibri-Instruct
Article https://huggingface.co/blog/vovaRL/hybernation-models-blog
Model vovaRL/NanoColibri-Instruct
Article https://huggingface.co/blog/vovaRL/hybernation-models-blog
Banaxi-Techย
posted an update 8 days ago
Post
3256
Our preview of BananaMind 2 Pro will release on August 3.
Before we do that we want to hit a goal
Lets get 150 followers on my account and 75 on BananaMind!
Me: @Banaxi-Tech
BananaMind:
BananaMind
Full Release on August 10-14
Would really appreciate it!
Before we do that we want to hit a goal
Lets get 150 followers on my account and 75 on BananaMind!
Me: @Banaxi-Tech
BananaMind:
Full Release on August 10-14
Would really appreciate it!
Banaxi-Techย
posted an update 10 days ago
Banaxi-Techย
posted an update 12 days ago
Post
3585
We're excited to release BananaMindBench Leaderboard, our leaderboard for BananaMind Base Bench 1.1.
It measures model performance on a variety of different tasks:
Language Completion
Common sense too
World Knowledge
Context Tracking
Quantitative
Logical Reasoning
Code Completion
Each has a different score and 1 overall score.
Submit your own model:
BananaMind/BananaMindBench-Leaderboard
Check it out:
BananaMind/BananaMindBench-Leaderboard
It measures model performance on a variety of different tasks:
Language Completion
Common sense too
World Knowledge
Context Tracking
Quantitative
Logical Reasoning
Code Completion
Each has a different score and 1 overall score.
Submit your own model:
BananaMind/BananaMindBench-Leaderboard
Check it out:
BananaMind/BananaMindBench-Leaderboard
Banaxi-Techย
posted an update 13 days ago
Post
2864
We're excited to release BananaMind Base Bench 1.1 A new benchmark for base language models with 350 text-completion examples across seven categories. Models are scored using continuation likelihood and receive an Overall Elo score.
Initial results:
BananaMind-2-Medium: 1034
BananaMind-2-Mini: 974
Supra-50M-Base: 973
Supra-1.5-50M-Base-exp: 948
BananaMind-2-Nano: 910
The official script downloads the gated dataset directly from Hugging Face. The dataset is for benchmarking only and may not be used for model training.
BananaMind/BananaMind-Base-Bench-1.1
Initial results:
BananaMind-2-Medium: 1034
BananaMind-2-Mini: 974
Supra-50M-Base: 973
Supra-1.5-50M-Base-exp: 948
BananaMind-2-Nano: 910
The official script downloads the gated dataset directly from Hugging Face. The dataset is for benchmarking only and may not be used for model training.
BananaMind/BananaMind-Base-Bench-1.1
Banaxi-Techย
posted an update 14 days ago
Post
3612
We're excited to announce BananaMind 2V, our small vision model series!
These models are NOT released yet.
We will release them in mid-august!
BananaMind 2V will include:
BananaMind 2V 256M, the flagship based on BananaMind 2 Pro (BananaMind 2 Pro is not released yet).
BananaMind 2V 100M, our mid model, based on BananaMind 2 Medium.
BananaMind 2V 50M, our smallest vision model, based on BananaMind 2 Mini.
These are currently unreleased and will release in mid-august.
Our training will start after BananaMind 2 Pro has finished training.
These models are NOT released yet.
We will release them in mid-august!
BananaMind 2V will include:
BananaMind 2V 256M, the flagship based on BananaMind 2 Pro (BananaMind 2 Pro is not released yet).
BananaMind 2V 100M, our mid model, based on BananaMind 2 Medium.
BananaMind 2V 50M, our smallest vision model, based on BananaMind 2 Mini.
These are currently unreleased and will release in mid-august.
Our training will start after BananaMind 2 Pro has finished training.
Banaxi-Techย
posted an update 16 days ago
Post
2745
We're excited to release BananaMind 2 Medium and BananaMind 2 Medium Chat!
Theyโre both 50M parameter models trained on 50B tokens from FineWeb-Edu, DCLM, Cosmopedia v2, FineMath-4+ and NPSet-2 Python-Edu.
The base model reached 61.86% on PIQA, 43.81% on ARC Easy and 32.43% on HellaSwag. The Chat version was fine-tuned on Smol-SmolTalk and scored 38% overall on our internal instruction benchmark, with 56% on multi-turn, 60% on context recall and 80% on code.
The full details are in the model repos.
Check it out at
BananaMind/BananaMind-2-Medium
BananaMind/BananaMind-2-Medium-Chat
Theyโre both 50M parameter models trained on 50B tokens from FineWeb-Edu, DCLM, Cosmopedia v2, FineMath-4+ and NPSet-2 Python-Edu.
The base model reached 61.86% on PIQA, 43.81% on ARC Easy and 32.43% on HellaSwag. The Chat version was fine-tuned on Smol-SmolTalk and scored 38% overall on our internal instruction benchmark, with 56% on multi-turn, 60% on context recall and 80% on code.
The full details are in the model repos.
Check it out at
BananaMind/BananaMind-2-Medium
BananaMind/BananaMind-2-Medium-Chat
Banaxi-Techย
posted an update 17 days ago
Post
1465
Lets get 80 followers on my account and 30 on BananaMind.
When we hit that we are going to release BananaMind 2 Medium tomorrow.
Me: @Banaxi-Tech
BananaMind:
BananaMind
Early checkpoint shows #1 for <50M on the Open SLM Leaderboard!
Keep Shipping! ๐
When we hit that we are going to release BananaMind 2 Medium tomorrow.
Me: @Banaxi-Tech
BananaMind:
Early checkpoint shows #1 for <50M on the Open SLM Leaderboard!
Keep Shipping! ๐
Banaxi-Techย
posted an update 18 days ago
Post
118
๐ BananaMind 2 Medium is coming soon, in about 2 days.
Release Status BananaMind/BananaMind-2-Medium-Status
Release Status BananaMind/BananaMind-2-Medium-Status
Banaxi-Techย
posted an update 20 days ago
Post
2553
Introducing BananaMind 2 Nano
BananaMind 2 Nano is the smallest member of the BananaMind 2.0 family โ a 10M-parameter language model that shows how much you can squeeze out of a tiny footprint. It uses the family's digit-isolated tokenizer, so it keeps solid arithmetic despite its size, and it's small enough to run just about anywhere.
Trained on 30B tokens in about a day on a single RTX 5070 Ti (16GB), 4096-token context.
Benchmarks:
Average 35.77
ARC Easy 36.20
PIQA 55.98
ARC Challenge 23.38
HellaSwag 27.50
That 35.77 average edges out Pythia-31M (~34.79) at roughly a third the parameters.
Released under Apache 2.0 on Hugging Face: BananaMind/BananaMind-2-Nano โ weights, tokenizer, and config included.
BananaMind 2 Nano is the smallest member of the BananaMind 2.0 family โ a 10M-parameter language model that shows how much you can squeeze out of a tiny footprint. It uses the family's digit-isolated tokenizer, so it keeps solid arithmetic despite its size, and it's small enough to run just about anywhere.
Trained on 30B tokens in about a day on a single RTX 5070 Ti (16GB), 4096-token context.
Benchmarks:
Average 35.77
ARC Easy 36.20
PIQA 55.98
ARC Challenge 23.38
HellaSwag 27.50
That 35.77 average edges out Pythia-31M (~34.79) at roughly a third the parameters.
Released under Apache 2.0 on Hugging Face: BananaMind/BananaMind-2-Nano โ weights, tokenizer, and config included.
Banaxi-Techย
posted an update 21 days ago
Post
149
We're excited to release BananaMind 2 MoE, a new addition to the BananaMind 2 series!
BananaMind 2 MoE is a sparse mixture-of-experts model with 25M total parameters but only 2M active per token, Like the rest of the series, it uses our custom digit-aware BPE tokenizer that keeps every digit isolated, fixing the core arithmetic weakness of our earlier models. It's trained on 30B tokens from FineWeb-Edu, DCLM, Cosmopedia-v2 and FineMath-4+, and outperforms Pythia-31M on average despite activating just 2M parameters.
Check it out at
BananaMind/BananaMind-2-MoE
BananaMind 2 Nano is Coming Next. Already training. Apache 2.0.
BananaMind 2 MoE is a sparse mixture-of-experts model with 25M total parameters but only 2M active per token, Like the rest of the series, it uses our custom digit-aware BPE tokenizer that keeps every digit isolated, fixing the core arithmetic weakness of our earlier models. It's trained on 30B tokens from FineWeb-Edu, DCLM, Cosmopedia-v2 and FineMath-4+, and outperforms Pythia-31M on average despite activating just 2M parameters.
Check it out at
BananaMind/BananaMind-2-MoE
BananaMind 2 Nano is Coming Next. Already training. Apache 2.0.