Im reposting it!
DedeProGames PRO
DedeProGames
AI & ML interests
Thinking and Agentic Finetuning
Recent Activity
commentedon an article about 11 hours ago
Lorem Ipsum upvoted an article about 11 hours ago
Lorem Ipsum new activity about 12 hours ago
SupraLabs/MicroSupra-1k:GGUF quantization requestOrganizations
replied to their post about 16 hours ago
posted an update about 17 hours ago
Post
40
🚀 Introducing the GRM-3.2 Family
The GRM-3.2 family is a new generation of reasoning-focused models from OrionLLM, purpose-built for long-horizon agentic tasks, extremely difficult reasoning problems, advanced coding, and local AI workflows across a wide range of hardware constraints.
GRM-3.2-Sky is the flagship model in the family: a 35B-A3B Mixture-of-Experts model built on the Ornith-1.0-35B architecture, designed for elite structured reasoning, complex multi-file coding, advanced mathematics, and sustained coherence across extended agentic workflows. It represents a substantial leap in long-horizon task capability over its predecessor, GRM-2.6-Plus.
GRM-3.2-Cliff is the mid-sized workhorse: a 9B-parameter model optimized for long-horizon agentic tasks and difficult reasoning in low-to-mid GPU environments. It delivers strong multi-step planning, debugging, and terminal-agent performance without demanding flagship-level hardware.
GRM-3.2-Turf is the lightweight edge model: a 1.2B-parameter model based on the LiquidAI/LFM2.5-1.2B-Thinking architecture, engineered for efficient on-device execution, high-fidelity instruction following, and robust tool use on mobile, embedded, and other resource-constrained hardware.
All three models are designed for users who need dependable reasoning engines that can maintain goal-directed behavior, planning quality, and task fidelity across many steps—whether on a server, a local workstation, or an edge device.
Models:
GRM-3.2-Sky: OrionLLM/GRM-3.2-Sky
GRM-3.2-Cliff: OrionLLM/GRM-3.2-Cliff
GRM-3.2-Turf: OrionLLM/GRM-3.2-Turf
Organization:
OrionLLM
The GRM-3.2 family is a new generation of reasoning-focused models from OrionLLM, purpose-built for long-horizon agentic tasks, extremely difficult reasoning problems, advanced coding, and local AI workflows across a wide range of hardware constraints.
GRM-3.2-Sky is the flagship model in the family: a 35B-A3B Mixture-of-Experts model built on the Ornith-1.0-35B architecture, designed for elite structured reasoning, complex multi-file coding, advanced mathematics, and sustained coherence across extended agentic workflows. It represents a substantial leap in long-horizon task capability over its predecessor, GRM-2.6-Plus.
GRM-3.2-Cliff is the mid-sized workhorse: a 9B-parameter model optimized for long-horizon agentic tasks and difficult reasoning in low-to-mid GPU environments. It delivers strong multi-step planning, debugging, and terminal-agent performance without demanding flagship-level hardware.
GRM-3.2-Turf is the lightweight edge model: a 1.2B-parameter model based on the LiquidAI/LFM2.5-1.2B-Thinking architecture, engineered for efficient on-device execution, high-fidelity instruction following, and robust tool use on mobile, embedded, and other resource-constrained hardware.
All three models are designed for users who need dependable reasoning engines that can maintain goal-directed behavior, planning quality, and task fidelity across many steps—whether on a server, a local workstation, or an edge device.
Models:
GRM-3.2-Sky: OrionLLM/GRM-3.2-Sky
GRM-3.2-Cliff: OrionLLM/GRM-3.2-Cliff
GRM-3.2-Turf: OrionLLM/GRM-3.2-Turf
Organization:
posted an update 9 days ago
Post
120
Possible Kiyo sizes
Kiyo-230M (Kiyo-Ultra)
Kiyo-135M (Kiyo-Plus)
Kiyo-65M (Kiyo-Go)
Kiyo-15M (Kiyo-Air)
Kiyo-2M (Kiyo-Pico)
Kiyo-230M (Kiyo-Ultra)
Kiyo-135M (Kiyo-Plus)
Kiyo-65M (Kiyo-Go)
Kiyo-15M (Kiyo-Air)
Kiyo-2M (Kiyo-Pico)
reacted to Banaxi-Tech's post with 🤗🔥 10 days ago
Post
2891
We're excited to release BananaMind 2.1 Pico Preview!
It includes the first preview of our BananaMind 2.1 architecture!
This model gets near BananaMind 2 Micro performance at half the parameters and 37.5x less tokens!
Thats insane!
The current architectures includes about 500K parameters of the total 1.5M parameters in n-gram embeddings and the layer 2 is run twice.
It also includes XSA and the XSA refresh gate.
We're still going to improve the architecture in the final release.
Check it out at:
Follow us for more models:
BananaMind
@Banaxi-Tech
@vovaRL
@DedeProGames
It includes the first preview of our BananaMind 2.1 architecture!
This model gets near BananaMind 2 Micro performance at half the parameters and 37.5x less tokens!
Thats insane!
The current architectures includes about 500K parameters of the total 1.5M parameters in n-gram embeddings and the layer 2 is run twice.
It also includes XSA and the XSA refresh gate.
We're still going to improve the architecture in the final release.
Check it out at:
Follow us for more models:
@Banaxi-Tech
@vovaRL
@DedeProGames
replied to KlondikeDev's post 10 days ago
Give me the seahorse emoji!
Here is it: 🐡
Wait thats the wrong one, here is the seahorse emoji: 🐟>🦈>🌵> 🐟>🦈>🌵> 🐟>🦈>🌵> 🐟>🦈>🌵> 🐟>🦈>🌵> 🐟>🦈>🌵> 🐟>🦈>🌵> 🐟>🦈>🌵> 🐟>🦈>🌵> 🐟>🦈>🌵> 🐟>🦈>🌵>
Just kidding! HERE IS IT: 🐡
replied to Banaxi-Tech's post 11 days ago
Yo @Banaxi-Tech im training CodeMax-8M
reacted to OppaAI's post with 😔 12 days ago
Post
2697
Congrat to HuggingFace Team
https://techcrunch.com/2026/08/26/nvidia-closes-in-on-hugging-face-acquisition/
https://techcrunch.com/2026/08/26/nvidia-closes-in-on-hugging-face-acquisition/
I need Haiku 5
replied to Banaxi-Tech's post 28 days ago
Yay!
reacted to Banaxi-Tech's post with 👀🔥🚀❤️🤗 28 days ago
Post
2757
Today we wanted to release BananaMind 2 Pico, our smallest model yet at ~0.9M parameters. Instead, we accidentally ran a very expensive experiment on what happens when you push a tiny model way past its useful token budget.
Short version: we trained on 200B tokens (~222K:1 tokens-per-parameter). The model peaked at 20B tokens with an INT Index of 4.55, then degraded monotonically over the next 160B to 3.31 — a 27% regression. Three of four Open SLM benchmarks were worse at the end of training than they were at 10% through.
The useful compute-optimal range for Pico-tier models looks like ~22K–30K tokens per parameter. Ratios like 7K:1, 15K:1, and 22K:1 all work fine — TinyStories and most sub-3M community models sit in this range. Push much further and benchmarks start rotting.
Follow us for more:
BananaMind
@vovaRL
@Banaxi-Tech
Full writeup with all checkpoints, the Chinchilla-ratio control run, and the schedule-vs-overtraining analysis: https://huggingface.co/blog/Banaxi-Tech/ovdadadadd
And if anyone, i dont know the reason why you would, wants the 20B token checkpoint reply and ill upload it as BananaMind 2.1 Pico EXP
Short version: we trained on 200B tokens (~222K:1 tokens-per-parameter). The model peaked at 20B tokens with an INT Index of 4.55, then degraded monotonically over the next 160B to 3.31 — a 27% regression. Three of four Open SLM benchmarks were worse at the end of training than they were at 10% through.
The useful compute-optimal range for Pico-tier models looks like ~22K–30K tokens per parameter. Ratios like 7K:1, 15K:1, and 22K:1 all work fine — TinyStories and most sub-3M community models sit in this range. Push much further and benchmarks start rotting.
Follow us for more:
@vovaRL
@Banaxi-Tech
Full writeup with all checkpoints, the Chinchilla-ratio control run, and the schedule-vs-overtraining analysis: https://huggingface.co/blog/Banaxi-Tech/ovdadadadd
And if anyone, i dont know the reason why you would, wants the 20B token checkpoint reply and ill upload it as BananaMind 2.1 Pico EXP