ConReq.
AI & ML interests
Recent Activity
Organizations
OOOOO good idea i might try to make it work on cpu
These are the models we are releasing:
- Pebble 10M
- Pebble 25M
- Pebble 50M
Each model will use a Mamba-Transformer 3:1 hybrid architecture and will be pretrained on 25 billion tokens before IFT and SFT.
Depending on development time and resources, we may also release:
- Pebble 5M
- Pebble 75M
- Pebble 1M (possibly)
We hope you're excited and enjoy the models!
Follow for more:
@Hoglet-33
I’ll likely get the 25M models converted tonight or tomorrow AEST
We’re excited to release Pebble-25M and Pebble-25M-Chat!
Both models use our 3:1 Mamba2/Transformer hybrid architecture and were pretrained on 25B tokens. Pebble-25M-Chat was then further fine-tuned on an additional 250M tokens from smol-smoltalk, following the same approach used for the Pebble-10M models.
We hope you enjoy experimenting with them!
Pebble-50M is coming in a few days.
Models
Pebble-25M: basically-ai/Pebble-25M
Pebble-25M-Chat: basically-ai/Pebble-25M-Chat
Pebble-10M GGUFs
In case you missed it, our friend @ContextReq made GGUF versions of the Pebble-10M models:
ContextReq/Pebble-10M-GGUF
ContextReq/Pebble-10M-Chat-GGUF
Follow us if you don’t want to miss future releases and updates!
@Hoglet-33
you know, i'll actually test all of this. i can run some experiments comparing:
Architecture Mamba : Attention
Mamba-only ∞:1
Mostly Mamba 7:1
Current model 3:1
More attention 1:1
Transformer-only 0:1Then we can actually see which ratio performs best at this scale.
As far as I know, I don't think the attention layers were dominating the gradients, but I'll measure that when I run these experiments.
When I get to them, the models will probably be released under https://huggingface.co/basically-experimental
Currently working on testing adamw4 and adafactor aswell as a custom .cu if you want the results, shoot me an email. i got a whole readme of experiment results. I have to rebuild the ternary.cu trainer though lmao