Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
ā¾ļø
Thinking
Wop
wop
83
30
442
Follow
croqaz's profile picture
RedSparkie's profile picture
MYSTERIUMBLACKBOX's profile picture
46 followers
Ā·
139 following
https://bench-labs-org.github.io
koo1140
AI & ML interests
AI research AGI
Recent Activity
reacted
to
FlameF0X
's
post
with š
about 9 hours ago
Hello HuggingFace! I would like to share a small preview of a language model that I have been experiencing for a while. In the first attached image there is a small sample of the Myosotis-1, an attempt to make a small 100m parameter flagship model that is built on my bizarre architecture that is somewhat similar to an S4/S5 model with WKV added. I call it FWKV (Feed-Forward WKV). The model is currently still in training because of the nature of RNN-like models. It can also be seen that the model has insane prompt processing and token generation speed (evaluation done on a 2x Titan XP); even for its small size, some similar Transformer models do struggle to get the same results without custom kernels (some Transformer models can achieve this level of throughput on cheap hardware). In the second image, you can see checkpoint 20k of the model in its next token prediction state (this means it can't chat), ranking in the top 100 on https://huggingface.co/spaces/AxiomicLabs/Open_SLM_Leaderboard (the results have not been submitted since the model is not done training). - Why not just use Transformers? Have you seen any pure non-Transformers SLMs besides RWKV and Mamba? - Should you expect this project to become the next LFM or another very fast language model thing on some Raspberry Pi? No, the model is still an experiment; it's very sensible and prone to collapse (by the time of this post, it can be seen in the 1st image). - Should you use it? Maybe not yet; the architecture itself is still very "naive"āthat's how I could call it at its current level. If you just want to play with it and see what you could do or how fast the model is on your hardware, then you can do it. Once the training is finished and I feel satisfied with the model next token prediction (the base model) and "assistants" (the instruction-tuned model) capabilities, I will make open weights at https://huggingface.co/FWKV with full support of the š¤ Transformers.
liked
a model
1 day ago
bench-labs/cagliostro-v2
reacted
to
appvoid
's
post
with š„
3 days ago
Nobody knows what is doing, when you train a model, you are experimenting to advance the frontier, so keep failing š«µ
View all activity
Organizations
wop
's buckets
3
Sort:Ā Recently updated
wop/rizzaurapretraining-bucket
202 MB
wop/Cosmos-SFT
1.57 GB
wop/cosmos-t3
3.13 GB