yano
AI & ML interests
Recent Activity
Organizations
23 thinking modes? What?
Suprisingly smart
๐ค We trained a ~20k-parameter Transformer that can actually write stories!
raincandy-u/MacroStories
โ ~50ร smaller than the 1M-parameter TinyStories model
โ ~3,000ร smaller than AlexNet
โ 81 KB in FP32
yayyy the whole model. เซฎ หถแต แต แตหถ แ
She has a 32-dimensional hidden state, a 378-token vocabulary, and just one decoder block โ recurrently applied 4 times with shared weights.
Despite having only 19,969 parameters, she can maintain a simple narrative across 100โ300 words: establish a goal, encounter a problem, take relevant actions, and reach an outcome.
She runs extremely fast on CPU โ no GPU required. The entire model is tiny enough to load almost instantly! โบ๏ธ
Very Promising!
@Banaxi-Tech oh ok. I donโt think anyone uses Em dashes. Like our school donโt even teach them
I use them only when OCRing text that uses an Mdash. Otherwise i use a hyphen.
Also you can likely remove mdashes by using the bias to remove or make it 90% less likely.
Heimdallr is now my standart :)
First time comment user here
I wonder how useful this would be for me.
Doing a few different projects.
Generating images for cards that lack an image (RP cards, think Character Tavern, Character Hub or SillyTavern), and a lot of images have messups (like 10%-20%, the usual, screwed up hands, perspective where a bike is behind them 20 feet yet their hand is resting on the bike, extra limbs, etc) enough to qualify a regen. Wonder if this would do for finding those. Likely with the description maybe also an answer between 1-10 of how much it follows the prompt.
Manga optimization. I need to identify color, black/white, dithered and grayscale. I have optimized scripts to get really good compression for each case. But if they are in the wrong set they either take up a lot more space, or get really bad downsampling results. A handful of cases the page is just text, so identifying that to OCR instead of keeping separate images.
Text alignment. I got a couple OCR methods where paragraphs don't always get newlines, which needs done beforehand. Could i use it to compare against the OCR to identify which line(s) in the input should have extra spacing added?
A small demo built on imajev-4b: a closet stylist ๐
Request you to star it here so we can make it better - https://github.com/mohit67890/imajev.
Tap a piece and it reads the photo (red 75%, checked 99%), then the app picks bottoms, shoes and a bag from your own closet in the colours you like. Change your colours and the outfit changes.
Under the hood it's one request with one photo and 4 typed questions. Every option gets a probability, so the app applies its rules (one pattern per outfit) and ranks what's left. About 1.1 s per outfit on a Mac (MLX). Every % in the video is the model's real answer.
Also, thanks to @zenmagnets for the FP8 version for Blackwell GPUs: same calibration, 0.1 pt less accuracy on all 23,900 DecisionBench rows, ~1.6ร faster ๐
zenmagnets/Imajev-4B-FP8-SM120
๐ง Weights: mohit67890/imajev-4b
๐ Demo: mohit67890/imajev
๐ Leaderboard: https://benchmarkheaven.com/image-jev-bench
I find this too complex
I remember playing the Gamecube version back in.... ~2004, using triggers to get it to connect and then using gravity to speed up before switching connections to keep it going...
Ahh back when life was less confusing.