From my experience, modern machine learning models don't scale well <i>down</i> at all, and it's almost certainly better to just use a simple Markov chain variant of some sort - like Niall on the Amiga, or whatever Terry Pratchett used to come up with Foul Ol' Ron's catchphrase, "Millennium hand and shrimp".
> The model weights and inference code need to be contained within 25KB of user-space memory<p>Wouldn’t it be era-appropriate to allow relying on banked memory? You’d still need to hold the inference code, but you could effectively stream(ing page) the weights as you compute on them.
From a look at <a href="https://en.wikipedia.org/wiki/BBC_Micro#Specifications" rel="nofollow">https://en.wikipedia.org/wiki/BBC_Micro#Specifications</a> I think the 6502 versions of the beeb didn't have banked RAM so to keep it loadable from tape the limits might be as stated.<p>But with substantial additional effort, maybe some banked ROMs could be added..?
Cool to think this demo would have been possible over fifty years ago. I wonder what someone from 1975 would have said if you had shown this to them back then.
This is super cool!
As someone who's worked a little with NES programming and tried out cc65, I'm surprised he didn't just hand write some assembly, he likely couldve saved a lot of space if I had to guess.
The 6502 is notoriously unfit for a C compiler, so probably there is room for more performance in the future. :)
The biggest win for AI dev efficiency is cutting down what gets loaded into context. Semantically matching tasks to the top tools helps a lot.
This is amazing project! I hope it will result in real miniaturization of AI - for example, edge LLM inside of glasses. That will be awesome.
really great work
[flagged]