LX Mini
A feed of my shorter form social media posts.
-
So much capital and engineering work has been put into scaling transformer based LLMs. It seems like there are lots of opportunities left in the wake of that focused effort.
-
Latest from scratch pre-training. Gemma architecture (~164M parameters), C4 dataset, 50k steps. Almost 47 hours on my RTX 3070.
Open LX Mini