Recent
-
It's data time. I've been getting the practice and mechanics of training models off the ground, with just basic eval (slice of generic pre-training data). My intentions are to pre-train and fine-tune toward more specific tasks. So it's time to get a first round of some simple ev…
-
Did some light optimization work for a 30% decrease in the time to train my baseline model. Notes in my logbook: https://codycollier.com/lx/2026/2026-08-01-a-model-training-perf-baseline-and-some-optimizations.html
-
So much capital and engineering work has been put into scaling transformer based LLMs. It seems like there are lots of opportunities left in the wake of that focused effort.
-
Latest from scratch pre-training. Gemma architecture (~164M parameters), C4 dataset, 50k steps. Almost 47 hours on my RTX 3070.
-
Overnight HVAC lends a hand for GPU temp on a long training run :)
-
I wonder what proportion of the training tokens come from scanned books versus the Internet at the frontier training shops.
-
Always Be Training Working out some problems with pre-training a GPT-2 model with C4 dataset on a Google TPU v6e1.
-
Having fun with some of the output of my alien ink library :)