Skip to post

I wonder what it would mean to release a base model which doesn’t score as well on benchmarks but is more suitable for effective downstream fine-tuning.