TX · Utils
LLM Training Data Calculator
How many tokens, words, and Wikipedia-scale articles common GPT-style models need — Chinchilla (~20×) and legacy (~80×) scaling heuristics.
data.tokens(N)
| Model | Params | BF16 | Tokens · 20× | Tokens · 80× | Words · 20× | Wiki arts · 20× | EnWiki · 20× | Books · 20× |
|---|---|---|---|---|---|---|---|---|
| Loading… | ||||||||
References
- Hoffmann et al. (2022) — Training Compute-Optimal Large Language Models (Chinchilla)
- Kaplan et al. (2020) — Scaling Laws for Neural Language Models
- Radford et al. (2019) — Language Models are Unsupervised Multitask Learners (GPT-2)
- Wikipedia: Size of Wikipedia — article and word counts
- Transformer GPU Memory Calculator — BF16 weight memory