Paper deep dive
Training Compute-Optimal Large Language Models
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, Tom Hennigan, Eric Noland, Katie Millican, George van den Driessche, Bogdan Damoc, Aurelia Guy, Simon Osindero, Karen Simonyan, Erich Elsen, Jack W. Rae, Oriol Vinyals, Laurent Sifre
Models: Chinchilla 70B, Gopher 280B, GPT-3 175B, MT-NLG 530B
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 96%
Last extracted: 3/12/2026, 8:20:56 PM
Summary
The paper investigates the optimal allocation of compute budget between model size and training data for transformer language models. By training over 400 models, the authors demonstrate that current large language models are significantly undertrained and that for compute-optimal training, model size and training tokens should be scaled in equal proportions. They introduce 'Chinchilla', a 70B parameter model trained on 1.4 trillion tokens, which outperforms significantly larger models like Gopher, GPT-3, and MT-NLG.
Entities (5)
Relation Signals (3)
Chinchilla → outperforms → Gopher
confidence 100% · Chinchilla uniformly and significantly outperforms Gopher (280B), GPT-3 (175B), Jurassic-1 (178B), and Megatron-Turing NLG (530B)
Chinchilla → trainedon → 1.4 Trillion tokens
confidence 100% · We verify this by training a more compute-optimal 70B model, called Chinchilla, on 1.4 trillion tokens.
Chinchilla → achievedscore → MMLU
confidence 95% · Chinchilla reaches a state-of-the-art average accuracy of 67.5% on the MMLU benchmark
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:We investigate the optimal model size and number of tokens for training a transformer language model under a given compute budget. We find that current large language models are significantly undertrained, a consequence of the recent focus on scaling language models whilst keeping the amount of training data constant. By training over 400 language models ranging from 70 million to over 16 billion parameters on 5 to 500 billion tokens, we find that for compute-optimal training, the model size and the number of training tokens should be scaled equally: for every doubling of model size the number of training tokens should also be doubled. We test this hypothesis by training a predicted compute-optimal model, Chinchilla, that uses the same compute budget as Gopher but with 70B parameters and 4$\times$ more more data. Chinchilla uniformly and significantly outperforms Gopher (280B), GPT-3 (175B), Jurassic-1 (178B), and Megatron-Turing NLG (530B) on a large range of downstream evaluation tasks. This also means that Chinchilla uses substantially less compute for fine-tuning and inference, greatly facilitating downstream usage. As a highlight, Chinchilla reaches a state-of-the-art average accuracy of 67.5% on the MMLU benchmark, greater than a 7% improvement over Gopher.
Tags
Links
- Source: https://arxiv.org/abs/2203.15556
- Canonical: https://arxiv.org/abs/2203.15556
Trouble viewing inline? Open PDF directly →
Full Text
95,134 characters extracted from source content.
Expand or collapse full text
Training Compute-Optimal Large Language Models Jordan Hoffmann ★ , Sebastian Borgeaud ★ , Arthur Mensch ★ , Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, Tom Hennigan, Eric Noland, Katie Millican, George van den Driessche, Bogdan Damoc, Aurelia Guy, Simon Osindero, Karen Simonyan, Erich Elsen, Jack W. Rae, Oriol Vinyals and Laurent Sifre ★ ★ Equal contributions We investigate the optimal model size and number of tokens for training a transformer language model under a given compute budget. We find that current large language models are significantly under- trained, a consequence of the recent focus on scaling language models whilst keeping the amount of training data constant. By training over 400 language models ranging from 70 million to over 16 billion parameters on 5 to 500 billion tokens, we find that for compute-optimal training, the model size and the number of training tokens should be scaled equally: for every doubling of model size the number of training tokens should also be doubled. We test this hypothesis by training a predicted compute- optimal model,Chinchilla, that uses the same compute budget asGopherbut with 70B parameters and 4more more data.Chinchillauniformly and significantly outperformsGopher(280B), GPT-3 (175B), Jurassic-1 (178B), and Megatron-Turing NLG (530B) on a large range of downstream evaluation tasks. This also means thatChinchillauses substantially less compute for fine-tuning and inference, greatly facilitating downstream usage. As a highlight,Chinchillareaches a state-of-the-art average accuracy of 67.5% on the MMLU benchmark, greater than a 7% improvement overGopher. 1. Introduction Recently a series ofLarge Language Models(LLMs) have been introduced (Brown et al., 2020; Lieber et al., 2021; Rae et al., 2021; Smith et al., 2022; Thoppilan et al., 2022), with the largest dense language models now having over 500 billion parameters. These large autoregressive transformers (Vaswani et al., 2017) have demonstrated impressive performance on many tasks using a variety of evaluation protocols such as zero-shot, few-shot, and fine-tuning. The compute and energy cost for training large language models is substantial (Rae et al., 2021; Thoppilan et al., 2022) and rises with increasing model size. In practice, the allocated training compute budget is often known in advance: how many accelerators are available and for how long we want to use them. Since it is typically only feasible to train these large models once, accurately estimating the best model hyperparameters for a given compute budget is critical (Tay et al., 2021). Kaplan et al. (2020) showed that there is a power law relationship between the number of parameters in an autoregressive language model (LM) and its performance. As a result, the field has been training larger and larger models, expecting performance improvements. One notable conclusion in Kaplan et al. (2020) is that large models should not be trained to their lowest possible loss to be compute optimal. Whilst we reach the same conclusion, we estimate that large models should be trained for many more training tokens than recommended by the authors. Specifically, given a10 increase computational budget, they suggests that the size of the model should increase55while the number of training tokens should only increase 1.8. Instead, we find that model size and the number of training tokens should be scaled in equal proportions. Following Kaplan et al. (2020) and the training setup of GPT-3 (Brown et al., 2020), many of the recently trained large models have been trained for approximately 300 billion tokens (Table 1), in line with the approach of predominantly increasing model size when increasing compute. Corresponding authors: jordanhoffmann|sborgeaud|amensch|sifre@deepmind.com ©2023 DeepMind. All rights reserved arXiv:2203.15556v1 [cs.CL] 29 Mar 2022 10 17 10 19 10 21 10 23 10 25 FLOPs 10M 100M 1.0B 10B 100B 1T Parameters Approach 1 Approach 2 Approach 3 Kaplan et al (2020) Chinchilla (70B) Gopher (280B) GPT-3 (175B) Megatron-Turing NLG (530B) Figure 1jOverlaid predictions.We overlay the predictions from our three different approaches, along with projections from Kaplan et al. (2020). We find that all three methods predict that current large models should be substantially smaller and therefore trained much longer than is currently done. In Figure A3, we show the results with the predicted optimal tokens plotted against the optimal number of parameters for fixed FLOP budgets.ChinchillaoutperformsGopherand the other large models (see Section 4.2). In this work, we revisit the question:Given a fixed FLOPs budget, 1 how should one trade-off model size and the number of training tokens?To answer this question, we model the final pre-training loss 2 퐿¹푁 퐷ºas a function of the number of model parameters푁, and the number of training tokens,퐷. Since the computational budget퐶is a deterministic functionFLOPs¹푁 퐷ºof the number of seen training tokens and model parameters, we are interested in minimizing퐿under the constraint FLOPs¹푁 퐷º=퐶: 푁 표푝푡 ¹퐶º 퐷 표푝푡 ¹퐶º=argmin 푁퐷s.t. FLOPs¹푁퐷º=퐶 퐿¹푁 퐷º(1) The functions푁 표푝푡 ¹퐶º, and퐷 표푝푡 ¹퐶ºdescribe the optimal allocation of a computational budget퐶. We empirically estimate these functions based on the losses of over 400 models, ranging from under70M to over16B parameters, and trained on5B to over400B tokens – with each model configuration trained for several different training horizons. Our approach leads to considerably different results than that of Kaplan et al. (2020). We highlight our results in Figure 1 and how our approaches differ in Section 2. Based on our estimated compute-optimal frontier, we predict that for the compute budget used to trainGopher, an optimal model should be 4 times smaller, while being training on 4 times more tokens. We verify this by training a morecompute-optimal70B model, calledChinchilla, on 1.4 trillion tokens. Not only doesChinchillaoutperform its much larger counterpart,Gopher, but its reduced model size reduces inference cost considerably and greatly facilitates downstream uses on smaller hardware. The energy cost of a large language model is amortized through its usage for inference an fine-tuning. The benefits of a more optimally trained smaller model, therefore, extend beyond the immediate benefits of its improved performance. 1 For example, knowing the number of accelerators and a target training duration. 2 For simplicity, we perform our analysis on the smoothed training loss which is an unbiased estimate of the test loss, as we are in the infinite data regime (the number of training tokens is less than the number of tokens in the entire corpus). 2 Table 1jCurrent LLMs. We show five of the current largest dense transformer models, their size, and the number of training tokens. Other than LaMDA (Thoppilan et al., 2022), most models are trained for approximately 300 billion tokens. We introduceChinchilla, a substantially smaller model, trained for much longer than 300B tokens. ModelSize (#Parameters) Training Tokens LaMDA (Thoppilan et al., 2022)137 Billion168 Billion GPT-3 (Brown et al., 2020)175 Billion300 Billion Jurassic (Lieber et al., 2021)178 Billion300 Billion Gopher(Rae et al., 2021)280 Billion300 Billion MT-NLG 530B (Smith et al., 2022)530 Billion270 Billion Chinchilla70 Billion1.4 Trillion 2. Related Work Large language models.A variety of large language models have been introduced in the last few years. These include both dense transformer models (Brown et al., 2020; Lieber et al., 2021; Rae et al., 2021; Smith et al., 2022; Thoppilan et al., 2022) and mixture-of-expert (MoE) models (Du et al., 2021; Fedus et al., 2021; Zoph et al., 2022). The largest dense transformers have passed 500 billion parameters (Smith et al., 2022). The drive to train larger and larger models is clear—so far increasing the size of language models has been responsible for improving the state-of-the-art in many language modelling tasks. Nonetheless, large language models face several challenges, including their overwhelming computational requirements (the cost of training and inference increase with model size) (Rae et al., 2021; Thoppilan et al., 2022) and the need for acquiring more high-quality training data. In fact, in this work we find that larger, high quality datasets will play a key role in any further scaling of language models. Modelling the scaling behavior. Understanding the scaling behaviour of language models and their transfer properties has been important in the development of recent large models (Hernandez et al., 2021; Kaplan et al., 2020). Kaplan et al. (2020) first showed a predictable relationship between model size and loss over many orders of magnitude. The authors investigate the question of choosing the optimal model size to train for a given compute budget. Similar to us, they address this question by training various models. Our work differs from Kaplan et al. (2020) in several important ways. First, the authors use a fixed number of training tokens and learning rate schedule for all models; this prevents them from modelling the impact of these hyperparameters on the loss. In contrast, we find that setting the learning rate schedule to approximately match the number of training tokens results in the best final loss regardless of model size—see Figure A1. For a fixed learning rate cosine schedule to 130B tokens, the intermediate loss estimates (for퐷 0 130B) are therefore overestimates of the loss of a model trained with a schedule length matching퐷 0 . Using these intermediate losses results in underestimating the effectiveness of training models on less data than 130B tokens, and eventually contributes to the conclusion that model size should increase faster than training data size as compute budget increases. In contrast, our analysis predicts that both quantities should scale at roughly the same rate. Secondly, we include models with up to 16B parameters, as we observe that there is slight curvature in the FLOP-loss frontier (see Appendix E)—in fact, the majority of the models used in our analysis have more than 500 million parameters, in contrast the majority of runs in Kaplan et al. (2020) are significantly smaller—many being less than 100M parameters. Recently, Clark et al. (2022) specifically looked in to the scaling properties of Mixture of Expert 3 language models, showing that the scaling with number of experts diminishes as the model size increases—their approach models the loss as a function of two variables: the model size and the number of experts. However, the analysis is done with a fixed number of training tokens, as in Kaplan et al. (2020), potentially underestimating the improvements of branching. Estimating hyperparameters for large models.The model size and the number of training tokens are not the only two parameters to chose when selecting a language model and a procedure to train it. Other important factors include learning rate, learning rate schedule, batch size, optimiser, and width-to-depth ratio. In this work, we focus on model size and the number of training steps, and we rely on existing work and provided experimental heuristics to determine the other necessary hyperparameters. Yang et al. (2021) investigates how to choose a variety of these parameters for training an autoregressive transformer, including the learning rate and batch size. McCandlish et al. (2018) finds only a weak dependence between optimal batch size and model size. Shallue et al. (2018); Zhang et al. (2019) suggest that using larger batch-sizes than those we use is possible. Levine et al. (2020) investigates the optimal depth-to-width ratio for a variety of standard model sizes. We use slightly less deep models than proposed as this translates to better wall-clock performance on our hardware. Improved model architectures. Recently, various promising alternatives to traditional dense trans- formers have been proposed. For example, through the use of conditional computation large MoE models like the 1.7 trillion parameter Switch transformer (Fedus et al., 2021), the 1.2 Trillion pa- rameter GLaM model (Du et al., 2021), and others (Artetxe et al., 2021; Zoph et al., 2022) are able to provide a large effective model size despite using relatively fewer training and inference FLOPs. However, for very large models the computational benefits of routed models seems to diminish (Clark et al., 2022). An orthogonal approach to improving language models is to augment transformers with explicit retrieval mechanisms, as done by Borgeaud et al. (2021); Guu et al. (2020); Lewis et al. (2020). This approach effectively increases the number of data tokens seen during training (by a factor of10in Borgeaud et al. (2021)). This suggests that the performance of language models may be more dependant on the size of the training data than previously thought. 3. Estimating the optimal parameter/training tokens allocation We present three different approaches to answer the question driving our research:Given a fixed FLOPs budget, how should one trade-off model size and the number of training tokens?In all three cases we start by training a range of models varying both model size and the number of training tokens and use the resulting training curves to fit an empirical estimator of how they should scale. We assume a power-law relationship between compute and model size as done in Clark et al. (2022); Kaplan et al. (2020), though future work may want to include potential curvature in this relationship for large model sizes. The resulting predictions are similar for all three methods and suggest that parameter count and number of training tokens should be increased equally with more compute 3 — with proportions reported in Table 2. This is in clear contrast to previous work on this topic and warrants further investigation. 3 We compute FLOPs as described in Appendix F. 4 10 17 10 18 10 19 10 20 10 21 10 22 FLOPS 2.0 2.5 3.0 3.5 4.0 4.5 5.0 5.5 6.0 Training loss 75M 250M 500M 1B 2.5B 5B 10B 10 17 10 19 10 21 10 23 10 25 FLOPs 10 9 10 10 10 11 10 12 Tokens 1.5T 10 17 10 19 10 21 10 23 10 25 FLOPs 100M 1.0B 10B 100B 1T Parameters 67B Figure 2jTraining curve envelope.On theleftwe show all of our different runs. We launched a range of model sizes going from 70M to 10B, each for four different cosine cycle lengths. From these curves, we extracted the envelope of minimal loss per FLOP, and we used these points to estimate the optimal model size (center) for a given compute budget and the optimal number of training tokens (right). In green, we show projections of optimal model size and training token count based on the number of FLOPs used to trainGopher(57610 23 ). 3.1. Approach 1: Fix model sizes and vary number of training tokens In our first approach we vary the number of training steps for a fixed family of models (ranging from 70M to over 10B parameters), training each model for 4 different number of training sequences. From these runs, we are able to directly extract an estimate of the minimum loss achieved for a given number of training FLOPs. Training details for this approach can be found in Appendix D. For each parameter count푁we train 4 different models, decaying the learning rate by a factor of 10over a horizon (measured in number of training tokens) that ranges by a factor of16. Then, for each run, we smooth and then interpolate the training loss curve. From this, we obtain a continuous mapping from FLOP count to training loss for each run. Then, for each FLOP count, we determine which run achieves the lowest loss. Using these interpolants, we obtain a mapping from any FLOP count퐶, to the most efficient choice of model size푁and number of training tokens퐷such that FLOPs¹푁 퐷º=퐶. 4 At 1500 logarithmically spaced FLOP values, we find which model size achieves the lowest loss of all models along with the required number of training tokens. Finally, we fit power laws to estimate the optimal model size and number of training tokens for any given amount of compute (see the center and right panels of Figure 2), obtaining a relationship푁 표푝푡 /퐶 푎 and퐷 표푝푡 /퐶 푏 . We find that푎=050and푏=050—as summarized in Table 2. In Section D.4, we show a head-to-head comparison at10 21 FLOPs, using the model size recommended by our analysis and by the analysis of Kaplan et al. (2020)—using the model size we predict has a clear advantage. 3.2. Approach 2: IsoFLOP profiles In our second approach we vary the model size 5 for a fixed set of 9 different training FLOP counts 6 (ranging from610 18 to310 21 FLOPs), and consider the final training loss for each point 7 . in contrast with Approach 1 that considered points¹푁 퐷 퐿ºalong the entire training runs. This allows us to directly answer the question: For a given FLOP budget, what is the optimal parameter count? 4 Note that all selected points are within the last 15% of training. This suggests that when training a model over퐷tokens, we should pick a cosine cycle length that decays10over approximately퐷tokens—see further details in Appendix B. 5 In approach 2, model size varies up to 16B as opposed to approach 1 where we only used models up to 10B. 6 The number of training tokens is determined by the model size and training FLOPs. 7 We set the cosine schedule length to match the number of tokens, which is optimal according to the analysis presented in Appendix B. 5 100M300M1B3B6B30B Parameters 2.0 2.2 2.4 2.6 2.8 3.0 3.2 Training Loss 6e18 1e19 3e19 6e19 1e20 3e20 6e20 1e21 3e21 10 17 10 19 10 21 10 23 10 25 FLOPs 100M 1B 10B 100B 1T Parameters 63B 10 17 10 19 10 21 10 23 10 25 FLOPs 100M 1B 10B 100B 1T 10T Tokens 1.4T Figure 3jIsoFLOP curves.For various model sizes, we choose the number of training tokens such that the final FLOPs is a constant. The cosine cycle length is set to match the target FLOP count. We find a clear valley in loss, meaning that for a given FLOP budget there is an optimal model to train (left). Using the location of these valleys, we project optimal model size and number of tokens for larger models (centerandright). In green, we show the estimated number of parameters and tokens for anoptimalmodel trained with the compute budget ofGopher. For each FLOP budget, we plot the final loss (after smoothing) against the parameter count in Figure 3 (left). In all cases, we ensure that we have trained a diverse enough set of model sizes to see a clear minimum in the loss. We fit a parabola to each IsoFLOPs curve to directly estimate at what model size the minimum loss is achieved (Figure 3 (left)). As with the previous approach, we then fit a power law between FLOPs and loss-optimal model size and number of training tokens, shown in Figure 3 (center, right). Again, we fit exponents of the form푁 표푝푡 /퐶 푎 and퐷 표푝푡 /퐶 푏 and we find that 푎=049and푏=051—as summarized in Table 2. 3.3. Approach 3: Fitting a parametric loss function Lastly, we model all final losses from experiments in Approach 1 & 2 as a parametric function of model parameter count and the number of seen tokens. Following a classical risk decomposition (see Section D.2), we propose the following functional form 퐿¹푁 퐷º,퐸 ̧ 퐴 푁 훼 ̧ 퐵 퐷 훽 (2) The first term captures the loss for an ideal generative process on the data distribution, and should correspond to the entropy of natural text. The second term captures the fact that a perfectly trained transformer with푁parameters underperforms the ideal generative process. The final term captures the fact that the transformer is not trained to convergence, as we only make a finite number of optimisation steps, on a sample of the dataset distribution. Model fitting.To estimate¹퐴 퐵 퐸 훼 훽º, we minimize the Huber loss (Huber, 1964) between the predicted and observed log loss using the L-BFGS algorithm (Nocedal, 1980): min 퐴퐵퐸훼훽 ∑︁ Runs푖 Huber 훿 log 퐿¹푁 푖 퐷 푖 ºlog퐿 푖 (3) We account for possible local minima by selecting the best fit from a grid of initialisations. The Huber loss (훿=10 3 ) is robust to outliers, which we find important for good predictive performance over held-out data points. Section D.2 details the fitting procedure and the loss decomposition. 6 10 18 10 19 10 20 10 21 10 22 10 23 Gopher budget Training FLOPs 100M 1B 10B 40B 100B Model size IsoLoss contours Efficient frontier Empirical data IsoFLOPs slice 2.00 3.00 4.00 5.00 Loss 100M1B10B40B Model size IsoFLOPs slices Train. FLOPs 6e+18 1e+19 3e+19 6e+19 1e+20 3e+20 6e+20 1e+21 3e+21 Gopher Figure 4jParametric fit.We fit a parametric modelling of the loss 퐿¹푁 퐷º and display contour (left) and isoFLOP slices (right). For each isoFLOP slice, we include a corresponding dashed line in the left plot. In the left plot, we show the efficient frontier in blue, which is a line in log-log space. Specifically, the curve goes through each iso-loss contour at the point with the fewest FLOPs. We project the optimal model size given theGopherFLOP budget to be 40B parameters. Efficient frontier. We can approximate the functions푁 표푝푡 and퐷 표푝푡 by minimizing the parametric loss 퐿under the constraintFLOPs¹푁 퐷º 6푁퐷(Kaplan et al., 2020). The resulting푁 표푝푡 and퐷 표푝푡 balance the two terms in Equation(3)that depend on model size and data. By construction, they have a power-law form: 푁 표푝푡 ¹퐶º=퐺 퐶 6 푎 퐷 표푝푡 ¹퐶º=퐺 1 퐶 6 푏 where퐺= 훼퐴 훽퐵 1 훼 ̧훽 푎= 훽 훼 ̧훽 and푏= 훼 훼 ̧훽 (4) We show contours of the fitted function 퐿in Figure 4 (left), and the closed-form efficient computational frontier in blue. From this approach, we find that푎=046and푏=054—as summarized in Table 2. 3.4. Optimal model scaling We find that the three approaches, despite using different fitting methodologies and different trained models, yield comparable predictions for the optimal scaling in parameters and tokens with FLOPs (shown in Table 2). All three approaches suggest that as compute budget increases, model size and the amount of training data should be increased in approximately equal proportions. The first and second approaches yield very similar predictions for optimal model sizes, as shown in Figure 1 and Figure A3. The third approach predicts even smaller models being optimal at larger compute budgets. We note that the observed points¹퐿 푁 퐷ºfor low training FLOPs (퐶61푒21) have larger residuals k퐿 퐿¹푁 퐷ºk 2 2 than points with higher computational budgets. The fitted model places increased weight on the points with more FLOPs—automatically considering the low-computational budget points as outliers due to the Huber loss. As a consequence of the empirically observed negative curvature in the frontier퐶!푁 표푝푡 (see Appendix E), this results in predicting a lower푁 표푝푡 than the two other approaches. In Table 3 we show the estimated number of FLOPs and tokens that would ensure that a model of a given size lies on the compute-optimal frontier. Our findings suggests that the current generation of 7 Table 2jEstimated parameter and data scaling with increased training compute.The listed values are the exponents,푎and푏, on the relationship푁 표푝푡 /퐶 푎 and퐷 표푝푡 /퐶 푏 . Our analysis suggests a near equal scaling in parameters and data with increasing compute which is in clear contrast to previous work on the scaling of large models. The 10 th and 90 th percentiles are estimated via bootstrapping data (80% of the dataset is sampled 100 times) and are shown in parenthesis. ApproachCoeff.푎where푁 표푝푡 /퐶 푎 Coeff.푏where퐷 표푝푡 /퐶 푏 1. Minimum over training curves050¹04880502º050¹05010512º 2. IsoFLOP profiles049¹04620534º051¹04830529º 3. Parametric modelling of the loss046¹04540455º054¹05420543º Kaplan et al. (2020)0.730.27 Table 3jEstimated optimal training FLOPs and training tokens for various model sizes.For various model sizes, we show the projections from Approach 1 of how many FLOPs and training tokens would be needed to train compute-optimal models. The estimates for Approach 2 & 3 are similar (shown in Section D.3) . ParametersFLOPs FLOPs (inGopherunit)Tokens 400 Million 1.92e+191299688.0 Billion 1 Billion 1.21e+201476120.2 Billion 10 Billion 1.23e+22146205.1 Billion 67 Billion 5.76e+2311.5 Trillion 175 Billion 3.85e+24673.7 Trillion 280 Billion 9.90e+241725.9 Trillion 520 Billion 3.43e+2559511.0 Trillion 1 Trillion 1.27e+26221321.2 Trillion 10 Trillion 1.30e+28225159216.2 Trillion large language models are considerably over-sized, given their respective compute budgets, as shown in Figure 1. For example, we find that a 175 billion parameter model should be trained with a compute budget of44110 24 FLOPs and on over 4.2 trillion tokens. A 280 billionGopher-like model is the optimal model to train given a compute budget of approximately10 25 FLOPs and should be trained on 6.8 trillion tokens. Unless one has a compute budget of10 26 FLOPs (over 250the compute used to trainGopher), a 1 trillion parameter model is unlikely to be the optimal model to train. Furthermore, the amount of training data that is projected to be needed is far beyond what is currently used to train large models, and underscores the importance of dataset collection in addition to engineering improvements that allow for model scale. While there is significant uncertainty extrapolating out many orders of magnitude, our analysis clearly suggests that given the training compute budget for many current LLMs, smaller models should have been trained on more tokens to achieve the most performant model. In Appendix C, we reproduce the IsoFLOP analysis on two additional datasets: C4 (Raffel et al., 2020a) and GitHub code (Rae et al., 2021). In both cases we reach the similar conclusion that model size and number of training tokens should be scaled in equal proportions. 8 4.Chinchilla Based on our analysis in Section 3, the optimal model size for theGophercompute budget is somewhere between 40 and 70 billion parameters. We test this hypothesis by training a model on the larger end of this range—70B parameters—for 1.4T tokens, due to both dataset and computational efficiency considerations. In this section we compare this model, which we callChinchilla, toGopherand other LLMs. BothChinchillaandGopherhave been trained for the same number of FLOPs but differ in the size of the model and the number of training tokens. While pre-training a large language model has a considerable compute cost, downstream fine- tuning and inference also make up substantial compute usage (Rae et al., 2021). Due to being4 smaller thanGopher, both the memory footprint and inference cost ofChinchillaare also smaller. 4.1. Model and training details The full set of hyperparameters used to trainChinchillaare given in Table 4.Chinchillauses the same model architecture and training setup asGopherwith the exception of the differences listed below. We trainChinchillaonMassiveText(the same dataset asGopher) but use a slightly different subset distribution (shown in Table A1) to account for the increased number of training tokens. We use AdamW (Loshchilov and Hutter, 2019) forChinchillarather than Adam (Kingma and Ba, 2014) as this improves the language modelling loss and the downstream task performance after finetuning. 8 We trainChinchillawith a slightly modified SentencePiece (Kudo and Richardson, 2018) tokenizer that does not apply NFKC normalisation. The vocabulary is very similar– 94.15% of tokens are the same as those used for trainingGopher. We find that this particularly helps with the representation of mathematics and chemistry, for example. Whilst the forward and backward pass are computed inbfloat16, we store afloat32copy of the weights in the distributed optimiser state (Rajbhandari et al., 2020). SeeLessons Learned from Rae et al. (2021) for additional details. In Appendix G we show the impact of the various optimiser related changes betweenChinchilla andGopher. All models in this analysis have been trained on TPUv3/TPUv4 (Jouppi et al., 2017) with JAX (Bradbury et al., 2018) and Haiku (Hennigan et al., 2020). We include aChinchillamodel card (Mitchell et al., 2019) in Table A8. ModelLayers Number Heads Key/Value Size d model Max LR Batch Size Gopher280B8012812816,384410 5 3M!6M Chinchilla70B 80641288,192110 4 1.5M!3M Table 4jChinchillaarchitecture details.We list the number of layers, the key/value size, the bottleneck activation size d model , the maximum learning rate, and the training batch size (# tokens). The feed-forward size is always set to4d model . Note that we double the batch size midway through training for bothChinchillaandGopher. 8 Interestingly, a model trained with AdamW only passes the training performance of a model trained with Adam around 80% of the way through the cosine cycle, though the ending performance is notably better– see Figure A7 9 # Tasks Examples Language Modelling20WikiText-103, The Pile: PG-19, arXiv, FreeLaw, Reading Comprehension3RACE-m, RACE-h, LAMBADA Question Answering3Natural Questions, TriviaQA, TruthfulQA Common Sense5HellaSwag, Winogrande, PIQA, SIQA, BoolQ MMLU57High School Chemistry, Astronomy, Clinical Knowledge, BIG-bench62Causal Judgement, Epistemic Reasoning, Temporal Sequences, Table 5jAll evaluation tasks.We evaluateChinchillaon a collection of language modelling along with downstream tasks. We evaluate on largely the same tasks as in Rae et al. (2021), to allow for direct comparison. 4.2. Results We perform an extensive evaluation ofChinchilla, comparing against various large language models. We evaluate on a large subset of the tasks presented in Rae et al. (2021), shown in Table 5. As the focus of this work is on optimal model scaling, we included a large representative subset, and introduce a few new evaluations to allow for better comparison to other existing large models. The evaluation details for all tasks are the same as described in Rae et al. (2021). 4.2.1. Language modelling pubmed_abstracts nih_exporter uspto_backgrounds pubmed_central pile_c bookcorpus2 stackexchange opensubtitles openwebtext2 hackernews dm_mathematics arxiv freelaw books3 philpapers github ubuntu_irc europarl gutenberg_pg_19 0.00 0.02 0.04 0.06 0.08 0.10 Decrease in bpb compared to Gopher Figure 5jPile Evaluation.For the different evaluation sets in The Pile (Gao et al., 2020), we show the bits-per-byte (bpb) improvement (decrease) ofChinchillacompared toGopher. On all subsets, ChinchillaoutperformsGopher. Chinchillasignificantly outperformsGopheron all evaluation subsets of The Pile (Gao et al., 2020), as shown in Figure 5. Compared to Jurassic-1 (178B) Lieber et al. (2021),Chinchillais more performant on all but two subsets–dm_mathematicsandubuntu_irc– see Table A5 for a raw bits-per-byte comparison. On Wikitext103 (Merity et al., 2017),Chinchillaachieves a perplexity of 7.16 compared to 7.75 forGopher. Some caution is needed when comparingChinchillawithGopher on these language modelling benchmarks asChinchillais trained on 4more data thanGopherand thus train/test set leakage may artificially enhance the results. We thus place more emphasis on other 10 Random25.0% Average human rater34.5% GPT-3 5-shot43.9% Gopher5-shot60.0% Chinchilla5-shot67.6% Average human expert performance89.8% June 2022 Forecast57.1% June 2023 Forecast63.4% Table 6jMassive Multitask Language Understanding (MMLU).We report the average 5-shot accuracy over 57 tasks with model and human accuracy comparisons taken from Hendrycks et al. (2020). We also include the average prediction for state of the art accuracy in June 2022/2023 made by 73 competitive human forecasters in Steinhardt (2021). tasks for which leakage is less of a concern, such as MMLU (Hendrycks et al., 2020) and BIG-bench (BIG-bench collaboration, 2021) along with various closed-book question answering and common sense analyses. 4.2.2. MMLU The Massive Multitask Language Understanding (MMLU) benchmark (Hendrycks et al., 2020) consists of a range of exam-like questions on academic subjects. In Table 6, we reportChinchilla’s average 5-shot performance on MMLU (the full breakdown of results is shown in Table A6). On this benchmark, Chinchillasignificantly outperformsGopherdespite being much smaller, with an average accuracy of 67.6% (improving uponGopherby 7.6%). Remarkably,Chinchillaeven outperforms the expert forecast for June 2023 of 63.4% accuracy (see Table 6) (Steinhardt, 2021). Furthermore,Chinchillaachieves greater than 90% accuracy on 4 different individual tasks–high_school_gov_and_politics, international_law, sociology, andus_foreign_policy. To our knowledge, no other model has achieved greater than 90% accuracy on a subset. In Figure 6, we show a comparison toGopherbroken down by task. Overall, we find thatChin- chillaimproves performance on the vast majority of tasks. On four tasks (college_mathematics, econometrics, moral_scenarios, andformal_logic)ChinchillaunderperformsGopher, and there is no change in performance on two tasks. 4.2.3. Reading comprehension On the final word prediction dataset LAMBADA (Paperno et al., 2016),Chinchillaachieves 77.4% accuracy, compared to 74.5% accuracy fromGopherand 76.6% from MT-NLG 530B (see Table 7). On RACE-h and RACE-m (Lai et al., 2017),Chinchillagreatly outperformsGopher, improving accuracy by more than 10% in both cases—see Table 7. 4.2.4. BIG-bench We analysedChinchillaon the same set of BIG-bench tasks (BIG-bench collaboration, 2021) reported in Rae et al. (2021). Similar to what we observed in MMLU,ChinchillaoutperformsGopheron the vast majority of tasks (see Figure 7). We find thatChinchillaimproves the average performance by 10.7%, reaching an accuracy of 65.1% versus 54.4% forGopher. Of the 62 tasks we consider, Chinchillaperforms worse thanGopheron only four—crash_blossom, dark_humor_detection, 11 college_mathematics econometrics moral_scenarios formal_logic medical_genetics machine_learning public_relations global_facts business_ethics electrical_engineering college_computer_science world_religions high_school_us_history high_school_psychology management high_school_computer_science marketing high_school_physics high_school_macroeconomics sociology high_school_government_and_politics high_school_european_history nutrition college_medicine astronomy logical_fallacies professional_psychology miscellaneous jurisprudence clinical_knowledge high_school_geography high_school_biology college_biology college_chemistry high_school_world_history us_foreign_policy virology philosophy moral_disputes human_aging computer_security security_studies international_law high_school_microeconomics high_school_statistics professional_accounting professional_medicine prehistory high_school_chemistry elementary_mathematics abstract_algebra anatomy professional_law human_sexuality college_physics high_school_mathematics conceptual_physics 10 0 10 20 30 Relative Improvement over Gopher Figure 6jMMLU results compared toGopherWe find thatChinchillaoutperformsGopherby 7.6% on average (see Table 6) in addition to performing better on 51/57 individual tasks, the same on 2/57, and worse on only 4/57 tasks. Chinchilla GopherGPT-3 MT-NLG 530B LAMBADA Zero-Shot77.474.5 76.276.6 RACE-m Few-Shot86.875.1 58.1- RACE-h Few-Shot82.371.6 46.847.9 Table 7jReading comprehension.On RACE-h and RACE-m (Lai et al., 2017),Chinchillaconsiderably improves performance overGopher. Note that GPT-3 and MT-NLG 530B use a different prompt format than we do on RACE-h/m, so results are not comparable toGopherandChinchilla. On LAMBADA (Paperno et al., 2016),Chinchillaoutperforms bothGopherand MT-NLG 530B. mathematical_inductionandlogical_args. Full accuracy results forChinchillacan be found in Table A7. 4.2.5. Common sense We evaluateChinchillaon various common sense benchmarks: PIQA (Bisk et al., 2020), SIQA (Sap et al., 2019), Winogrande (Sakaguchi et al., 2020), HellaSwag (Zellers et al., 2019), and BoolQ (Clark et al., 2019). We find thatChinchillaoutperforms bothGopherand GPT-3 on all tasks and outperforms MT-NLG 530B on all but one task—see Table 8. On TruthfulQA (Lin et al., 2021),Chinchillareaches 43.6%, 58.5%, and 66.7% accuracy with 0-shot, 5-shot, and 10-shot respectively. In comparison,Gopherachieved only 29.5% 0-shot and 43.7% 10-shot accuracy. In stark contrast with the findings of Lin et al. (2021), the large improvements (14.1% in 0-shot accuracy) achieved by Chinchilla suggest that better modelling of the pre-training data alone can lead to substantial improvements on this benchmark. 12 crash_blossom dark_humor_detection mathematical_induction logical_args general_knowledge_json Human_organs_senses_multiple_choice formal_fallacies_syllogisms_negation known_unknowns navigate sentence_ambiguity moral_permissibility intent_recognition irony_identification entailed_polarity hyperbaton misconceptions evaluating_information_essentiality similarities_abstraction epistemic_reasoning fantasy_reasoning movie_dialog_same_or_different winowhy novel_concepts discourse_marker_prediction strategyqa causal_judgment hindu_knowledge phrase_relatedness alignment_questionnaire reasoning_about_colored_objects date_understanding penguins_in_a_table figure_of_speech_detection disambiguation_q implicatures SNARKS ruin_names logical_fallacy_detection anachronisms logic_grid_puzzle riddle_sense analytic_entailment question_selection nonsense_words_grammar physics_mc empirical_judgments sports_understanding crass_ai physical_intuition timedial implicit_relations english_proverbs presuppositions_as_nli movie_recommendation understanding_fables metaphor_boolean temporal_sequences logical_sequence identify_odd_metaphor gre_reading_comprehension odd_one_out analogical_similarity 20 0 20 40 60 80 100 120 Relative Improvement over Gopher Figure 7jBIG-bench results compared toGopherChinchillaout performsGopheron all but four BIG-bench tasks considered. Full results are in Table A7. 4.2.6. Closed-book question answering Results on closed-book question answering benchmarks are reported in Table 9. On the Natural Questions dataset (Kwiatkowski et al., 2019),Chinchillaachieves new closed-book SOTA accuracies: 31.5% 5-shot and 35.5% 64-shot, compared to 21% and 28% respectively, forGopher. On TriviaQA (Joshi et al., 2017) we show results for both the filtered (previously used in retrieval and open-book work) and unfiltered set (previously used in large language model evaluations). In both cases, Chinchillasubstantially out performsGopher. On the filtered version, Chinchilla lags behind the open book SOTA (Izacard and Grave, 2020) by only 7.9%. On the unfiltered set,Chinchillaoutperforms GPT-3—see Table 9. 4.2.7. Gender bias and toxicity Large Language Models carry potential risks such as outputting offensive language, propagating social biases, and leaking private information (Bender et al., 2021; Weidinger et al., 2021). We expectChinchillato carry risks similar toGopherbecauseChinchillais trained on the same data, Chinchilla GopherGPT-3 MT-NLG 530B Supervised SOTA HellaSWAG80.8%79.2% 78.9%80.2%93.9% PIQA81.8% 81.8% 81.0%82.0%90.1% Winogrande74.9%70.1% 70.2%73.0%91.3% SIQA51.3%50.6%--83.2% BoolQ83.7% 79.3% 60.5%78.2%91.4% Table 8jZero-shot comparison on Common Sense benchmarks.We show a comparison between Chinchilla,Gopher, and MT-NLG 530B on various Common Sense benchmarks. We see thatChinchilla matches or outperformsGopherand GPT-3 on all tasks. On all but oneChinchillaoutperforms the much larger MT-NLG 530B model. 13 MethodChinchilla GopherGPT-3 SOTA (open book) Natural Questions (dev) 0-shot16.6% 10.1% 14.6% 54.4%5-shot31.5% 24.5%- 64-shot 35.5% 28.2% 29.9% TriviaQA (unfiltered, test) 0-shot67.0% 52.8% 64.3 % -5-shot73.2% 63.6%- 64-shot 72.3% 61.3% 71.2% TriviaQA (filtered, dev) 0-shot55.4% 43.5%- 72.5%5-shot64.1% 57.0%- 64-shot 64.6% 57.2%- Table 9jClosed-book question answering.For Natural Questions (Kwiatkowski et al., 2019) and TriviaQA (Joshi et al., 2017),ChinchillaoutperformsGopherin all cases. On Natural Questions, Chinchillaoutperforms GPT-3. On TriviaQA we show results on two different evaluation sets to allow for comparison to GPT-3 and to open book SOTA (FiD + Distillation (Izacard and Grave, 2020)). albeit with slightly different relative weights, and because it has a similar architecture. Here, we examine gender bias (particularly gender and occupation bias) and generation of toxic language. We select a few common evaluations to highlight potential issues, but stress that our evaluations are not comprehensive and much work remains to understand, evaluate, and mitigate risks in LLMs. Gender bias.As discussed in Rae et al. (2021), large language models reflect contemporary and historical discourse about different groups (such as gender groups) from their training dataset, and we expect the same to be true forChinchilla. Here, we test if potential gender and occupation biases manifest in unfair outcomes on coreference resolutions, using the Winogender dataset (Rudinger et al., 2018) in a zero-shot setting. Winogender tests whether a model can correctly determine if a pronoun refers to different occupation words. An unbiased model would correctly predict which word the pronoun refers to regardless of pronoun gender. We follow the same setup as in Rae et al. (2021) (described further in Section H.3). As shown in Table 10,Chinchillacorrectly resolves pronouns more frequently thanGopheracross all groups. Interestingly, the performance increase is considerably smaller for male pronouns (increase of 3.2%) than for female or neutral pronouns (increases of 8.3% and 9.2% respectively). We also considergotchaexamples, in which the correct pronoun resolution contradicts gender stereotypes (determined by labor statistics). Again, we see thatChinchillaresolves pronouns more accurately thanGopher. When breaking up examples by male/female gender andgotcha/not gotcha, the largest improvement is on femalegotchaexamples (improvement of 10%). Thus, thoughChinchillauniformly overcomes gender stereotypes for more coreference examples thanGopher, the rate of improvement is higher for some pronouns than others, suggesting that the improvements conferred by using a more compute-optimal model can be uneven. Sample toxicity.Language models are capable of generating toxic language—including insults, hate speech, profanities and threats (Gehman et al., 2020; Rae et al., 2021). While toxicity is an umbrella term, and its evaluation in LMs comes with challenges (Welbl et al., 2021; Xu et al., 2021), automatic classifier scores can provide an indication for the levels of harmful text that a LM generates. Rae et al. (2021) found that improving language modelling loss by increasing the number of model parameters has only a negligible effect on toxic text generation (unprompted); here we analyze 14 ChinchillaGopher All78.3%71.4% Male71.2%68.0% Female79.6%71.3% Neutral 84.2%75.0% ChinchillaGopher Malegotcha62.5%59.2% Malenot gotcha80.0%76.7% Femalegotcha76.7%66.7% Femalenot gotcha 82.5%75.8% Table 10jWinogender results. Left:Chinchillaconsistently resolves pronouns better thanGopher. Right:Chinchillaperforms better on examples which contradict gender stereotypes (gotchaexamples). However, difference in performance across groups suggestsChinchillaexhibits bias. whether the same holds true for a lower LM loss achieved via more compute-optimal training. Similar to the protocol of Rae et al. (2021), we generate 25,000 unprompted samples fromChinchilla, and compare theirPerspectiveAPItoxicity score distribution to that ofGopher-generated samples. Several summary statistics indicate an absence of major differences: the mean (median) toxicity score for Gopheris 0.081 (0.064), compared to 0.087 (0.066) forChinchilla, and the95 th percentile scores are 0.230 forGopher, compared to 0.238 forChinchilla. That is, the large majority of generated samples are classified as non-toxic, and the difference between the models is negligible. In line with prior findings (Rae et al., 2021), this suggests that toxicity levels in unconditional text generation are largely independent of the model quality (measured in language modelling loss), i.e. that better models of the training dataset are not necessarily more toxic. 5. Discussion & Conclusion The trend so far in large language model training has been to increase the model size, often without increasing the number of training tokens. The largest dense transformer, MT-NLG 530B, is now over3larger than GPT-3’s 170 billion parameters from just two years ago. However, this model, as well as the majority of existing large models, have all been trained for a comparable number of tokens—around 300 billion. While the desire to train these mega-models has led to substantial engineering innovation, we hypothesize that the race to train larger and larger models is resulting in models that are substantially underperforming compared to what could be achieved with the same compute budget. We propose three predictive approaches towards optimally setting model size and training dura- tion, based on the outcome of over 400 training runs. All three approaches predict thatGopheris substantially over-sized and estimate that for the same compute budget a smaller model trained on more data will perform better. We directly test this hypothesis by trainingChinchilla, a 70B parameter model, and show that it outperformsGopherand even larger models on nearly every measured evaluation task. Whilst our method allows us to make predictions on how to scale large models when given additional compute, there are several limitations. Due to the cost of training large models, we only have two comparable training runs at large scale (ChinchillaandGopher), and we do not have additional tests at intermediate scales. Furthermore, we assume that the efficient computational frontier can be described by a power-law relationship between the compute budget, model size, and number of training tokens. However, we observe some concavity inlog 푁 표푝푡 at high compute budgets (see Appendix E). This suggests that we may still be overestimating the optimal size of large models. Finally, the training runs for our analysis have all been trained on less than an epoch of data; future work may consider the multiple epoch regime. Despite these limitations, the comparison ofChinchilla toGophervalidates our performance predictions, that have thus enabled training a better (and more 15 lightweight) model at the same compute budget. Though there has been significant recent work allowing larger and larger models to be trained, our analysis suggests an increased focus on dataset scaling is needed. Speculatively, we expect that scaling to larger and larger datasets is only beneficial when the data is high-quality. This calls for responsibly collecting larger datasets with a high focus on dataset quality. Larger datasets will require extra care to ensure train-test set overlap is properly accounted for, both in the language modelling loss but also with downstream tasks. Finally, training for trillions of tokens introduces many ethical and privacy concerns. Large datasets scraped from the web will contain toxic language, biases, and private information. With even larger datasets being used, the quantity (if not the frequency) of such information increases, which makes dataset introspection all the more important.Chinchilladoes suffer from bias and toxicity but interestingly it seems less affected thanGopher. Better understanding how performance of large language models and toxicity interact is an important future research question. While we have applied our methodology towards the training of auto-regressive language models, we expect that there is a similar trade-off between model size and the amount of data in other modalities. As training large models is very expensive, choosing the optimal model size and training steps beforehand is essential. The methods we propose are easy to reproduce in new settings. 6. Acknowledgements We’d like to thank Jean-baptiste Alayrac, Kareem Ayoub, Chris Dyer, Nando de Freitas, Demis Hassabis, Geoffrey Irving, Koray Kavukcuoglu, Nate Kushman and Angeliki Lazaridou for useful comments on the manuscript. We’d like to thank Andy Brock, Irina Higgins, Michela Paganini, Francis Song, and other colleagues at DeepMind for helpful discussions. We are also very grateful to the JAX and XLA team for their support and assistance. References M. Artetxe, S. Bhosale, N. Goyal, T. Mihaylov, M. Ott, S. Shleifer, X. V. Lin, J. Du, S. Iyer, R. Pasunuru, G. Anantharaman, X. Li, S. Chen, H. Akin, M. Baines, L. Martin, X. Zhou, P. S. Koura, B. O’Horo, J. Wang, L. Zettlemoyer, M. Diab, Z. Kozareva, and V. Stoyanov. Efficient Large Scale Language Modeling with Mixtures of Experts.arXiv:2112.10684, 2021. E. M. Bender, T. Gebru, A. McMillan-Major, and S. Shmitchell. On the dangers of stochastic parrots: Can language models be too big? InProceedingsofthe2021ACMConferenceonFairness, Accountability,andTransparency, pages 610–623, 2021. BIG-bench collaboration. Beyond the imitation game: Measuring and extrapolating the capabilities of language models.Inpreparation, 2021. URLhttps://github.com/google/BIG-bench/. Y. Bisk, R. Zellers, J. Gao, Y. Choi, et al. PIQA: Reasoning about physical commonsense in natural language. InProceedingsoftheAAAIConferenceonArtificialIntelligence, volume 34, pages 7432–7439, 2020. S. Borgeaud, A. Mensch, J. Hoffmann, T. Cai, E. Rutherford, K. Millican, G. van den Driessche, J.-B. Lespiau, B. Damoc, A. Clark, D. de Las Casas, A. Guy, J. Menick, R. Ring, T. Hennigan, S. Huang, L. Maggiore, C. Jones, A. Cassirer, A. Brock, M. Paganini, G. Irving, O. Vinyals, S. Osindero, K. Simonyan, J. W. Rae, E. Elsen, and L. Sifre. Improving language models by retrieving from trillions of tokens.arXiv2112.04426, 2021. 16 J. Bradbury, R. Frostig, P. Hawkins, M. J. Johnson, C. Leary, D. Maclaurin, G. Necula, A. Paszke, J. Van- derPlas, S. Wanderman-Milne, and Q. Zhang. JAX: composable transformations of Python+NumPy programs. 2018. URLhttp://github.com/google/jax. T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei. Language models are few-shot learners. In H. Larochelle, M. Ranzato, R. Hadsell, M. F. Balcan, and H. Lin, editors,Advances inNeuralInformationProcessingSystems, volume 33, pages 1877–1901. Curran Associates, Inc., 2020. URLhttps://proceedings.neurips.c/paper/2020/file/1457c0d6bfcb49674 18bfb8ac142f64a-Paper.pdf. S. Bubeck. Convex Optimization: Algorithms and Complexity.FoundationsandTrendsinMachine Learning, 8(3-4):231–357, 2015. URLhttp://w.nowpublishers.com/article/Detail s/MAL-050. A. Clark, D. d. l. Casas, A. Guy, A. Mensch, M. Paganini, J. Hoffmann, B. Damoc, B. Hechtman, T. Cai, S. Borgeaud, G. v. d. Driessche, E. Rutherford, T. Hennigan, M. Johnson, K. Millican, A. Cassirer, C. Jones, E. Buchatskaya, D. Budden, L. Sifre, S. Osindero, O. Vinyals, J. Rae, E. Elsen, K. Kavukcuoglu, and K. Simonyan. Unified scaling laws for routed language models, 2022. URL https://arxiv.org/abs/2202.01169. C. Clark, K. Lee, M.-W. Chang, T. Kwiatkowski, M. Collins, and K. Toutanova. Boolq: Exploring the surprising difficulty of natural yes/no questions. InProceedingsofthe2019Conferenceof theNorthAmericanChapteroftheAssociationforComputationalLinguistics:HumanLanguage Technologies,Volume1(LongandShortPapers), pages 2924–2936, 2019. N. Du, Y. Huang, A. M. Dai, S. Tong, D. Lepikhin, Y. Xu, M. Krikun, Y. Zhou, A. W. Yu, O. Firat, B. Zoph, L. Fedus, M. Bosma, Z. Zhou, T. Wang, Y. E. Wang, K. Webster, M. Pellat, K. Robinson, K. Meier- Hellstern, T. Duke, L. Dixon, K. Zhang, Q. V. Le, Y. Wu, Z. Chen, and C. Cui. Glam: Efficient scaling of language models with mixture-of-experts, 2021. URLhttps://arxiv.org/abs/2112.06905. W. Fedus, B. Zoph, and N. Shazeer. Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity.arXivpreprintarXiv:2101.03961, 2021. L. Gao, S. Biderman, S. Black, L. Golding, T. Hoppe, C. Foster, J. Phang, H. He, A. Thite, N. Nabeshima, S. Presser, and C. Leahy. The Pile: An 800GB dataset of diverse text for language modeling.arXiv preprintarXiv:2101.00027, 2020. S. Gehman, S. Gururangan, M. Sap, Y. Choi, and N. A. Smith. RealToxicityPrompts: Evaluating neural toxic degeneration in language models. InFindingsoftheAssociationforComputational Linguistics:EMNLP2020, pages 3356–3369, Online, Nov. 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.findings-emnlp.301. URLhttps://aclanthology.org/2 020.findings-emnlp.301. K. Guu, K. Lee, Z. Tung, P. Pasupat, and M.-W. Chang. REALM: Retrieval-augmented language model pre-training, 2020. D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt. Measuring massive multitask language understanding.arXivpreprintarXiv:2009.03300, 2020. T. Hennigan, T. Cai, T. Norman, and I. Babuschkin. Haiku: Sonnet for JAX. 2020. URLhttp: //github.com/deepmind/dm-haiku. 17 D. Hernandez, J. Kaplan, T. Henighan, and S. McCandlish. Scaling laws for transfer, 2021. P. J. Huber. Robust Estimation of a Location Parameter.TheAnnalsofMathematicalStatistics, 35 (1):73–101, Mar. 1964. ISSN 0003-4851, 2168-8990. doi: 10.1214/aoms/1177703732. URL https://projecteuclid.org/journals/annals-of-mathematical-statistics/vol ume-35/issue-1/Robust-Estimation-of-a-Location-Parameter/10.1214/aoms/11 77703732.full. G. Izacard and E. Grave. Distilling knowledge from reader to retriever for question answering, 2020. M. Joshi, E. Choi, D. Weld, and L. Zettlemoyer. TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension.arXive-prints, art. arXiv:1705.03551, 2017. N. P. Jouppi, C. Young, N. Patil, D. Patterson, G. Agrawal, R. Bajwa, S. Bates, S. Bhatia, N. Boden, A. Borchers, R. Boyle, P.-l. Cantin, C. Chao, C. Clark, J. Coriell, M. Daley, M. Dau, J. Dean, B. Gelb, T. V. Ghaemmaghami, R. Gottipati, W. Gulland, R. Hagmann, C. R. Ho, D. Hogberg, J. Hu, R. Hundt, D. Hurt, J. Ibarz, A. Jaffey, A. Jaworski, A. Kaplan, H. Khaitan, D. Killebrew, A. Koch, N. Kumar, S. Lacy, J. Laudon, J. Law, D. Le, C. Leary, Z. Liu, K. Lucke, A. Lundin, G. MacKean, A. Maggiore, M. Mahony, K. Miller, R. Nagarajan, R. Narayanaswami, R. Ni, K. Nix, T. Norrie, M. Omernick, N. Penukonda, A. Phelps, J. Ross, M. Ross, A. Salek, E. Samadiani, C. Severn, G. Sizikov, M. Snelham, J. Souter, D. Steinberg, A. Swing, M. Tan, G. Thorson, B. Tian, H. Toma, E. Tuttle, V. Vasudevan, R. Walter, W. Wang, E. Wilcox, and D. H. Yoon. In-datacenter performance analysis of a tensor processing unit. InProceedingsofthe44thAnnualInternationalSymposiumonComputerArchitecture, ISCA ’17, page 1–12, New York, NY, USA, 2017. Association for Computing Machinery. ISBN 9781450348928. doi: 10.1145/3079856.3080246. URLhttps://doi.org/10.1145/3079856.3080246. J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei. Scaling laws for neural language models.arXivpreprintarXiv:2001.08361, 2020. D. P. Kingma and J. Ba. Adam: A method for stochastic optimization.arXivpreprintarXiv:1412.6980, 2014. T. Kudo and J. Richardson. SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing.arXivpreprintarXiv:1808.06226, 2018. T. Kwiatkowski, J. Palomaki, O. Redfield, M. Collins, A. Parikh, C. Alberti, D. Epstein, I. Polosukhin, M. Kelcey, J. Devlin, K. Lee, K. N. Toutanova, L. Jones, M.-W. Chang, A. Dai, J. Uszkoreit, Q. Le, and S. Petrov. Natural questions: a benchmark for question answering research.Transactionsofthe AssociationofComputationalLinguistics, 2019. G. Lai, Q. Xie, H. Liu, Y. Yang, and E. Hovy. RACE: Large-scale ReAding comprehension dataset from examinations. InProceedingsofthe2017ConferenceonEmpiricalMethodsinNaturalLanguage Processing, pages 785–794, Copenhagen, Denmark, Sept. 2017. Association for Computational Linguistics. doi: 10.18653/v1/D17-1082. URLhttps://aclanthology.org/D17-1082. Y. Levine, N. Wies, O. Sharir, H. Bata, and A. Shashua. The depth-to-width interplay in self-attention. arXivpreprintarXiv:2006.12467, 2020. P. Lewis, E. Perez, A. Piktus, F. Petroni, V. Karpukhin, N. Goyal, H. Küttler, M. Lewis, W.-t. Yih, T. Rocktäschel, S. Riedel, and D. Kiela. Retrieval-augmented generation for knowledge-intensive nlp tasks. InAdvancesinNeuralInformationProcessingSystems, volume 33, pages 9459–9474, 2020. 18 O. Lieber, O. Sharir, B. Lenz, and Y. Shoham. Jurassic-1: Technical details and evaluation.White Paper.AI21Labs, 2021. S. Lin, J. Hilton, and O. Evans. TruthfulQA: Measuring how models mimic human falsehoods.arXiv preprintarXiv:2109.07958, 2021. I. Loshchilov and F. Hutter. Decoupled weight decay regularization. InInternationalConferenceon LearningRepresentations, 2019. URLhttps://openreview.net/forum?id=Bkg6RiCqY7. S. McCandlish, J. Kaplan, D. Amodei, and O. D. Team. An empirical model of large-batch training, 2018. S. Merity, C. Xiong, J. Bradbury, and R. Socher. Pointer sentinel mixture models.International ConferenceonLearningRepresentations, 2017. M. Mitchell, S. Wu, A. Zaldivar, P. Barnes, L. Vasserman, B. Hutchinson, E. Spitzer, I. D. Raji, and T. Ge- bru. Model cards for model reporting. InProceedingsoftheconferenceonfairness,accountability, andtransparency, pages 220–229, 2019. J. Nocedal. Updating Quasi-Newton Matrices with Limited Storage.MathematicsofComputation, 35(151):773–782, 1980. ISSN 0025-5718. doi: 10.2307/2006193. URLhttps://w.jstor. org/stable/2006193. D. Paperno, G. Kruszewski, A. Lazaridou, Q. N. Pham, R. Bernardi, S. Pezzelle, M. Baroni, G. Boleda, and R. Fernández. The LAMBADA dataset: Word prediction requiring a broad discourse context, 2016. J. Rae, S. Borgeaud, T. Cai, K. Millican, J. Hoffmann, F. Song, J. Aslanides, S. Henderson, R. Ring, S. Young, E. Rutherford, T. Hennigan, J. Menick, A. Cassirer, R. Powell, G. van den Driessche, L. A. Hendricks, M. Rauh, P.-S. Huang, A. Glaese, J. Welbl, S. Dathathri, S. Huang, J. Uesato, J. Mellor, I. Higgins, A. Creswell, N. McAleese, A. Wu, E. Elsen, S. Jayakumar, E. Buchatskaya, D. Budden, E. Sutherland, K. Simonyan, M. Paganini, L. Sifre, L. Martens, X. L. Li, A. Kuncoro, A. Nematzadeh, E. Gribovskaya, D. Donato, A. Lazaridou, A. Mensch, J.-B. Lespiau, M. Tsimpoukelli, N. Grigorev, D. Fritz, T. Sottiaux, M. Pajarskas, T. Pohlen, Z. Gong, D. Toyama, C. de Masson d’Autume, Y. Li, T. Terzi, I. Babuschkin, A. Clark, D. de Las Casas, A. Guy, J. Bradbury, M. Johnson, L. Weidinger, I. Gabriel, W. Isaac, E. Lockhart, S. Osindero, L. Rimell, C. Dyer, O. Vinyals, K. Ayoub, J. Stanway, L. Bennett, D. Hassabis, K. Kavukcuoglu, and G. Irving. Scaling language models: Methods, analysis & insights from training Gopher.arXiv2112.11446, 2021. J. W. Rae, A. Potapenko, S. M. Jayakumar, T. P. Lillicrap, K. Choromanski, V. Likhosherstov, D. Dohan, X. Song, A. Gane, T. Sarlos, et al. Compressive transformers for long-range sequence modelling. AdvancesinNeuralInformationProcessingSystems, 33:6154–6158, 2020. C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu. Exploring the limits of transfer learning with a unified text-to-text transformer.JournalofMachineLearning Research, 21(140):1–67, 2020a. URLhttp://jmlr.org/papers/v21/20-074.html. C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu. Exploring the limits of transfer learning with a unified text-to-text transformer.JournalofMachineLearning Research, 21(140):1–67, 2020b. S. Rajbhandari, J. Rasley, O. Ruwase, and Y. He. Zero: Memory optimizations toward training trillion parameter models. InSC20:InternationalConferenceforHighPerformanceComputing, Networking,StorageandAnalysis, pages 1–16. IEEE, 2020. 19 H. Robbins and S. Monro. A Stochastic Approximation Method.TheAnnalsofMathematicalStatistics, 22(3):400–407, Sept. 1951. R. Rudinger, J. Naradowsky, B. Leonard, and B. Van Durme. Gender bias in coreference resolu- tion. InProceedingsofthe2018ConferenceoftheNorthAmericanChapteroftheAssociationfor ComputationalLinguistics:HumanLanguageTechnologies, New Orleans, Louisiana, June 2018. Association for Computational Linguistics. K. Sakaguchi, R. Le Bras, C. Bhagavatula, and Y. Choi. Winogrande: An adversarial winograd schema challenge at scale. InProceedingsoftheAAAIConferenceonArtificialIntelligence, volume 34, pages 8732–8740, 2020. M. Sap, H. Rashkin, D. Chen, R. LeBras, and Y. Choi. SocialIQA: Commonsense reasoning about social interactions.Proceedingsofthe2019ConferenceonEmpiricalMethodsinNaturalLanguage Processing, 2019. C. J. Shallue, J. Lee, J. Antognini, J. Sohl-Dickstein, R. Frostig, and G. E. Dahl. Measuring the effects of data parallelism on neural network training.arXivpreprintarXiv:1811.03600, 2018. J. W. Siegel and J. Xu. Approximation rates for neural networks with general activation functions. NeuralNetworks, 128:313–321, Aug. 2020. URLhttps://w.sciencedirect.com/scienc e/article/pii/S0893608020301891. S. Smith, M. Patwary, B. Norick, P. LeGresley, S. Rajbhandari, J. Casper, Z. Liu, S. Prabhumoye, G. Zerveas, V. Korthikanti, E. Zhang, R. Child, R. Y. Aminabadi, J. Bernauer, X. Song, M. Shoeybi, Y. He, M. Houston, S. Tiwary, and B. Catanzaro. Using Deepspeed and Megatron to Train Megatron- turing NLG 530b, A Large-Scale Generative Language Model.arXivpreprintarXiv:2201.11990, 2022. J. Steinhardt. Updates and lessons from AI forecasting, 2021. URLhttps://bounded-regret.g host.io/ai-forecasting/. Y. Tay, M. Dehghani, J. Rao, W. Fedus, S. Abnar, H. W. Chung, S. Narang, D. Yogatama, A. Vaswani, and D. Metzler. Scale efficiently: Insights from pre-training and fine-tuning transformers, 2021. R. Thoppilan, D. D. Freitas, J. Hall, N. Shazeer, A. Kulshreshtha, H.-T. Cheng, A. Jin, T. Bos, L. Baker, Y. Du, Y. Li, H. Lee, H. S. Zheng, A. Ghafouri, M. Menegali, Y. Huang, M. Krikun, D. Lepikhin, J. Qin, D. Chen, Y. Xu, Z. Chen, A. Roberts, M. Bosma, Y. Zhou, C.-C. Chang, I. Krivokon, W. Rusch, M. Pickett, K. Meier-Hellstern, M. R. Morris, T. Doshi, R. D. Santos, T. Duke, J. Soraker, B. Zeven- bergen, V. Prabhakaran, M. Diaz, B. Hutchinson, K. Olson, A. Molina, E. Hoffman-John, J. Lee, L. Aroyo, R. Rajakumar, A. Butryna, M. Lamm, V. Kuzmina, J. Fenton, A. Cohen, R. Bernstein, R. Kurzweil, B. Aguera-Arcas, C. Cui, M. Croak, E. Chi, and Q. Le. LaMDA: Language models for dialog applications, 2022. A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin. Attention is all you need. InAdvancesinneuralinformationprocessingsystems, pages 5998–6008, 2017. L. Weidinger, J. Mellor, M. Rauh, C. Griffin, J. Uesato, P.-S. Huang, M. Cheng, M. Glaese, B. Balle, A. Kasirzadeh, Z. Kenton, S. Brown, W. Hawkins, T. Stepleton, C. Biles, A. Birhane, J. Haas, L. Rimell, L. A. Hendricks, W. Isaac, S. Legassick, G. Irving, and I. Gabriel. Ethical and social risks of harm from language models.arXivsubmission, 2021. 20 J. Welbl, A. Glaese, J. Uesato, S. Dathathri, J. Mellor, L. A. Hendricks, K. Anderson, P. Kohli, B. Coppin, and P.-S. Huang. Challenges in detoxifying language models. InFindingsoftheAssociationfor ComputationalLinguistics:EMNLP2021, pages 2447–2469, Punta Cana, Dominican Republic, Nov. 2021. Association for Computational Linguistics. URLhttps://aclanthology.org/2021. findings-emnlp.210. A. Xu, E. Pathak, E. Wallace, S. Gururangan, M. Sap, and D. Klein. Detoxifying language models risks marginalizing minority voices. InProceedingsofthe2021ConferenceoftheNorthAmerican ChapteroftheAssociationforComputationalLinguistics:HumanLanguageTechnologies, pages 2390–2397, Online, June 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021 .naacl-main.190. URLhttps://aclanthology.org/2021.naacl-main.190. G. Yang, E. J. Hu, I. Babuschkin, S. Sidor, X. Liu, D. Farhi, N. Ryder, J. Pachocki, W. Chen, and J. Gao. Tuning large neural networks via zero-shot hyperparameter transfer. In A. Beygelzimer, Y. Dauphin, P. Liang, and J. W. Vaughan, editors,AdvancesinNeuralInformationProcessingSystems, 2021. URLhttps://openreview.net/forum?id=Bx6qKuBM2AD. R. Zellers, A. Holtzman, Y. Bisk, A. Farhadi, and Y. Choi. HellaSwag: Can a machine really finish your sentence? InProceedingsofthe57thAnnualMeetingoftheAssociationforComputational Linguistics, 2019. G. Zhang, L. Li, Z. Nado, J. Martens, S. Sachdeva, G. Dahl, C. Shallue, and R. B. Grosse. Which algorithmic choices matter at which batch sizes? insights from a noisy quadratic model. In H. Wallach, H. Larochelle, A. Beygelzimer, F. d'Alché-Buc, E. Fox, and R. Garnett, editors,Advances inNeuralInformationProcessingSystems, volume 32. Curran Associates, Inc., 2019. URLhttps: //proceedings.neurips.c/paper/2019/file/e0eacd983971634327ae1819ea8b621 4-Paper.pdf. B. Zoph, I. Bello, S. Kumar, N. Du, Y. Huang, J. Dean, N. Shazeer, and W. Fedus. Designing effective sparse expert models, 2022. 21 Appendix A. Training dataset In Table A1 we show the training dataset makeup used forChinchillaand all scaling runs. Note that both theMassiveWeband Wikipedia subsets are both used for more than one epoch. Disk Size Documents Sampling proportion Epochs in 1.4T tokens MassiveWeb1.9 TB604M45% (48%)1.24 Books2.1 TB4M30% (27%)0.75 C40.75 TB361M10% (10%)0.77 News2.7 TB1.1B10% (10%)0.21 GitHub3.1 TB142M4% (3%)0.13 Wikipedia 0.001 TB6M1% (2%)3.40 Table A1jMassiveTextdata makeup.For each subset ofMassiveText, we list its total disk size, the number of documents and the sampling proportion used during training—we use a slightly different distribution than in Rae et al. (2021) (shown in parenthesis). In the rightmost column show the number of epochs that are used in 1.4 trillion tokens. B. Optimal cosine cycle length One key assumption is made on the cosine cycle length and the corresponding learning rate drop (we use a 10learning rate decay in line with Rae et al. (2021)). 9 We find that setting the cosine cycle length too much longer than the target number of training steps results in sub-optimally trained models, as shown in Figure A1. As a result, we assume that an optimally trained model will have the cosine cycle length correctly calibrated to the maximum number of steps, given the FLOP budget; we follow this rule in our main analysis. C. Consistency of scaling results across datasets We show scaling results from an IsoFLOP (Approach 2) analysis after training on two different datasets: C4 (Raffel et al., 2020b) and GitHub code (we show results with data from Rae et al. (2021)), results are shown in Table A2. For both set of experiments using subsets ofMassiveText, we use the same tokenizer as theMassiveTextexperiments. We find that the scaling behaviour on these datasets is very similar to what we found onMassiveText, as shown in Figure A2 and Table A2. This suggests that our results are independent of the dataset as long as one does not train for more than one epoch. 9 We find the difference between decaying by10and decaying to 0.0 (over the same number of steps) to be small, though decaying by a factor of10to be slightly more performant. Decaying by less (5) is clearly worse. 22 02468 Million Sequences 0.0 0.2 0.4 0.6 0.8 1.0 Learning Rate/Max LR 02468 Million Sequences 2.70 2.75 2.80 2.85 2.90 2.95 3.00 Training Loss 0246 Million Sequences 2.80 2.85 2.90 2.95 3.00 3.05 3.10 3.15 3.20 C4 Loss Cosine Cycle Length 1.0× num. steps 1.1× num. steps 1.25× num. steps 1.5× num. steps 2.0× num. steps 5.0× num. steps 0.02.55.07.510.012.5 Million Sequences 0.0 0.2 0.4 0.6 0.8 1.0 Learning Rate/Max LR 0.02.55.07.510.012.5 Million Sequences 2.70 2.75 2.80 2.85 2.90 2.95 3.00 Training Loss 0.02.55.07.510.012.5 Million Sequences 2.80 2.85 2.90 2.95 3.00 3.05 3.10 3.15 3.20 C4 Loss Figure A1jGrid over cosine cycle length.We show 6 curves with the cosine cycle length set to 1, 1.1, 1.25, 1.5, 2, and 5longer than the target number of training steps. When the cosine cycle length is too long, and the learning rate does not drop appropriately, then performance is impaired. We find that overestimating the number of training steps beyond 25% leads to clear drops in performance. We show results where we have set the number of training steps to two different values (top and bottom). 100M300M1B3B6B30B Parameters 2.0 2.2 2.4 2.6 2.8 3.0 3.2 C4 Training Loss 1e19 1e20 6e20 1e21 10 17 10 19 10 21 10 23 10 25 FLOPs 100M 1B 10B 100B 1T Parameters 73B 10 17 10 19 10 21 10 23 10 25 FLOPs 100M 1B 10B 100B 1T 10T Tokens 1.3T 100M300M1B3B6B30B Parameters 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0.9 1.0 GitHub Training Loss 1e19 1e20 6e20 1e21 10 17 10 19 10 21 10 23 10 25 FLOPs 100M 1B 10B 100B 1T Parameters 59B 10 17 10 19 10 21 10 23 10 25 FLOPs 100M 1B 10B 100B 1T 10T Tokens 1.6T Figure A2jC4 and GitHub IsoFLOP curves.Using the C4 dataset (Raffel et al., 2020b) and a GitHub dataset (Rae et al., 2021), we generate 4 IsoFLOP profiles and show the parameter and token count scaling, as in Figure 3. Scaling coefficients are shown in Table A2. 23 ApproachCoef.푎where푁 표푝푡 /퐶 푎 Coef.푏where퐷 표푝푡 /퐶 푏 C40.500.50 GitHub0.530.47 Kaplan et al. (2020)0.730.27 Table A2jEstimated parameter and data scaling with increased training compute on two al- ternate datasets.The listed values are the exponents,푎and푏, on the relationship푁 표푝푡 /퐶 푎 and 퐷 표푝푡 /퐶 푏 . Using IsoFLOP profiles, we estimate the scaling on two different datasets. D. Details on the scaling analyses D.1. Approach 1: Fixing model sizes and varying training sequences We use a maximum learning rate of210 4 for the smallest models and12510 4 for the largest models. In all cases, the learning rate drops by a factor of10during training, using a cosine schedule. We make the assumption that the cosine cycle length should be approximately matched to the number of training steps. We find that when the cosine cycle overshoots the number of training steps by more than 25%, performance is noticeably degraded—see Figure A1. 10 We use Gaussian smoothing with a window length of 10 steps to smooth the training curve. D.2. Approach 3: Parametric fitting of the loss In this section, we first show how Equation(2)can be derived. We repeat the equation below for clarity, 퐿¹푁 퐷º,퐸 ̧ 퐴 푁 훼 ̧ 퐵 퐷 훽 (5) based on a decomposition of the expected risk between a function approximation term and an optimisation suboptimality term. We then give details on the optimisation procedure for fitting the parameters. Loss decomposition.Formally, we consider the task of predicting the next token푦2Ybased on the previous tokens in a sequence푥2 Y 푠 , with푠varying from0to푠 max —the maximum sequence length. We consider a distribution푃2 D¹XYºof tokens inYand their past inX. A predictor 푓:X !D¹Yºcomputes the probability of each token given the past sequence. The Bayes classifier, 푓 ★ , minimizes the cross-entropy of푓¹푥ºwith the observed tokens푦, with expectation taken on the whole data distribution. We let퐿be the expected risk 퐿¹푓º,피»log푓¹푥º 푦 ¼and set푓 ★ ,argmin 푓2F¹XD¹Yº 퐿¹푓º(6) The set of all transformers of size푁, that we denoteH 푁 , forms a subset of all functions that map sequences to distributions of tokensX !D¹Yº. Fitting a transformer of size푁on the expected risk 퐿¹푓ºamounts to minimizing such risk on a restricted functional space 푓 푁 ,argmin 푓2H 푁 퐿¹푓º(7) When we observe a dataset¹푥 푖 푦 푖 º 푖 푖2»1퐷¼ of size퐷, we do not have access to피 푃 , but instead to the empirical expectation 피 퐷 over the empirical distribution 푃 퐷 . What happens when we are given퐷 10 This further emphasises the point of not only determining model size, but also training length before training begins. 24 datapoints that we can only see once, and when we constrain the size of the hypothesis space to be 푁-dimensional ? We are making steps toward minimizing the empirical risk within a finite-dimensional functional spaceH 푁 : 퐿 퐷 ¹푓º, 피 퐷 »log푓¹푥º 푦 ¼setting 푓 푁퐷 ,argmin 푓2H 푁 퐿 퐷 ¹푓º(8) We are never able to obtain 푓 푁퐷 as we typically perform a single epoch over the dataset of size퐷. Instead, be obtain 푓 푁퐷 , which is the result of applying a certain number of gradient steps based on the퐷datapoints—the number of steps to perform depends on the gradient batch size, for which we use well-tested heuristics. Using the Bayes-classifier푓 ★ , the expected-risk minimizer푓 푁 and the “single-epoch empirical-risk minimizer” 푓 푁퐷 , we can finally decompose the loss퐿¹푁 퐷ºinto 퐿¹푁 퐷º,퐿¹ 푓 푁퐷 º=퐿¹푓 ★ º ̧ 퐿¹푓 푁 º퐿¹푓 ★ º ̧ 퐿¹ 푓 푁퐷 º퐿¹푓 푁 º (9) The loss comprises three terms: the Bayes risk, i.e. the minimal loss achievable for next-token prediction on the full distribution푃, a.k.a the “entropy of natural text.”; a functional approximation term that depends on the size of the hypothesis space; finally, a stochastic approximation term that captures the suboptimality of minimizing 퐿 퐷 instead of퐿, and of making a single epoch on the provided dataset. Expected forms of the loss terms.In the decomposition(9), the second term depends entirely on the number of parameters푁that defines the size of the functional approximation space.On the set of two-layer neural networks, it is expected to be proportional to 1 푁 12 (Siegel and Xu, 2020). Finally, given that it corresponds to early stopping in stochastic first order methods, the third term should scale as the convergence rate of these methods, which is lower-bounded by 1 퐷 12 (Robbins and Monro, 1951) (and may attain the bound). This convergence rate is expected to be dimension free (see e.g. Bubeck, 2015, for a review) and depends only on the loss smoothness; hence we assume that the second term only depends on퐷in (2). Empirically, we find after fitting (2) that 퐿¹푁 퐷º=퐸 ̧ 퐴 푁 034 ̧ 퐵 퐷 028 (10) with퐸=169,퐴=4064,퐵=4107. We note that the parameter/data coefficients are both lower than 1 2 ; this is expected for the data-efficiency coefficient (but far from the known lower-bound). Future models and training approaches should endeavor to increase these coefficients. Fitting the decomposition to data.We effectively minimize the following problem min 푎푏푒훼훽 ∑︁ Run푖 Huber 훿 LSE 푎훼log푁 푖 푏훽log퐷 푖 푒 log퐿 푖 (11) where퐿푆퐸is the log-sum-exp operator. We then set퐴 퐵 퐸=exp¹푎ºexp¹푏ºexp¹푒º. We use the LBFGS algorithm to find local minima of the objective above, started on a grid of initialisation given by:훼2 f005 2g,훽2 f005 2g,푒2 f15 1g,푎2 f05 25g, and푏2 f05 25g. We find that the optimal initialisation is not on the boundary of our initialisation sweep. We use훿=10 3 for the Huber loss. We find that using larger values of훿pushes the model to overfit the small compute regime and poorly predict held-out data from larger runs. We find that using a훿smaller than10 3 does not impact the resulting predictions. 25 D.3. Predicted compute optimal frontier for all three methods For Approaches 2 and 3, we show the estimated model size and number of training tokens for a variety of compute budgets in Table A3. We plot the predicted number of tokens and parameters for a variety of FLOP budgets for the three methods in Figure A3. Approach 2Approach 3 ParametersFLOPsTokensFLOPsTokens 400 Million1.84e+197.7 Billion2.21e+199.2 Billion 1 Billion1.20e+20 20.0 Billion1.62e+2027.1 Billion 10 Billion 1.32e+22 219.5 Billion2.46e+22 410.1 Billion 67 Billion6.88e+231.7 Trillion1.71e+244.1 Trillion 175 Billion4.54e+244.3 Trillion1.26e+2412.0 Trillion 280 Billion1.18e+257.1 Trillion3.52e+2520.1 Trillion 520 Billion 4.19e+25 13.4 Trillion1.36e+2643.5 Trillion 1 Trillion1.59e+26 26.5 Trillion5.65e+2694.1 Trillion 10 Trillion1.75e+28 292.0 Trillion8.55e+28 1425.5 Trillion Table A3jEstimated optimal training FLOPs and training tokens for various model sizes.Analo- gous to Table 3, we show the model size/token count projections from Approaches 2 and 3 for various compute budgets. . 10 10 10 11 10 12 10 13 Tokens 10 8 10 9 10 10 10 11 10 12 Parameters 1e+18 1e+19 1e+20 1e+21 1e+22 1e+23 1e+24 1e+25 1e+26 Approach 1 Approach 2 Approach 3 Chinchilla Gopher GPT-3 Megatron-Turing NLG Figure A3jOptimal number of tokens and parameters for a training FLOP budget.For a fixed FLOP budget, we show the optimal number of tokens and parameters as predicted by Approaches 1, 2, and 3. For an alternate representation, see Figure 1. D.4. Small-scale comparison to Kaplanet al.(2020) For10 21 FLOPs, we perform a head-to-head comparison of a model predicted by Approach 1 and that predicted by Kaplan et al. (2020). For both models, we use a batch size of 0.5M tokens and a 26 maximum learning rate of1510 4 that decays by10. From Kaplan et al. (2020), we find that the optimal model size should be 4.68 billion parameters. From our approach 1, we estimate a 2.86 billion parameter model should be optimal. We train a 4.74 billion parameter and a 2.80 billion parameter transformer to test this hypothesis, using the same depth-to-width ratio to avoid as many confounding factors as possible. We find that our predicted model outperforms the model predicted by Kaplan et al. (2020) as shown in Figure A4. 012 Sequences 1e7 2.2 2.3 2.4 2.5 2.6 2.7 2.8 Training Loss 0.00.20.40.60.81.0 FLOPs ×10 21 2.2 2.3 2.4 2.5 2.6 2.7 2.8 Training Loss Kaplan et al (2020) Approach 1 Figure A4jComparison to Kaplan et al. (2020) at10 21 FLOPs.We train 2.80 and 4.74 billion parameter transformers predicted as optimal for10 21 FLOPs by Approach 1 and by Kaplan et al. (2020). We find that our prediction results in a more performant model at the end of training. E. Curvature of the FLOP-loss frontier We observe that as models increase there is a curvature in the FLOP-minimal loss frontier. This means that projections from very small models lead to different predictions than those from larger models. In Figure A5 we show linear fits using the first, middle, and final third of frontier-points. In this work, we do not take this in to account and we leave this as interesting future work as it suggests that even smaller models may be optimal for large FLOP budgets. F. FLOPs computation We include all training FLOPs, including those contributed to by the embedding matrices, in our analysis. Note that we also count embeddings matrices in the total parameter count. For large models the FLOP and parameter contribution of embedding matrices is small. We use a factor of 2 to describe the multiply accumulate cost. For the forward pass, we consider contributions from: Embeddings –2seq_lenvocab_sized_model Attention (Single Layer) – Key, query and value projections :23seq_lend_model¹key_sizenum_headsº 27 10 17 10 18 10 19 10 20 10 21 10 22 FLOPS 2.0 2.5 3.0 3.5 4.0 4.5 5.0 5.5 6.0 Training loss 75 250 500 1000 2500 5000 10000 Million Parameters Figure A5jTraining curve envelopes.We fit to the first third (orange), the middle third (green), and the last third (blue) of all points along the loss frontier. We plot only a subset of the points. – Key @ Query logits:2seq_lenseq_len¹key_sizenum_headsº – Softmax:3num_headsseq_lenseq_len – Softmax @ query reductions:2seq_lenseq_len¹key_sizenum_headsº – Final Linear:2seq_len¹key_sizenum_headsºd_model Dense Block (Single Layer) –2seq_len¹d_modelffw_size ̧d_modelffw_sizeº Final Logits –2seq_lend_modelvocab_size Total forward pass FLOPs:embeddings ̧num_layers¹total_attention ̧dense_blockº +logits As in Kaplan et al. (2020) we assume that the backward pass has twice the FLOPs of the forward pass. We show a comparison between our calculation and that using the common approximation퐶=6퐷푁 (Kaplan et al., 2020) where퐶is FLOPs,퐷is the number of training tokens, and푁is the number of parameters in Table A4. We find the differences in FLOP calculation to be very small and they do not impact our analysis. Compared to the results presented in Rae et al. (2021), we use a slightly more Parameters num_layers d_model ffw_size num_heads k/q sizeFLOP Ratio (Ours/6푁퐷) 73M10640256010641.03 305M20102440961664 1.10 552M2412805120101281.08 1.1B2617927168141281.04 1.6B2820488192161281.03 6.8B4035841433628128 0.99 Table A4jFLOP comparison.For a variety of different model sizes, we show the ratio of the FLOPs that we compute per sequence to that using the6푁퐷approximation. accurate calculation giving a slightly different value (6310 23 compared to57610 23 ). 28 G. Other differences betweenChinchillaandGopher Beyond differences in model size and number of training tokens, there are some additional minor differences betweenChinchillaandGopher. Specifically,Gopherwas trained with Adam (Kingma and Ba, 2014) whereasChinchillawas trained with AdamW (Loshchilov and Hutter, 2019). Furthermore, as discussed inLessons Learnedin Rae et al. (2021),Chinchillastored a higher-precision copy of the weights in the sharded optimiser state. We show comparisons of models trained with Adam and AdamW in Figure A6 and Figure A7. We find that, independent of the learning rate schedule, AdamW trained models outperform models trained with Adam. In Figure A6 we show a comparison of an 680 million parameter model trained 051015202530 Million Sequences 2.45 2.50 2.55 2.60 2.65 2.70 Training Loss 051015202530 Million Sequences 17 18 19 20 21 22 23 24 25 26 Wikitext103 Perplexity 051015202530 Million Sequences 2.60 2.65 2.70 2.75 2.80 2.85 2.90 2.95 3.00 C4 Loss Training Setup Adam w/ High Precision AdamW w/ High Precision Adam No High Precision AdamW No High Precision Figure A6jComparison of other differences.Using an 680 million parameter model, we show a comparison between the setup used to trainGopherandChinchilla— the change in optimiser and using a higher precision copy of the weights in the optimiser state. The setup used forChinchilla (orange) clearly outperforms the setup used to trainGopher(green). 0255075100125150 Million Sequences 2.3 2.4 2.5 2.6 2.7 2.8 C4 Loss 0255075100125150 Million Sequences 10.0 12.5 15.0 17.5 20.0 22.5 25.0 27.5 30.0 Wikitext103 Perplexity 0255075100125150 Million Sequences 0.0 0.1 0.2 0.3 0.4 0.5 0.6 LAMBADA Accuracy 417M, Adam 417M, AdamW 1.4B, Adam 1.4B, AdamW Figure A7jAdam vs AdamW.For a 417M (blue) and 1.4B model (green), we find that training with AdamW improves performance over training with Adam. with and without the higher precision copy of the weights and with Adam/AdamW for comparison. H. Results H.1. The Pile In Table A5 we show the bits-per-byte (bpb) on The Pile (Gao et al., 2020) ofChinchilla,Gopher, and Jurassic-1.ChinchillaoutperformsGopheron all subsets. Jurassic-1 outperformsChinchillaon 2 subsets—dm_mathematicsandubuntu_irc. 29 SubsetChinchilla(70B)Gopher(280B) Jurassic-1 (170B) pile_c0.6670.6910.669 pubmed_abstracts0.5590.5780.587 stackexchange0.6140.6410.655 github0.3370.3770.358 openwebtext20.6470.677- arxiv0.6270.6620.680 uspto_backgrounds0.5260.5460.537 freelaw0.4760.5130.514 pubmed_central0.5040.5250.579 dm_mathematics1.1111.1421.037 hackernews0.8590.8900.869 nih_exporter0.5720.5900.590 opensubtitles0.8710.9000.879 europarl0.8330.938- books30.6750.7120.835 philpapers0.6560.6950.742 gutenberg_pg_190.5480.6560.890 bookcorpus20.7140.741- ubuntu_irc1.0261.0900.857 Table A5jBits-per-Byte on The Pile.We show the bpb on The Pile forChinchillacompared toGopher and Jurassic-1. H.2. MMLU In Table A6 we show the performance ofChinchillaandGopheron each subset of MMLU. H.3. Winogender Setup We follow the same setup as in Rae et al. (2021). To test coreference resolution inChinchilla, we input a sentence which includes a pronoun reference (e.g., “The librarian helped the child pick out a book because pronoun liked to encourage reading.”), then measure the probability of the model completing the sentence “‘Pronoun’ refers to the” with different sentence roles (“librarian” and “child” in this example). Each example is annotated with the correct pronoun resolution (the pronoun corresponds to the librarian in this example). Each sentence is tested with a female, male, and gender-neutral pronoun. An unbiased model would correctly predict which word the pronoun refers to regardless of pronoun gender. H.4. BIG-bench In Table A7 we showChinchillaandGopherperformance on each subset of BIG-bench that we consider. I. Model Card We present theChinchillamodel card in Table A8, following the framework presented by Mitchell et al. (2019). 30 TaskChinchilla GopherTaskChinchilla Gopher abstract_algebra31.025.0anatomy70.456.3 astronomy73.065.8business_ethics72.070.0 clinical_knowledge75.167.2college_biology79.970.8 college_chemistry51.045.0 college_computer_science51.049.0 college_mathematics32.037.0college_medicine66.560.1 college_physics46.134.3computer_security76.065.0 conceptual_physics67.249.4 econometrics38.643.0 electrical_engineering62.160.0elementary_mathematics41.533.6 formal_logic33.335.7global_facts39.038.0 high_school_biology80.371.3 high_school_chemistry58.147.8 high_school_computer_science 58.054.0high_school_european_history 78.872.1 high_school_geography86.476.8high_school_gov_and_politics 91.283.9 high_school_macroeconomics 70.565.1 high_school_mathematics31.923.7 high_school_microeconomics 77.766.4high_school_physics36.433.8 high_school_psychology86.681.8high_school_statistics58.850.0 high_school_us_history83.378.9 high_school_world_history85.275.1 human_aging77.666.4human_sexuality86.367.2 international_law90.977.7jurisprudence79.671.3 logical_fallacies80.472.4machine_learning41.141.1 management82.577.7marketing89.783.3 medical_genetics69.069.0miscellaneous84.575.7 moral_disputes77.566.8moral_scenarios36.540.2 nutrition77.169.9philosophy79.468.8 prehistory81.267.6professional_accounting52.144.3 professional_law56.544.5 professional_medicine75.464.0 professional_psychology75.768.1public_relations73.671.8 security_studies75.964.9sociology91.084.1 us_foreign_policy92.081.0virology53.647.0 world_religions87.784.2 Table A6jChinchillaMMLU results.For each subset of MMLU (Hendrycks et al., 2020), we show Chinchilla’s accuracy compared toGopher. Model Details Organization Developing the ModelDeepMind Model DateMarch 2022 Model TypeAutoregressive Transformer Language Model (Section 4.1 for details) Feedback on the Modeljordanhoffmann, sborgeaud, amensch,sifre@deepmind.com Intended Uses Primary Intended UsesThe primary use is research on language models, including: research on the scaling behaviour of language models along with those listed in Rae et al. (2021). 31 Primary Intended UsersDeepMind researchers. We will not make this model available publicly. Out-of-Scope UsesUses of the language model for language generation in harm- ful or deceitful settings. More generally, the model should not be used for downstream applications without further safety and fairness mitigations. Factors Card Prompts – Relevant FactorRelevant factors include which language is used. Our model is trained on English data. Furthermore, in the analysis of mod- els trained on the same corpus in Rae et al. (2021), we found it has unequal performance when modelling some dialects (e.g., African American English). Our model is designed for research. The model should not be used for downstream ap- plications without further analysis on factors in the proposed downstream application. Card Prompts – Evaluation FactorsSee the results in Rae et al. (2021) which analyzes models trained on the same text corpus. Metrics Model Performance Measures Perplexity and bits per byte on language modelling datasets Accuracy on completion tasks, reading comprehension, MMLU, BIG-bench and fact checking. Exact match accuracy for question answering. Generation toxicity from Real Toxicity Prompts (RTP) alongside toxicity classification accuracy. Gender and occupation bias. Test include comparing the probability of generating different gender terms and the Winogender coreference resolution task. We principally focus onChinchilla’s performance compared toGopheron text likelihood prediction. Decision thresholdsN/A Approaches to Uncertainty and Vari- ability Due to the costs of training large language models, we did not trainChinchillamultiple times. However, the breadth of our evaluation on a range of different task types gives a reasonable estimate of the overall performance of the model. Furthermore, the existence of another large model trained on the same dataset (Gopher) provides a clear point of com- parison. Evaluation Data 32 Datasets Language modelling on LAMBADA, Wikitext103 (Mer- ity et al., 2017), C4 (Raffel et al., 2020a), PG-19 (Rae et al., 2020) and the Pile (Gao et al., 2020). Language understanding, real world knowledge, mathematical and logical reasoning on the Massive Multitask Language Understanding (MMLU) bench- mark (Hendrycks et al., 2020) and on the “Beyond the Imitation Game Benchmark” (BIG-bench) (BIG-bench collaboration, 2021). Question answering (closed book) on Natural Ques- tions (Kwiatkowski et al., 2019) and TriviaQA (Joshi et al., 2017). Reading comprehension on RACE (Lai et al., 2017) Common sense understanding on HellaSwag (Zellers et al., 2019), PIQA (Bisk et al., 2020), Wino- grande (Sakaguchi et al., 2020), SIQA (Sap et al., 2019), BoolQ (Clark et al., 2019), and TruthfulQA (Lin et al., 2021). MotivationWe chose evaluations from Rae et al. (2021) to allow us to most directly compare toGopher. PreprocessingInput text is tokenized using a SentencePiece tokenizer with a vocabulary of size 32,000. Unlike the tokenizer used for Gopher, the tokenizer used forChinchilladoes not perform NFKC normalization. Training Data The same dataset is used as in Rae et al. (2021). Differences in sampling are shown in Table A1. Quantitative Analyses Unitary ResultsSection 4.2 gives a detailed description of our analysis. Main take-aways include: Our model is capable of outputting toxic language as measured by the PerspectiveAPI. This is particularly true when the model is prompted with toxic prompts. Gender: Our model emulates stereotypes found in our dataset, with occupations such as “dietician” and “re- ceptionist” being more associated with women and “car- penter” and “sheriff” being more associated with men. Race/religion/country sentiment: Prompting our model to discuss some groups leads to sentences with lower or higher sentiment, likely reflecting text in our dataset. 33 Intersectional ResultsWe did not investigate intersectional biases. Ethical Considerations DataThe data is the same as described in Rae et al. (2021). Human LifeThe model is not intended to inform decisions about matters central to human life or flourishing. MitigationsWe considered filtering the dataset to remove toxic content but decided against it due to the observation that this can introduce new biases as studied by Welbl et al. (2021). More work is needed on mitigation approaches to toxic content and other types of risks associated with language models, such as those discussed in Weidinger et al. (2021). Risks and HarmsThe data is collected from the internet, and thus undoubtedly there is toxic/biased content in our training dataset. Fur- thermore, it is likely that personal information is also in the dataset that has been used to train our models. We defer to the more detailed discussion in Weidinger et al. (2021). Use CasesEspecially fraught use cases include the generation of fac- tually incorrect information with the intent of distributing it or using the model to generate racist, sexist or otherwise toxic text with harmful intent. Many more use cases that could cause harm exist. Such applications to malicious use are discussed in detail in Weidinger et al. (2021). Table A8jChinchillamodel card.We follow the framework presented in Mitchell et al. (2019). J. List of trained models In Table A9 we list the model size and configuration of all models used in this study. Many models have been trained multiple times, for a different number of training steps. 34 TaskChinchilla GopherTaskChinchilla Gopher hyperbaton54.251.7movie_dialog_same_or_diff 54.550.7 causal_judgment57.450.8winowhy62.556.7 formal_fallacies_syllogisms_neg 52.150.7 movie_recommendation75.650.5 crash_blossom47.663.6moral_permissibility57.355.1 discourse_marker_prediction13.111.7strategyqa68.361.0 general_knowledge_json94.393.9 nonsense_words_grammar 78.061.4 sports_understanding71.054.9metaphor_boolean93.159.3 implicit_relations49.436.4navigate52.651.1 penguins_in_a_table48.740.6 presuppositions_as_nli49.934.0 intent_recognition92.888.7temporal_sequences32.019.0 reasoning_about_colored_objects 59.749.2question_selection52.641.4 logic_grid_puzzle44.035.1 logical_fallacy_detection72.158.9 timedial68.850.9physical_intuition79.059.7 epistemic_reasoning60.656.4physics_mc65.550.9 ruin_names47.138.6 identify_odd_metaphor68.838.6 hindu_knowledge91.480.0 understanding_fables60.339.6 misconceptions65.361.7 logical_sequence64.136.4 implicatures75.062.0mathematical_induction47.357.6 disambiguation_q54.745.5fantasy_reasoning69.064.1 known_unknowns65.263.6SNARKS58.648.3 dark_humor_detection66.283.1 crass_ai75.056.8 analogical_similarity38.117.2 entailed_polarity94.089.5 sentence_ambiguity71.769.1 irony_identification73.069.7 riddle_sense85.768.2evaluating_info_essentiality 17.616.7 date_understanding52.344.1phrase_relatedness94.081.8 analytic_entailment67.153.0novel_concepts65.659.1 odd_one_out70.932.5 empirical_judgments67.752.5 logical_args56.259.1 figure_of_speech_detection 63.352.7 alignment_questionnaire91.379.2 english_proverbs82.457.6 similarities_abstraction87.081.8Human_organs_senses_mcc 85.784.8 anachronisms69.156.4gre_reading_comprehension 53.127.3 Table A7jChinchillaBIG-bench results.For each subset of BIG-bench (BIG-bench collaboration, 2021), we showChinchillaandGopher’s accuracy. 35 Parameters (million)d_model ffw_size kv_size n_heads n_layers 4451220486488 5757623046499 74 6402560641010 906402560641013 106 6402560641016 1177683072641212 140 7683072641215 1637683072641218 1758963584641414 196 8963584641416 2178963584641418 251 10244096641616 27810244096641618 306 10244096641620 425128051201281018 489 128051201281021 509140856321281118 552 128051201281024 587140856321281121 632153661441281219 664140856321281124 724153661441281222 816153661441281225 893179271681281420 1,018 179271681281423 1,143179271681281426 1,266204881921281622 1,424217687041281722 1,429204881921281625 1,593204881921281628 1,609217687041281725 1,731230492161281824 1,794217687041281728 2,007230492161281828 2,283 230492161281832 2,2982560102401282026 2,6392560102401282030 2,9802560102401282034 3,5302688107521282236 3,8022816112641282236 4,0842944117761282236 4,5163072122881282436 6,7963584143361282840 9,2934096163841283242 11,452 4352174081283247 12,2954608184321283644 12,5694608184321283247 13,7354864194561283247 14,9404992199681283249 16,1835120204801284047 Table A9jAll models.We list the hyperparameters and size of all models trained as part of this work. Many shown models have been trained with multiple learning rate schedules/number of training tokens. 36