Compute-optimal training: scale data with parameters.
Read the paper ↗
Most large models were undertrained. Compute-optimal training scales tokens with parameters.