Showed loss falls predictably with model size, data and compute.
Read the paper ↗
Loss falls as a smooth power law in parameters, data and compute, turning model building into a forecastable investment.