Study: Improvements in Pre-training Efficiency from 2019 to 2025 Mainly Come from Data

By: www.dwarkesh.com|2026/09/09 09:12:10

The study analyzed the contributions of data and model improvements to pre-training progress from 2019 to 2025. Conducted on a smaller scale, the research utilized publicly available model recipes and data corpora released each year, combining training at a computational scale of up to 1e19 FLOPs. The results indicated that under a 1e19 FLOPs computational budget, the efficiency gain from data improvements was 12.0 times, while model improvements contributed 3.7 times, with the data side gain being approximately 3.24 times that of the model side. The gains from data and model improvements are largely independent, with minimal interaction; under a linear model, the additive effects of both can explain 88% of the variance in OLMES scores. The model side evolved from GPT-2 to OLMo-2, covering optimizers, positional encoding, normalization, activation functions, and initialization, among others. On the data side, the corpus evolved from OpenWebText with about 9 billion tokens in 2019 to larger and more finely filtered datasets like UltraFineWeb by 2025.

This content is provided for general informational purposes only and doesn't constitute financial, investment, legal, or tax advice. Any events, rewards, online promotions, or related information mentioned herein should not be considered a recommendation, solicitation, or invitation to purchase, sell, trade, or otherwise deal in any crypto assets. Crypto assets are highly volatile and may result in loss. The availability of WEEX services, products, and related events may vary by region. You are responsible for ensuring that your participation is in accordance with applicable local laws and regulations.

You may also like

iconiconiconiconiconiconicon
Customer Support:@weikecs
Business Cooperation:@weikecs
Quant Trading & MM:bd@weex.com
VIP Program:support@weex.com