From Acceleration to Saturation: Scaling Behavior of Bootstrapped Language Model Pretraining

2510.06548v1 cs.CL, cs.LG 2025-10-10

Авторы:

Seng Pei Liew, Takuya Kato

Abstract

Bootstrapped pretraining, i.e., the reuse of a pretrained base model for further pretraining, such as continual pretraining or model growth, is promising at reducing the cost of training language models from scratch. However, its effectiveness remains unclear, especially when applied to overtrained base models. In this work, we empirically study the scaling behavior of bootstrapped pretraining and find that its scaling efficiency diminishes in a predictable manner: The scaling exponent with respect to second-stage pretraining tokens decreases logarithmically with the number of tokens used to pretrain the base model. The joint dependence on first- and second-stage tokens is accurately modeled by a simple scaling law. Such saturation effect reveals a fundamental trade-off in multi-stage pretraining strategies: the more extensively a model is pretrained, the less additional benefit bootstrapping provides. Our findings provide practical insights for efficient language model training and raise important considerations for the reuse of overtrained models.

Ссылки и действия

Читать на arXiv Скачать PDF

Дополнительные ресурсы:

From Acceleration to Saturation: Scaling Behavior of Bootstrapped Language Model Pretraining

Авторы:

Abstract

Ссылки и действия

Связанные статьи

Automated Data Enrichment using Confidence-Aware Fine-Grained Debate among Open-...

Exposing Pink Slime Journalism: Linguistic Signatures and Robust Detection Again...

A Preliminary Study on the Promises and Challenges of Native Top-$k$ Sparse Atte...

Computational Linguistics Meets Libyan Dialect: A Study on Dialect Identificatio...

Sarcasm Detection on Reddit Using Classical Machine Learning and Feature Enginee...

Навигация