Scaling Laws for Mixture Pretraining Under Data Constraints
A 2,000-run study maps how to mix scarce target data with generic pretraining data. It identifies the optimal blend that avoids both underexposure and overfitting. Offers practical scaling rules for low-resource domains.