Scaling Laws for Mixture Pretraining Under Data Constraints
Apple researchers have analyzed how language models should balance limited, high-quality data with abundant generic datasets during pretraining. Their findings provide a framework for optimizing model performance when scaling in fields with constrained information, such as specialized technical domains or languages with fewer available resources.
Covered by 1 source
- AApple Machine Learning Blog↗2d ago