← Back to Model Beat
Research·8h ago·all news from September 25, 2026

Self-Play Pretraining with Zero Data

Researchers have introduced a method called Self-Play Pretraining that allows language models to generate their own training data rather than relying on curated external datasets. This approach challenges the prevailing reliance on massive human-collected archives by testing whether models can improve through internal generation. By potentially removing the bottleneck of data scarcity, this technique could shift how foundational models are developed and scaled.

Covered by 1 source

  • AarXiv CS.AI↗Aditya Cowsik, Kfir Dolev, Michael Y. Li, G. Bruno De Luca, Nourya Cohen, Noah D. Goodman, Yoav Levine8h ago

Related stories

ResearchParallel cut research time and cost in half with GPT‑6 AstraSep 22ResearchFrontier Red Team ResearchSep 23ResearchTop AI experts badly underestimated how fast the field is moving, study findsSep 24ResearchLearning to Discover Interesting MathematicsSep 25