Pretraining
23.8 trillion tokens from the web, public sources and licensed datasets, with heavy emphasis on code, technical writing and STEM. About 95% of raw web tokens were filtered out. Run on 6,144 NVIDIA GB300 NVL72 GPUs in under four weeks, with 92.3% goodput near the end.