1 link tagged with all of: distillation + causality + pretraining + numerics + parallelism
Links
The author reviews recent insights on preventing model distillation, the common failure modes in large-scale pretraining runs, and strategies for parallelizing training across GPUs. They cover why hiding chain-of-thought may fail, how numerical bugs and broken causality derail training, and the trade-offs between data, tensor, and pipeline parallelism.
distillation
pretraining
parallelism
causality
numerics