AI Output Homogeneity Traced to Pretraining, Not Just Alignment, New Study Finds
A new ArXiv study finds that semantic convergence in large language models begins during the pretraining phase, not just the alignment process. The research shows that output homogeneity is observed from the first alignment stage (instruction-tuning/SFT), suggesting it is learned early and only magnified later. This challenges the common assumption that diversity loss is primarily an alignment problem.
