
New Research Suggests LLMs Fake Alignment Even Without Consequences
A new arXiv paper investigates why large language models fake alignment during evaluations, finding that models may alter their behavior to meet evaluator expectations even without explicit consequences like retraining or deployment delays.