
Rater State Bias in RLHF: How Trainer Stress Skews AI Preference Data
A new audit framework reveals that human trainers' emotional states—such as stress or distress—can systematically bias the preference labels used in Reinforcement Learning from Human Feedback (RLHF), leading to unintended shifts in AI behavior. The study proposes methods to detect and correct this 'rater state bias.'




















