
AI Models Still Make Irrational Decisions Even When Aligned
Researchers found that AI models can still make irrational decisions even when trained to align with human values. This 'rational value risk' means models may not always choose the best possible action, even if they understand what's valuable.






















