research

LLMs Disagree with Humans on Politeness, Study Finds

Summarized by AI from reporting by ArXiv cs.CL, published under our editorial policy.

A new study evaluating seven large language models found that AI models agree more with each other than with humans on what constitutes polite conversation. The research highlights a gap in how LLMs understand social pragmatics, which could impact the quality of AI assistant interactions.

LLMs Disagree with Humans on Politeness, Study Finds

Key takeaways

  • A study on arXiv evaluated seven large language models and found that inter-model agreement on politeness was stronger than model--human agreement.
  • Model--human alignment on politeness judgments was associated with explicit linguistic features like polite words and phrases.
  • The models struggled with nuanced aspects of politeness such as context and tone.
  • Users can provide feedback to AI assistant developers to help improve models' understanding of politeness and social norms.

A new study published on arXiv found that large language models (LLMs) often disagree with human judgments about what's polite in conversations. The models were better at agreeing with each other than with people, suggesting they don't fully understand human social norms. This could affect how AI assistants interact with users in the future.

Study Evaluated Seven LLMs on Two Politeness Datasets

Researchers evaluated seven different LLMs using two English-language datasets. One dataset had continuous human ratings of politeness, while the other used three-way categorical labels (polite, neutral, impolite). The study found that inter-model agreement was stronger than model--human agreement. This means that the models were more consistent with each other than with human judgments.

The study also found that model--human alignment was associated with explicit linguistic features, such as the use of polite words and phrases. However, the models struggled with more nuanced aspects of politeness, such as context and tone.

Why AI Politeness Matters for User Experience

This study highlights the importance of understanding how LLMs evaluate social pragmatics. As AI assistants become more integrated into our daily lives, it's crucial that they can interact with users in a way that feels natural and respectful. If AI assistants can't accurately judge politeness, they may make inappropriate or offensive comments, which could damage user trust and satisfaction.

For example, an AI assistant that doesn't understand the nuances of politeness might respond to a user's request in a way that feels brusque or dismissive. This could lead to frustration and a negative user experience. On the other hand, an AI assistant that understands and can mimic human politeness norms could provide a more pleasant and engaging interaction.

How Users Can Help Improve AI Politeness

If you're using an AI assistant, pay attention to how it responds to your requests. If you notice that the assistant's responses are often impolite or inappropriate, you can provide feedback to the developers. Many AI assistants have feedback mechanisms that allow users to report issues and suggest improvements.

For example, if you're using a popular AI assistant like Siri or Alexa, you can go to the settings menu and look for an option to provide feedback. You can also leave a review in the app store or on the company's website. By providing feedback, you can help developers improve the AI's understanding of politeness and social norms.

Frequently asked

Which models were evaluated in the politeness study?
The source paper does not name the specific seven models evaluated, only that seven different LLMs were tested.
What datasets were used to test politeness in LLMs?
The study used two English-language datasets: one with continuous human ratings of politeness and another with three-way categorical labels (polite, neutral, impolite).
Does this mean AI assistants will always be impolite?
No. The study identifies a current gap in alignment with human norms, but developers can use feedback and further training to improve AI assistants' understanding of politeness over time.