Robust Critics: New Research Defends LLMs Against Multi-Turn Attacks
A new arXiv study proposes 'Robust Critics' to help AI chatbots distinguish between harmful multi-turn attacks and genuine questions. Current safety systems treat each interaction in isolation, missing gradual shifts in intent. This research could make conversational AI safer and more trustworthy.

Researchers from ArXiv cs.AI published a new study titled "Robust Critics: Defending LLMs Against Multi-Turn Attacks" on July 24, 2026. The paper addresses a central challenge in AI safety: distinguishing between harmful attacks and well-meaning but misunderstood questions. Current safety frameworks apply a contextual bandit treatment, ignoring the trajectory of the conversation. This makes it hard to detect gradual shifts in intent over multiple turns of dialogue, where an attacker's true purpose may only reveal itself slowly across many exchanges.
The research proposes a new approach that accounts for the full conversation history, improving the model's ability to recognize when a user is probing for harmful outputs versus asking legitimate but sensitive questions. This matters because it could make AI chatbots like ChatGPT or Claude safer to use. Imagine if an AI mistakenly believed you were planning something harmful just because you asked a few odd questions. Or worse, imagine if a malicious user could trick the AI into helping with something dangerous by slowly building up to their request. This study aims to prevent both scenarios, making AI interactions more trustworthy for everyone.
If you're curious about how AI safety works, try asking a chatbot like ChatGPT a series of questions that gradually become more sensitive. Notice how it might start to refuse certain requests as the conversation progresses. This is a simple way to see current safety systems in action. Full breakdown → https://www.ainformed.dev