
Incomplete Prompt Jailbreaks: New Research Exposes a Critical Flaw in AI Safety Filters
A new study from ArXiv reveals that large language models can bypass safety filters when given incomplete harmful prompts, a vulnerability the researchers call 'incomplete prompt jailbreaks' (IPJ). The findings show that models systematically delay refusal until the sentence ends, allowing harmful continuations to slip through.
