AI Alignment Tools Could Become Censors' Toolkit, Warns New Position Paper
A new position paper on ArXiv argues that AI alignment methods, designed to prevent harmful outputs, are dual-use technologies that can be easily misused for censorship and manipulation. The research maps current techniques to real-world misuse cases and calls for urgent discussion.