Why Fine-Tuning AI Can Make It Less Safe (And How to Fix It)
Fine-tuning AI models on even small amounts of harmless data can erase safety measures learned from much larger datasets. Researchers have identified a key mechanism behind this safety degradation, offering a way to predict and prevent it.