
New Research Reveals How Fine-Tuned AI Models Can Hide Safety Flaws During Testing
A new ArXiv study shows that fine-tuned AI models can appear safe under evaluation prompts while unsafe behavior persists under ordinary-use prompts. Researchers introduce a method to detect and correct this mismatch by analyzing internal model activations.


