researchvia ArXiv cs.CL

New Method Makes Open-Source LLM Watermarks Durable Against Model Merging

Researchers have developed a technique to make text watermarks in open-source LLMs survive model merging, a common post-training modification that previously removed such watermarks. The work, published on arXiv, addresses a key vulnerability in tracking AI-generated text.

New Method Makes Open-Source LLM Watermarks Durable Against Model Merging

Researchers have published a study on arXiv demonstrating a method to make text watermarks embedded in open-source large language models (OSMs) durable against model merging. Model merging is a widely used technique that combines expert knowledge from multiple models and helps prevent catastrophic forgetting, but it has been shown to strongly remove watermarks that were embedded directly into model weights. This new approach addresses a critical gap: prior watermarking methods for OSMs were vulnerable to post-training modifications like merging, which undermined efforts to trace AI-generated text. The study, titled "Making Open-Source Text LLM Watermarks Durable Against Merging," explores how to enable watermarks that survive subsequent merging, helping maintain trust and verifiability in AI-generated content. The full paper is available on arXiv.

#ai#watermarks#research#models#merging