Making Open-Source Text LLM Watermarks Durable Against Merging
Researchers have introduced a new technique to maintain the effectiveness of watermarks in open-source large language models even after the models undergo weight merging or fine-tuning. By integrating watermarking directly into the model parameters, this method prevents the common issue where post-training modifications inadvertently strip away the digital signals used to identify AI-generated text. This advancement provides a more reliable way to track content provenance as collaborative development and model editing become more frequent in the open-source community.
Covered by 1 source
- AarXiv CS.AI↗Luisa Scharff, Thibaud Gloaguen, Robin Staab, Martin VechevJul 24