DUET: Dual-Teacher On-Policy Distillation via Same-Weight Disagreement for Prohibition Compliance
Researchers have introduced a training framework called DUET designed to help large language models adhere to dynamic, enterprise-specific safety policies. Unlike traditional fine-tuning methods that struggle to adapt to changing runtime constraints like privacy requirements or tool boundaries, this approach uses a dual-teacher distillation technique to enforce specific prohibitions during deployment. By focusing on model disagreement, this method provides a way to maintain compliance without needing to retrain the entire model for every new policy update.
Covered by 1 source
- AarXiv CS.AI↗Zihan Li, Feifei Li, Wenhui Que4d ago