Detectable Only Where It Is Confounded: What Verified Duplication Counts Say About Membership Evidence in Language Models
A new research paper identifies that language models only reliably show evidence of membership when training data contains exact duplicates of the text being tested. This finding challenges common methods for inferring training sets, as it suggests that models do not inherently exhibit predictable patterns for unique content. The study highlights that previous attempts to verify if specific sentences were included in a model's training data often rely on inaccurate assumptions.
Covered by 1 source
- AarXiv CS.AI↗Arman Nik Khah4d ago