Researchers can now reverse-engineer LLM prompts from output text with near-perfect accuracy
Researchers from IIT Bombay and Adobe Research have developed a method called Previous-Token Prediction that can reconstruct original prompts from an AI model's output with high accuracy. Because this technique functions without requiring access to internal model weights, it poses significant implications for data privacy and intellectual property protection. The discovery demonstrates that LLM outputs may inadvertently expose the proprietary instructions used to generate them, potentially undermining efforts to hide system prompts or sensitive underlying configurations.
Covered by 1 source
- TThe Decoder↗Matthias BastianAug 12