Safeguards Based on Copyable Context Cannot Provide Reliable Safety for LLMs
The paper argues that open-access safeguards fail when their evidence of intent can be copied.
For dual-use requests, the authors separate what the model gives from what it can know about downstream use. If an attacker can imitate benign context, they derive a worst-case floor for attacker assistance that remains while useful answers are preserved. That leads to their trilemma: useful capability, reliable safety, and open access cannot all hold together. They point to trusted credentials as hard-to-copy evidence that can strengthen safeguards. HF Daily Papers' note
For dual-use requests, the authors separate what the model gives from what it can know about downstream use. If an attacker can imitate benign context, they derive a worst-case floor for attacker assistance that remains while useful answers are preserved. That leads to their trilemma: useful capability, reliable safety, and open access cannot all hold together. They point to trusted credentials as hard-to-copy evidence that can strengthen safeguards. HF Daily Papers' note
score 5