LongPIBench: A Long-Context Benchmark for Prompt Injection
Long-context prompt injection defenses look much weaker when the malicious instruction is buried in realistic, lengthy inputs.
The paper introduces LongPIBench, a benchmark spanning peer review, resume screening, code review, and email summarization. Each scenario includes synthetic and real-world datasets, with contexts running from thousands to tens of thousands of tokens. In the authors’ evaluation, even simple heuristic attacks often succeeded and bypassed state-of-the-art defenses. ArXiv · AI/CL/LG's note
The paper introduces LongPIBench, a benchmark spanning peer review, resume screening, code review, and email summarization. Each scenario includes synthetic and real-world datasets, with contexts running from thousands to tens of thousands of tokens. In the authors’ evaluation, even simple heuristic attacks often succeeded and bypassed state-of-the-art defenses. ArXiv · AI/CL/LG's note
score 5