RiskChainBench: A Benchmark for Obfuscated Platform Message Restoration and Evidence-Grounded Web Investigation
RiskChainBench tests whether models can recover hidden abuse instructions and then verify the linked site with evidence.
The benchmark pairs 3,600 synthetic restoration inputs with 600 human-labeled local web environments. Models must first reconstruct the obfuscated message, intent, and destination, then investigate the associated website and produce an evidence-cited risk report. Across ten models, Entry Top-1 ranged from 35.2% to 95.2%, while web decision accuracy ranged from 26.3% to 62.8%. The paper says execution failures made up 31.9% of web runs, pointing to exploration and risk judgment as the main bottlenecks. HF Daily Papers' note
The benchmark pairs 3,600 synthetic restoration inputs with 600 human-labeled local web environments. Models must first reconstruct the obfuscated message, intent, and destination, then investigate the associated website and produce an evidence-cited risk report. Across ten models, Entry Top-1 ranged from 35.2% to 95.2%, while web decision accuracy ranged from 26.3% to 62.8%. The paper says execution failures made up 31.9% of web runs, pointing to exploration and risk judgment as the main bottlenecks. HF Daily Papers' note
score 4