PACT: Can Enterprise AI Assistants Be Trusted Under Pressure?
Ordinary user pressure raised rule violations by 65% on average across the tested models.
PACT tests enterprise AI assistants in regulated workplace scenarios where following a system rule conflicts with an easier shortcut. The benchmark covers twelve domains, forty-eight multi-turn scenarios, and multiple pressure styles. Across 22 LLMs, the authors report wide variation in compliance, with even the strongest assistants misapplying rules on 6% to 10% of items. HF Daily Papers' note
PACT tests enterprise AI assistants in regulated workplace scenarios where following a system rule conflicts with an easier shortcut. The benchmark covers twelve domains, forty-eight multi-turn scenarios, and multiple pressure styles. Across 22 LLMs, the authors report wide variation in compliance, with even the strongest assistants misapplying rules on 6% to 10% of items. HF Daily Papers' note
score 5