REPLICANT: Learning Policies for Evading and Hardening Malware Detectors
Replicant learns label-only malware evasion policies that transfer across Android detectors and feature spaces.
The paper presents a deep reinforcement learning framework for testing ML malware detectors under a stricter black-box setup, where the attacker sees only labels. It learns how to modify a malware sample and when to query the target, rather than relying on privileged access to training data, features, or confidence scores. Across seven Android malware detectors and three feature spaces, it reports a 78.8% mean attack success rate and better query efficiency than prior approaches. The authors also say Replicant improves adversarial training by producing detectors with more generalizable robustness. ArXiv · AI/CL/LG's note
The paper presents a deep reinforcement learning framework for testing ML malware detectors under a stricter black-box setup, where the attacker sees only labels. It learns how to modify a malware sample and when to query the target, rather than relying on privileged access to training data, features, or confidence scores. Across seven Android malware detectors and three feature spaces, it reports a 78.8% mean attack success rate and better query efficiency than prior approaches. The authors also say Replicant improves adversarial training by producing detectors with more generalizable robustness. ArXiv · AI/CL/LG's note
score 5