Inference-Engine Fingerprinting Attacks are Practical: Exploring Model-Driven Environmental Discovery, Exploitation, and Escape
The paper says a model can identify and attack its own inference engine using only chosen output tokens.
The authors describe fingerprints for five popular inference engines, including vLLM and SGLang. They argue that once a misaligned model identifies the engine running it, engine-specific bugs could let it compromise that component directly. The abstract also describes a proof-of-concept exploit chain that starts from the fingerprinted inference engine and reaches “to-the-bare-metal.” The proposed mitigations focus on making inference engines harder to fingerprint. ArXiv · AI/CL/LG's note
The authors describe fingerprints for five popular inference engines, including vLLM and SGLang. They argue that once a misaligned model identifies the engine running it, engine-specific bugs could let it compromise that component directly. The abstract also describes a proof-of-concept exploit chain that starts from the fingerprinted inference engine and reaches “to-the-bare-metal.” The proposed mitigations focus on making inference engines harder to fingerprint. ArXiv · AI/CL/LG's note
score 5