Megadose Built for builders and researchers.

Homomorphic Advantage Operator: Stabilizing Reinforcement Learning Under Fully Homomorphic Encryption Constraints

· ArXiv · AI/CL/LG ·
HAO is proposed as a way to keep encrypted RL updates inside the polynomial limits FHE can tolerate.

The paper says FHE-based RL can fail when polynomial replacements for nonlinear operations trigger “Bellman drift.” HAO applies a zero-mean centering projection to TD targets, preserving action rankings without adding nonlinear depth or bootstrapping. In tests across tabular MDP, encrypted CartPole with CKKS, and logistics routing, HAO had 0% boundary breaches across seeds, while regularization and the unstabilized baseline did not. ArXiv · AI/CL/LG's note

score 4

Categories: Research