Variance-Optimal Off-Policy Evaluation with Conjunct Effect Modeling
VOCEM is presented as an unbiased bridge between cluster-weighted OffCEM and action-weighted doubly robust estimation, with the interpolation chosen to minimize variance.
The paper says standard doubly robust off-policy evaluation can stay unbiased but suffer from high-variance action-level importance weights. OffCEM lowers that variance with cluster-level weights, but depends on local reward-model correctness. VOCEM chooses a closed-form optimal interpolation coefficient and is shown to have variance no larger than either endpoint under the stated assumptions. In experiments on synthetic settings and two large-action benchmarks, it beats both OffCEM and DR across all 23 evaluated conditions. ArXiv · AI/CL/LG's note
The paper says standard doubly robust off-policy evaluation can stay unbiased but suffer from high-variance action-level importance weights. OffCEM lowers that variance with cluster-level weights, but depends on local reward-model correctness. VOCEM chooses a closed-form optimal interpolation coefficient and is shown to have variance no larger than either endpoint under the stated assumptions. In experiments on synthetic settings and two large-action benchmarks, it beats both OffCEM and DR across all 23 evaluated conditions. ArXiv · AI/CL/LG's note
score 4