Wasserstein Filtering: A Sample Selection Method for Robust Distribution Learning
The paper proposes filtering contaminated datasets by removing samples that most distort the Wasserstein geometry of the empirical distribution.
Wasserstein Filtering selects a retained subset whose empirical distribution is farthest from the fully contaminated empirical measure, aiming to isolate influential outliers. The authors give three tractable variants: marginal screening, SinkWF, and SlicedWF. They also define the FELP contamination model and prove minimax optimality for bounded-covariance distribution families under that model. Experiments cover synthetic data, anomaly detection benchmarks, and diffusion-model preprocessing under heavy contamination. ArXiv · AI/CL/LG's note
Wasserstein Filtering selects a retained subset whose empirical distribution is farthest from the fully contaminated empirical measure, aiming to isolate influential outliers. The authors give three tractable variants: marginal screening, SinkWF, and SlicedWF. They also define the FELP contamination model and prove minimax optimality for bounded-covariance distribution families under that model. Experiments cover synthetic data, anomaly detection benchmarks, and diffusion-model preprocessing under heavy contamination. ArXiv · AI/CL/LG's note
score 5