Megadose AI progress, ranked and analyzed.

Wasserstein Filtering: A Sample Selection Method for Robust Distribution Learning

· ArXiv · AI/CL/LG ·
The paper proposes filtering contaminated datasets by removing samples that most distort the Wasserstein geometry of the empirical distribution.

Wasserstein Filtering selects a retained subset whose empirical distribution is farthest from the fully contaminated empirical measure, aiming to isolate influential outliers. The authors give three tractable variants: marginal screening, SinkWF, and SlicedWF. They also define the FELP contamination model and prove minimax optimality for bounded-covariance distribution families under that model. Experiments cover synthetic data, anomaly detection benchmarks, and diffusion-model preprocessing under heavy contamination. ArXiv · AI/CL/LG's note

score 5

Categories: Research