Prototype-Rectified Iterative Self-supervised Manifold Denoising under Severe Acoustic Shift
PRISM claims a training-free way to undo severe noise shifts in audio-text model embeddings with a single projection at inference.
The paper says heavy acoustic noise pushes multimodal representations through a low-rank affine shift, with most distortion concentrated in leading principal components. PRISM estimates that shift from an unlabeled target batch, using frozen text prototypes as anchors, then applies closed-form corrections as one static matrix operation. On UrbanSound8K, it reports a 12.94-point gain over zero-shot and a 9.41-point gain over an oracle-assisted TTA baseline. The authors also describe a “Polyphonic Trap” failure case for broadband classes and say Confidence-Aware Regression recovers up to 8.16 points for the worst-hit class. HF Daily Papers' note
The paper says heavy acoustic noise pushes multimodal representations through a low-rank affine shift, with most distortion concentrated in leading principal components. PRISM estimates that shift from an unlabeled target batch, using frozen text prototypes as anchors, then applies closed-form corrections as one static matrix operation. On UrbanSound8K, it reports a 12.94-point gain over zero-shot and a 9.41-point gain over an oracle-assisted TTA baseline. The authors also describe a “Polyphonic Trap” failure case for broadband classes and say Confidence-Aware Regression recovers up to 8.16 points for the worst-hit class. HF Daily Papers' note
score 4