GS-Codec: A Gaussian-Splatting Bottleneck for Neural Audio Coding
The paper replaces codec codebooks with a Gaussian-splatting-style decomposition of audio latents.
GS-Codec fits encoder segments as weighted sums of 1D Gaussian primitives, then reconstructs audio from the rendered sum. A lightweight predictor network is trained to estimate those primitive parameters in one forward pass, avoiding the iterative fitting loop at inference. The authors say one trained checkpoint can trade bitrate against quality after training by changing primitive count and parameter bit depth. They report results matching or beating EnCodec and DAC on SIM, STOI, and UTMOS at comparable bitrates, with comparable WER. ArXiv · AI/CL/LG's note
GS-Codec fits encoder segments as weighted sums of 1D Gaussian primitives, then reconstructs audio from the rendered sum. A lightweight predictor network is trained to estimate those primitive parameters in one forward pass, avoiding the iterative fitting loop at inference. The authors say one trained checkpoint can trade bitrate against quality after training by changing primitive count and parameter bit depth. They report results matching or beating EnCodec and DAC on SIM, STOI, and UTMOS at comparable bitrates, with comparable WER. ArXiv · AI/CL/LG's note
score 5