IQ-JEPA: A Joint-Embedding Predictive Architecture with a Hermitian Vision Transformer for Sound Speed and Attenuation Estimation from Ultrasound IQ Data
IQ-JEPA uses unlabeled ultrasound IQ data to cut the number of simulated labels needed for sound-speed estimation.
The paper pretrains an encoder to predict masked complex IQ regions from surrounding context, then fine-tunes it on simulated sound-speed maps. Its Hermitian vision transformer is built to operate directly on complex ultrasound signals and respect phase properties the authors say matter for the task. On 79,293 Fullwave simulations, pretraining reached 15.60 m/s error with 10,000 labels, about a threefold label-efficiency gain over supervised training. The authors also report transferable frozen features for sound speed and attenuation, framing the result as an early step toward a quantitative-ultrasound foundation model. ArXiv · AI/CL/LG's note
The paper pretrains an encoder to predict masked complex IQ regions from surrounding context, then fine-tunes it on simulated sound-speed maps. Its Hermitian vision transformer is built to operate directly on complex ultrasound signals and respect phase properties the authors say matter for the task. On 79,293 Fullwave simulations, pretraining reached 15.60 m/s error with 10,000 labels, about a threefold label-efficiency gain over supervised training. The authors also report transferable frozen features for sound speed and attenuation, framing the result as an early step toward a quantitative-ultrasound foundation model. ArXiv · AI/CL/LG's note
score 4