Megadose AI progress, ranked and analyzed.

A report-grounded vision-language foundation model for colonoscopy from 280000 routine reports

· ArXiv · AI/CL/LG ·
EndoCLIP was trained by recovering lesion-level image-text pairs from routine colonoscopy reports.

The authors built the model from 125,756 lesion-level pairs drawn from 280,476 records. It beat general-purpose and biomedical vision-language encoders on retrieval, report generation, and six multi-centre classification tasks. In benign-versus-malignant testing, a linear probe came close to expert readers in a blinded study with 12 endoscopists. The paper argues that routine reports can become scalable supervision when findings are linked back to frames. ArXiv · AI/CL/LG's note

score 5

Categories: Research