Megadose AI progress, ranked and analyzed.

Bridging the Confidence Gap: Temperature Scaling for Calibrating Test-Time Prompt Tuning

· ArXiv · AI/CL/LG ·
CoTS uses temperature scaling to pull adapted prompt-tuned predictions back toward zero-shot confidence without giving up accuracy.

The paper says test-time prompt tuning can improve accuracy while making predictions less calibrated. Its proposed post-hoc method minimizes the confidence gap between adapted and zero-shot outputs, then extends that to a weak-strong augmentation ensemble called E-CoTS. On ImageNet variants, E-CoTS cuts average expected calibration error from 11.90% to 5.38% and raises accuracy from 60.74% to 62.95%. ArXiv · AI/CL/LG's note

score 4

Categories: Research