Megadose Built for builders and researchers.

MARGIN: Runtime Confidence Calibration for Multi-Agent Foundation Model Coordination

· HF Daily Papers ·
MARGIN treats model confidence as something to recalibrate at runtime, not take at face value.

The method learns per-model confidence corrections from observed outcomes, without retraining or a held-out calibration set. It tracks recent accuracy against stated confidence in bands, then uses corrected scores to weight answers in multi-model coordination. In the reported tests, raw confidence could be actively misleading: on BigCodeBench, higher average confidence was negatively related to accuracy. Calibration improved correct-answer ranking and lifted answer-selection accuracy on two of three code-generation benchmarks. HF Daily Papers' note

score 4

Categories: Research