Hierarchical Group-Conditional Conformal Risk Control for Selective Prediction in Language Models
The paper’s claim is that group-level risk control can stop selective LMs from hiding errors in subgroups.
HG-CRC calibrates abstention thresholds across a user-defined hierarchy of domains, difficulty, prompts, and related groupings, without retraining the model. In the authors’ tests, standard CRC met the overall budget while failing under group-composition shift in up to 47% of trials. HG-CRC reached empirical zero violations on ARC Challenge for Qwen3-4B and Llama-3.1-8B, though the paper notes those zeros are bootstrap bounds, not certified proof. The cost was steep: participation fell 22 to 37 points versus global CRC, and MMLU-Pro results were more limited. ArXiv · AI/CL/LG's note
HG-CRC calibrates abstention thresholds across a user-defined hierarchy of domains, difficulty, prompts, and related groupings, without retraining the model. In the authors’ tests, standard CRC met the overall budget while failing under group-composition shift in up to 47% of trials. HG-CRC reached empirical zero violations on ARC Challenge for Qwen3-4B and Llama-3.1-8B, though the paper notes those zeros are bootstrap bounds, not certified proof. The cost was steep: participation fell 22 to 37 points versus global CRC, and MMLU-Pro results were more limited. ArXiv · AI/CL/LG's note
score 4