Megadose AI progress, ranked and analyzed.

When Many Answers Are Valid, Voting Fails: Symbolic Verification for Best-of-K Causal Reasoning in LLMs

· HF Daily Papers ·
CALVER picks answers by checking causal validity, not by counting which sampled answer appears most.

The paper argues that plurality voting can lose when several different answers are valid and one repeated confounding error dominates the pool. CALVER scores structured traces against Pearl-style criteria such as d-separation, backdoor adjustment, and intervention, without using a reference answer. On CLEAR find-one-valid tasks, it reports 42.1% accuracy while plurality, reward models, LLM judges, and confidence-based selection stay near 30% on the same frozen samples. The authors say the gain holds across published Bayesian networks, another model family, text-built graphs, and related ATE and logic checks. HF Daily Papers' note

score 5

Categories: Research