GUT: Quantifying and Optimizing the Reasoning Uncertainty of LLMs via Graph Complexity
GUT models an LLM’s possible reasoning paths as a directed acyclic graph, then uses graph complexity as a proxy for uncertainty.
The paper splits the method into GUT-Q, which estimates reasoning uncertainty, and GUT-O, which tries to reduce it. The optimization module uses reinforcement learning with negative uncertainty as the reward. The authors say tests on four LLMs and five datasets support the method’s effectiveness. ArXiv · AI/CL/LG's note
The paper splits the method into GUT-Q, which estimates reasoning uncertainty, and GUT-O, which tries to reduce it. The optimization module uses reinforcement learning with negative uncertainty as the reward. The authors say tests on four LLMs and five datasets support the method’s effectiveness. ArXiv · AI/CL/LG's note
score 4