Jev thinks "I don't know'', but doesn't say it: Introducing Sys1Cal-v1 Dataset for Probability Calibration
Jev’s binary probabilities appear to hide a missing “unknown” mass.
The paper introduces Sys1Cal-v1, a benchmark where True/False questions have exact probabilities known by construction. It tests Jev’s Noul, Choice and Score primitives against those ground-truth distributions using total variation distance. The author reports that Jev’s Choice outputs force \(P(A)\) and \(P(\neg A)\) to sum to 1, while an inferred third term \(P(U)\) is absent. Recovering that missing term raises median soft accuracy for Choice answers from 0.771 to 0.978 in the paper’s evaluation. Source: HF Daily Papers' note
The paper introduces Sys1Cal-v1, a benchmark where True/False questions have exact probabilities known by construction. It tests Jev’s Noul, Choice and Score primitives against those ground-truth distributions using total variation distance. The author reports that Jev’s Choice outputs force \(P(A)\) and \(P(\neg A)\) to sum to 1, while an inferred third term \(P(U)\) is absent. Recovering that missing term raises median soft accuracy for Choice answers from 0.771 to 0.978 in the paper’s evaluation. Source: HF Daily Papers' note
score 4