A Missing Piece for Trustworthy AI Reviewers: From Benchmarking Rhetorical Robustness to SciCore Review
The paper argues AI reviewers need to be tested for whether wording changes sway their judgments.
The authors define “rhetorical robustness” as staying stable across meaning-preserving rewrites while still distinguishing between different papers. They introduce RobustReview, a 1,260-version full-manuscript benchmark, and report that 30 reviewer setups exposed cases where apparent stability came from score collapse. Their SciCore reviewer combines a full-manuscript judgment with one based on an extracted structured science core, aiming to reduce sensitivity to presentation. In a GPT-5.5 comparison, SciCore led on the paper’s joint stability-discrimination measure while remaining competitive with human alignment. Source: HF Daily Papers' note.
The authors define “rhetorical robustness” as staying stable across meaning-preserving rewrites while still distinguishing between different papers. They introduce RobustReview, a 1,260-version full-manuscript benchmark, and report that 30 reviewer setups exposed cases where apparent stability came from score collapse. Their SciCore reviewer combines a full-manuscript judgment with one based on an extracted structured science core, aiming to reduce sensitivity to presentation. In a GPT-5.5 comparison, SciCore led on the paper’s joint stability-discrimination measure while remaining competitive with human alignment. Source: HF Daily Papers' note.
score 4