ActReview: Rebuttal-Guided Training Data and Rubric Rewards for Actionable Peer Review Generation
ActReview trains review models to turn paper-specific weaknesses into grounded revision plans using author rebuttals as supervision.
The paper frames actionable peer review as two linked tasks: diagnosing weaknesses and suggesting concrete revisions. Its ActReview-40K dataset aligns OpenReview reviewer concerns with author responses, then grounds the feedback in localized evidence from the paper. The authors post-train Qwen3-8B-Base with supervised fine-tuning and GRPO rewards tied to weakness-specific rubrics. Results show gains in actionability and grounding over prior review-generation models, though human evaluation still finds a gap in technical accuracy. HF Daily Papers' note
The paper frames actionable peer review as two linked tasks: diagnosing weaknesses and suggesting concrete revisions. Its ActReview-40K dataset aligns OpenReview reviewer concerns with author responses, then grounds the feedback in localized evidence from the paper. The authors post-train Qwen3-8B-Base with supervised fine-tuning and GRPO rewards tied to weakness-specific rubrics. Results show gains in actionability and grounding over prior review-generation models, though human evaluation still finds a gap in technical accuracy. HF Daily Papers' note
score 4