ActReview: Rebuttal-Guided Training Data and Rubric Rewards for Actionable Peer Review Generation
ActReview uses author rebuttals as supervision for review feedback that points to concrete paper revisions.
The paper frames the task as generating both diagnostic claims and revision suggestions. Its ActReview-40K dataset aligns reviewer weaknesses with author responses from OpenReview threads, grounded in localized evidence from the paper. The authors post-train Qwen3-8B-Base with supervised fine-tuning and GRPO using weakness-specific rubric rewards. They report stronger actionability and grounding than prior specialized review-generation models, while human evaluation still finds a gap in technical accuracy. ArXiv · AI/CL/LG's note
The paper frames the task as generating both diagnostic claims and revision suggestions. Its ActReview-40K dataset aligns reviewer weaknesses with author responses from OpenReview threads, grounded in localized evidence from the paper. The authors post-train Qwen3-8B-Base with supervised fine-tuning and GRPO using weakness-specific rubric rewards. They report stronger actionability and grounding than prior specialized review-generation models, while human evaluation still finds a gap in technical accuracy. ArXiv · AI/CL/LG's note
score 4