Rubric-to-Code Credit Assignment for Reinforcement Learning
RCCA trains app-generation models by tying rubric failures back to the code spans most responsible.
The paper targets HTML/CSS/JavaScript app generation, where a single overall reward can blur which part of the output caused a functional failure. RCCA uses explicit functional rubrics, hierarchical failure signals, and evaluator attributions aligned to generated tokens. Its Ling-RCCA-Flash model reports 41.25 on MiniAppBench, 32.20 points above Ling-3.0-Flash and slightly ahead of Claude Opus 4.5. On ArtifactsBench, it reports 76.19, above the listed GPT-5 score under the official leaderboard setting. Source: HF Daily Papers' note.
The paper targets HTML/CSS/JavaScript app generation, where a single overall reward can blur which part of the output caused a functional failure. RCCA uses explicit functional rubrics, hierarchical failure signals, and evaluator attributions aligned to generated tokens. Its Ling-RCCA-Flash model reports 41.25 on MiniAppBench, 32.20 points above Ling-3.0-Flash and slightly ahead of Claude Opus 4.5. On ArtifactsBench, it reports 76.19, above the listed GPT-5 score under the official leaderboard setting. Source: HF Daily Papers' note.
score 5