ARC: Fair Relative Advantage Comparison in Open-Ended Real-World Interaction
ARC targets a fairness failure in RL training when agents with different valid interaction styles are compared as if they were equivalent.
The paper argues that open-ended agent tasks can reward style preferences instead of context-appropriate behavior when rollouts are grouped too broadly. ARC conditions rollout grouping by strategy, then combines that with hybrid rewards and entropy regularization. The authors test it inside their INTER interaction setup and report stronger tool-use benchmark results, plus a drop in time-to-first-token from 4.91s to 1.27s versus a think-style baseline. Implementation and the INTER-86K training data are slated for release. HF Daily Papers' note
The paper argues that open-ended agent tasks can reward style preferences instead of context-appropriate behavior when rollouts are grouped too broadly. ARC conditions rollout grouping by strategy, then combines that with hybrid rewards and entropy regularization. The authors test it inside their INTER interaction setup and report stronger tool-use benchmark results, plus a drop in time-to-first-token from 4.91s to 1.27s versus a think-style baseline. Implementation and the INTER-86K training data are slated for release. HF Daily Papers' note
score 5