Megadose AI progress, ranked and analyzed.

onPanda: Efficient Annotation of On-Policy Alignment Data for LLMs and Agents via Token-Level Correction

· ArXiv · AI/CL/LG ·
onPanda turns alignment annotation into a token-level correction loop, then uses the model to regenerate from each fix.

Annotators mark the first bad token, replace it with a candidate token or free-form text, and let the system continue generation from the corrected prefix. A small controlled study in the paper reports a 52% reduction in median annotation time versus manual post-editing. Because most final tokens still come from the model, the authors argue the data remains close to the model’s own sampling distribution for on-policy SFT and preference data. The work also releases Panda-CVL and a token-level correction benchmark. ArXiv · AI/CL/LG's note

score 4

Categories: Research