Efficient Test-Time Adaptation through Human-AI Interaction
The paper says agents can adapt to an individual user’s standards from interaction history within only tens of tasks.
The authors propose TAHI, a test-time adaptation setup that folds human-agent interaction signals into agent context, weights, and an evolving rubric. They test it with 30 people across writing and visual creation, covering 600 tasks. Reported solo task success improves by 4.5% to 20.9%. The rubric module also finds 16.0% to 22.3% more failures than LM-only or human-only rubrics. ArXiv · AI/CL/LG's note
The authors propose TAHI, a test-time adaptation setup that folds human-agent interaction signals into agent context, weights, and an evolving rubric. They test it with 30 people across writing and visual creation, covering 600 tasks. Reported solo task success improves by 4.5% to 20.9%. The rubric module also finds 16.0% to 22.3% more failures than LM-only or human-only rubrics. ArXiv · AI/CL/LG's note
score 5