The Handoff Tax: Continuing Non-Native Trajectories in LLM Agents
Switching a coding agent to a stronger model mid-run buys back less than half the lost quality, while still adding meaningful cost.
The paper calls that penalty the “handoff tax”: the receiving model has to continue work shaped by another model’s prior calls, tool use, and edits. The authors test handoffs between cheaper lower-capability and costlier higher-capability models from Claude and GPT families. Escalation performs worse when the stronger model inherits the full weaker-model trajectory; giving it less of that trajectory improves results. Downshifting looks more favorable, and in that direction removing the stronger model’s trajectory hurts quality. HF Daily Papers' note
The paper calls that penalty the “handoff tax”: the receiving model has to continue work shaped by another model’s prior calls, tool use, and edits. The authors test handoffs between cheaper lower-capability and costlier higher-capability models from Claude and GPT families. Escalation performs worse when the stronger model inherits the full weaker-model trajectory; giving it less of that trajectory improves results. Downshifting looks more favorable, and in that direction removing the stronger model’s trajectory hurts quality. HF Daily Papers' note
score 5