Should We Type or Talk to LLM Agents? A Comprehensive Study of Voice and Keyboard Input Perturbations
Voice input hurt tested LLM agents more than keyboard noise, especially when answers had to be reasoned out.
The paper introduces HIVE, a perturbation suite for voice transcription and QWERTY keyboard errors. Voice transcription lowered accuracy across every instruction-tuned model tested, with the structural changes in transcripts doing more damage than filler words. Keyboard perturbations were less costly, and models tolerated many before accuracy dropped. The authors tie both effects to token survival: losing original question tokens hurts, while added tokens matter less. ArXiv · AI/CL/LG's note
The paper introduces HIVE, a perturbation suite for voice transcription and QWERTY keyboard errors. Voice transcription lowered accuracy across every instruction-tuned model tested, with the structural changes in transcripts doing more damage than filler words. Keyboard perturbations were less costly, and models tolerated many before accuracy dropped. The authors tie both effects to token survival: losing original question tokens hurts, while added tokens matter less. ArXiv · AI/CL/LG's note
score 4