PACE: Perceived-Latency-Aware Cascading Service Routing and Filler Control for QoE-Efficient Retrieval-Augmented Dialogue Serving
PACE targets the delay users feel before a dialogue system starts answering, not just backend response time.
The paper defines Perceived Time-to-First-Response as the serving objective for retrieval-augmented dialogue. Its system jointly chooses the answer path and the filler shown during the wait, using a load-adaptive router, path-filler controller, and volatility-aware cache admission. In a humanoid-robot sales deployment on 75,000 CarQA requests, the cascade cut pure-LLM P95 PTFR from 0.53s to 0.29s at c16. The authors also report fewer filler calls, no filler-answer conflicts, and stale answers reduced to zero under their cache rule. ArXiv · AI/CL/LG's note
The paper defines Perceived Time-to-First-Response as the serving objective for retrieval-augmented dialogue. Its system jointly chooses the answer path and the filler shown during the wait, using a load-adaptive router, path-filler controller, and volatility-aware cache admission. In a humanoid-robot sales deployment on 75,000 CarQA requests, the cascade cut pure-LLM P95 PTFR from 0.53s to 0.29s at c16. The authors also report fewer filler calls, no filler-answer conflicts, and stale answers reduced to zero under their cache rule. ArXiv · AI/CL/LG's note
score 5