EAServe: Encode-Aware Disaggregated Serving for Multimodal Large Language Models
EAServe treats multimodal encoding as the control point for the whole serving pipeline.
The paper argues that image, video, and audio inputs break the usual prefill-decode split used for text-only LLM serving. Its system coordinates encode batching, prefill placement, and GPU sharing so downstream workers are not starved while encode GPUs sit underused. In tests across image, video, and audio MLLMs, EAServe reports up to 4.3x higher goodput than NVIDIA Dynamo and 1.7x higher than vLLM under the same SLO constraints. ArXiv · AI/CL/LG's note
The paper argues that image, video, and audio inputs break the usual prefill-decode split used for text-only LLM serving. Its system coordinates encode batching, prefill placement, and GPU sharing so downstream workers are not starved while encode GPUs sit underused. In tests across image, video, and audio MLLMs, EAServe reports up to 4.3x higher goodput than NVIDIA Dynamo and 1.7x higher than vLLM under the same SLO constraints. ArXiv · AI/CL/LG's note
score 5