Pre-Compiled Pipeline Shards for Distributed LLM Inference on Intel AI PC Fleets
A few Intel AI PCs can jointly run an LLM too large for any one machine by serving pre-compiled layer shards in sequence.
The paper splits models into OpenVINO pipeline stages, with each PC running one shard and passing activations to the next. It reports that adding a `beam_idx` Gather restores a missed OpenVINO GPU optimization, bringing shard performance back to monolithic parity. With speculative decoding and interleaved multi-user requests, a two-node Llama 3.1 8B INT4 setup reaches 1.79x single-user throughput. The authors also describe a four-node Lunar Lake deployment serving a 70B model at interactive speed. ArXiv · AI/CL/LG's note
The paper splits models into OpenVINO pipeline stages, with each PC running one shard and passing activations to the next. It reports that adding a `beam_idx` Gather restores a missed OpenVINO GPU optimization, bringing shard performance back to monolithic parity. With speculative decoding and interleaved multi-user requests, a two-node Llama 3.1 8B INT4 setup reaches 1.79x single-user throughput. The authors also describe a four-node Lunar Lake deployment serving a 70B model at interactive speed. ArXiv · AI/CL/LG's note
score 4