FreeToken: Efficient Edge-Native MoE Serving with Bandwidth-Adaptive Execution
FreeToken is pitched as a way to serve frontier-scale MoE models on ordinary local hardware by adapting execution to available bandwidth and memory.
The system co-designs model layout, loading, expert residency, CPU-GPU execution, agent state reuse, and runtime memory management for edge machines. It continuously remaps computation and model state instead of using a fixed offloading plan. The authors say it supports more than 20 MoE models and agent workloads, scaling from an 8GB laptop GPU to a single workstation GPU. They claim this enables serving a 35B model on a laptop, a 284B model on a gaming desktop, and GLM-5.2 753B on one workstation GPU. HF Daily Papers' note
The system co-designs model layout, loading, expert residency, CPU-GPU execution, agent state reuse, and runtime memory management for edge machines. It continuously remaps computation and model state instead of using a fixed offloading plan. The authors say it supports more than 20 MoE models and agent workloads, scaling from an 8GB laptop GPU to a single workstation GPU. They claim this enables serving a 35B model on a laptop, a 284B model on a gaming desktop, and GLM-5.2 753B on one workstation GPU. HF Daily Papers' note
score 5