Nehanth/pooled
Pooled splits one open LLM across browser tabs so multiple devices can run a model none could fit alone.
The repo says each device holds some layers, runs them with WebGPU/WGSL, and passes small hidden states over direct WebRTC links. Its demo shows Qwen 3.6 35B MoE running across a Mac and an iPhone, then using Code mode to build and revise a small Tetris app in the tab. There is no install or account, and the project says no server performs inference. Benchmarks emphasize that speculative decoding is what keeps multi-device rooms usable once network latency is added. GitHub · LLM repos' note
The repo says each device holds some layers, runs them with WebGPU/WGSL, and passes small hidden states over direct WebRTC links. Its demo shows Qwen 3.6 35B MoE running across a Mac and an iPhone, then using Code mode to build and revise a small Tetris app in the tab. There is no install or account, and the project says no server performs inference. Benchmarks emphasize that speculative decoding is what keeps multi-device rooms usable once network latency is added. GitHub · LLM repos' note
score 5