WebLLM: high-performance in-browser LLM inference engine
WebLLM runs open-source LLMs locally in the browser with WebGPU acceleration and an OpenAI-style API.
The project says inference needs no server support, with models loaded and cached in the browser. It supports streaming, JSON mode, logit controls, seeding, and preliminary function calling through an OpenAI-compatible interface. Built-in model families include Llama, Phi, Gemma, Mistral, and Qwen, with custom MLC-format models supported. It also includes worker, service worker, Chrome extension, cache-backend, and optional integrity-verification paths for browser deployments. HN · LLMs' note
The project says inference needs no server support, with models loaded and cached in the browser. It supports streaming, JSON mode, logit controls, seeding, and preliminary function calling through an OpenAI-compatible interface. Built-in model families include Llama, Phi, Gemma, Mistral, and Qwen, with custom MLC-format models supported. It also includes worker, service worker, Chrome extension, cache-backend, and optional integrity-verification paths for browser deployments. HN · LLMs' note
score 6