Megadose AI progress, ranked and analyzed.

WebLLM: high-performance in-browser LLM inference engine

· HN · LLMs ·
WebLLM runs open-source LLMs locally in the browser with WebGPU acceleration and an OpenAI-style API.

The project says inference needs no server support, with models loaded and cached in the browser. It supports streaming, JSON mode, logit controls, seeding, and preliminary function calling through an OpenAI-compatible interface. Built-in model families include Llama, Phi, Gemma, Mistral, and Qwen, with custom MLC-format models supported. It also includes worker, service worker, Chrome extension, cache-backend, and optional integrity-verification paths for browser deployments. HN · LLMs' note

score 6

Categories: OSS & Tools

Discussions

  • hn · 102 points · 17 comments
  • hn · 102 points · 17 comments
  • hn · 104 points · 17 comments
  • hn · 104 points · 17 comments