Launch HN: Magnitude (YC S25) – Self-optimizing inference engine for agents
Magnitude says it speeds up local open-model inference by compiling and tuning kernels on the user’s own hardware.
The project is an open-source desktop inference engine for agents, with macOS, Windows, and Linux builds. It claims up to 2x faster performance than llama.cpp, citing 92% faster decode on Metal and 19% on CUDA. Magnitude supports Apple Silicon, NVIDIA, AMD, and CPU-only setups, and connects to agents including Pi, OpenCode, Hermes, Codex, Claude Code, and Cline. It says prompts, files, and models stay on-device after models are downloaded. HN · Inference's note
The project is an open-source desktop inference engine for agents, with macOS, Windows, and Linux builds. It claims up to 2x faster performance than llama.cpp, citing 92% faster decode on Metal and 19% on CUDA. Magnitude supports Apple Silicon, NVIDIA, AMD, and CPU-only setups, and connects to agents including Pi, OpenCode, Hermes, Codex, Claude Code, and Cline. It says prompts, files, and models stay on-device after models are downloaded. HN · Inference's note
score 5