Megadose AI progress, ranked and analyzed.

Qwen 3.8 27B is excellent, but it defaults to wildly overthinking things

Simon Willison ·
Qwen 3.8 27B can run serious local LLM work from a 17GB file, but its default `xhigh` reasoning setting makes even simple prompts sprawl.

Willison says the Apache 2 27B vision model produced strong SVGs, accurate pelican bounding boxes, useful code, and a working coding-agent session on his own machines. The problem is the default reasoning effort: it burned 22,276 reasoning tokens and 21 minutes on a pelican SVG, then overbuilt a simple circle prompt. Turning reasoning down is his practical recommendation, though he notes reasoning helped with a one-shot bounding-box labeling tool. Speed remains the main blocker: LM Studio gave him about 15-30 tokens per second, while llama.cpp’s MTP setup improved his Spark benchmark by about 72%. Simon Willison's note

score 8

Categories: Model Releases, OSS & Tools