Qwen3.8-Flash-Next
Qwen released an open-weights multimodal MoE model previewing Qwen4 architecture with 125B total parameters and 6B active parameters.
Excerpt
Qwen3.8-Flash-Next
Another open weights model from Qwen. This one is "a multimodal MoE model that also serves as an early preview of the architecture used in Qwen4".
It's pretty big: 125B tokens, but only 6B active which means it gets a pretty big performance boost.
I've been trying it out on a DGX Spark using these Unsloth quantized models. I'm still exploring the model - so far I've tried the 72.5GB UD-IQ1_S one (producing these pelicans) and the 78.9GB UD-Q2_K_XL (producing these).
My favorite so far was this xhigh reasoning effort one from UD-Q2_K_XL:
Via Hacker News
Tags: ai, generative-ai, llms, qwen, pelican-riding-a-bicycle, ai-in-china, nvidia-spark</a
Read at source: https://simonwillison.net/2026/Aug/26/qwen38-flash-next/