Megadose AI progress, ranked daily.

Alibaba releases Qwen3.8-Flash, an open-weight, 125B-parameter model built on its next-gen Qwen 4 architecture, saying it rivals Opus 4.6 and V4-Flash (Luz Ding/Bloomberg)

Techmeme ·
Alibaba is pitching Qwen3.8-Flash as a cheaper open-weight preview of its next Qwen architecture.

The model is a multimodal MoE with 125B parameters and 6B active, with native 256K context extendable to 1M. Alibaba says its new GDN and Qwen Sparse Attention setup cuts long-context costs, including faster prefill and decode at 1M tokens. The company says a production version will come to QwenCloud at $0.16 per 1M input tokens and $0.47 per 1M output tokens. Techmeme's note

score 8

Categories: Model Releases, OSS & Tools