Alibaba presented Qwen3.8-Flash-Next as a preliminary showcase of the Qwen4 architecture. According to The Decoder, it is a mixture-of-experts model that activates 6 out of 125illion parameters for each token.

The source description states that training the model cost nine times less than a comparable baseline, and it outperformed larger models like DeepSeek-V4-Flash and Claude Opus 4.6 on programming and office task benchmarks.

If these results are confirmed, model developers will face increased pressure regarding the cost of their offerings. However, the source consists only of metadata and a synopsis, so independent verification of the results and comparison details is currently unavailable.