Qwen3.8-Flash-Next Previews Qwen4 Architecture With 6B Active Parameters

Covered by 5 sources
Read full post
Alibaba's Qwen team unveiled Qwen3.8-Flash-Next, an experimental 125B-parameter model activating only 6B per token, previewing the Qwen4 architecture focused on cost-efficient inference with hybrid attention and long context support.

Covered by 5 sources

More on this story


More in LLM & Text Generation

DeepSeek V4.1 Flash now available on AI Gateway

Covered by 2 sources

Cohere Debuts Open-Weight 218B Mixture-of-Experts Machine Translation Model

Unite.AI