Alibaba just released Qwen3.8-Flash: “An early preview of the architecture in Qwen4”

The New Stack (AI)
Read full post
Alibaba has released Qwen3.8-Flash, a 125-billion-parameter multimodal Mixture-of-Experts AI model serving as an early preview of the architecture planned for Qwen4. This model introduces architectural innovations like hybrid Gated DeltaNet and Gated Attention to improve efficiency and capability, and is positioned as a cost-effective solution for coding and office tasks.

More on this story


More in Machine Learning

Machine Learning3 min read

Nvidia and Palantir fine-tune a 30B Nemotron model for Nvidia’s supply chain. It beats a model 18 times its size.

Covered by 3 sources
Machine Learning6 min read

CoreWeave Puts Field Engineers Inside Customer Teams for Physical AI

Covered by 2 sources
Machine Learning2 min read

Weatherwatch: AI model beats standard methods at predicting cyclones

The Guardian