Qwen3.8 Flash Next now powers the standard mode, which Alibaba says is twice as fast and uses 75% fewer tokens per task than the previous default.
Alibaba's Tongyi Office AI assistant collapsed its model lineup into two tiers on 2026-08-26 (Asia): a default "standard mode" powered by the new Qwen3.8-Flash-Next model, and a paid "advanced mode" for harder work.
Standard mode is positioned to handle ~95% of daily office tasks; the advanced tier covers the remaining 5% of complex work. The Qwen LLM and Tongyi Office teams co-tuned the standard build for multi-step planning, tool selection, and context compression, with a custom inference harness for throughput.
Alibaba's own office-scenario numbers: about 100% faster single-task generation and about 75% lower average token consumption on standard mode vs. the previous default. The model is a 125-billion-parameter multimodal mixture-of-experts with 6 billion active, per a MarkTechPost walkthrough and Alibaba's tech report, which frames the architecture as a preview of Qwen4.
Alibaba also claims Qwen3.8-Flash-Next outperforms Anthropic's Claude Opus 4.6 — a vendor assertion with no independent reproduction cited. The 100% and 75% numbers are the same category: vendor-tested, not independently measured.
What is not yet public: third-party benchmark scores, sustained-load latency, and advanced-tier pricing.