Alibaba's newest downloadable AI model, Qwen3.8 Max, is the latest in its Qwen large language model family. It says the model ran a human free coding loop for ten days.
Alibaba's Qwen lab says its new flagship AI model, Qwen3.8-Max, spent more than ten days writing and rewriting its own coding harness without a human in the loop, and that a separate 125-hour research run beat a published paper's benchmark by 2.71 points. Both are demos, not audits, and one detail from the same source that surfaced the release tempers the celebration: this is not the largest publicly downloadable model in the world. Moonshot's Kimi K3 still holds that title.
Qwen3.8-Max is a 2.4-trillion-parameter mixture-of-experts model, meaning it has 2.4 trillion stored weights but only activates about 95 billion of them, roughly 4%, for any single token. That keeps per-query compute closer to a much smaller model while letting the lab draw on a much larger knowledge base. A 27-billion-parameter sibling, Qwen3.8-27B, ships alongside it for cheaper coding and cowork workloads. Both were promised as open-weights downloads, meaning the trained model files can be downloaded and run by anyone, not closed APIs, in a release window of July 25–27, 2026. The closed API line carries its own economics: Alibaba lists $2 per million input tokens, $6 per million output, with OpenAI- and Anthropic-compatible protocols and a 1-million-token context window.
The release's most novel claims are the long-horizon ones. In one demo, the lab says Qwen3.8-Max ran a coding agent unattended for more than 240 hours, building a self-evolving harness from scratch. In another, it ran a 125-hour autonomous data-selection loop on a published paper and improved on that paper's reported benchmark by 2.71 points. A third demo put it in the top 13% of 526 human teams in the WWW2025 Multimodal Dialogue Intent Recognition Challenge within 24 hours, ahead of 87% of the field.
Those numbers, plus a reported 93.0 on PaperBench, 74.8 on CoWorkBench, and 81.9 on WideSearch, are the lab's own. None have been independently reproduced in third-party leaderboards that this story could verify, and the same ZhihuFrontier summary that lists them is the only cited source for the figures.
The chip-design result is the most concrete. Alibaba says Qwen3.8-Max produced a complete digital logic flow for a cryptographic accelerator, a hardware block that performs the modular arithmetic used in public-key encryption, and reduced the gate count, the number of basic logic transistors in the design, from 8,298 to 678. That is a 91% reduction. The flow is reported to have closed physical timing at 500 MHz, meaning the resulting layout can run at that clock speed without violating timing constraints. An 81% die-area reduction, the share of silicon the block occupies on a chip, came with it.
None of the demos answer the question the release actually opens: whether open-weights Chinese labs can match closed US labs on tasks that run for days, where the public evidence is still mostly the lab's own. The E-Commerce Bench result, a 4.16x return on a 365-day simulated store ending with a balance of ¥416,252 (roughly $58,000 at typical mid-2026 exchange rates), is a separate long-horizon claim, again from the lab.
Deployment is already in motion. Baseten, a US-based inference platform, confirmed it would support both Qwen3.8-Max and the 27B sibling on release day. The Max also landed on Venice and Command Code, two consumer-facing chat surfaces, in the same window.
The watch item is concrete: the open-weights downloads Alibaba promised for "within a week" of the late-July announcement were due around the first week of August 2026. Whether they land on schedule, and whether the long-horizon demos survive contact with independent reproduction, are the two facts that will determine whether Qwen3.8-Max is the start of a reliability story or just another parameter-count race.