Four of the five labs in this series are American. The fifth builds some of the most capable open weights on Earth from Hangzhou — and understanding it matters precisely because LIWARSE’s safety framework has to work regardless of which government, or which language, a model was trained under.
This closes LIWARSE’s five-part review of the leading closed and open models from the world’s most consequential AI providers. Alibaba’s Qwen team, part of DAMO Academy, has run the same playbook as several Western labs — ship broadly open, then peel off a closed flagship at the top — but compressed into a matter of months rather than years, and against a genuinely different regulatory backdrop.
The Closed Flagship: Qwen3.7-Max
Unveiled at the Apsara Summit on May 20, 2026, Qwen3.7-Max was Alibaba’s first Qwen flagship released as API-only, through Alibaba Cloud’s DashScope platform, with no accompanying open weights — a first for the line. It carries a one-million-token context window, scores competitively against Western closed flagships on reasoning benchmarks including GPQA Diamond, and is explicitly positioned as an “agent frontier” model, with Alibaba demonstrating autonomous runs lasting over thirty hours and more than a thousand tool calls in a single session. A lighter multimodal sibling, Qwen3.7-Plus, followed on June 1 at roughly one-sixth the cost.
For teams outside China, Qwen3.7-Max is accessed exclusively through Alibaba Cloud, which raises the same data-residency and sovereignty questions any organization should ask before routing sensitive data — clinical, research, or otherwise — through any nation’s cloud infrastructure, regardless of which country it belongs to.
The Open Counterpart: Qwen 3.6
Released in April 2026 under the fully permissive Apache 2.0 license, Qwen 3.6 ships as a dense 27B model and a 35B mixture-of-experts variant with only 3 billion active parameters. Both use a hybrid attention architecture, support a native 256,000-token context extensible to roughly one million, and accept text, image, and video input. Its standout claim is genuine: the compact 27B model outperforms Alibaba’s own much larger Qwen 3.5 flagship on agentic coding benchmarks while running on a single consumer GPU — a real efficiency gain, not a marketing one, and it now sits among the strongest self-hostable coding models available from any provider in this series.
Qwen’s open tier carries broad multilingual coverage — reported across roughly 200 languages and dialects — which makes it a particularly relevant option for LIWARSE’s global mission: a low-resource clinic or research institution outside the world’s wealthiest countries can self-host a genuinely capable model without depending on any single nation’s cloud or export policy.
Future Outlook
Alibaba announced Qwen3.8-Max on August 3, 2026 — a roughly 2.4-trillion-parameter model that Alibaba itself describes as “second only to Fable 5” on its own benchmarks, though those figures remain vendor-reported and unverified by independent evaluators as of this writing. Alibaba has separately said it plans to publish open weights for a companion Qwen3.8-27B during the same week, which, if it holds to the Apache 2.0 pattern set by 3.6, would mark the first open release at genuine Max-class scale from any provider in this series. That commitment is not yet fulfilled, and this movement will judge it once the weights, and their license, actually appear — not before.
Risks and Benefits Through the LIWARSE Lens
Benefits
- Qwen 3.6’s genuine efficiency gain — flagship-adjacent performance on a single consumer GPU — does more than any pricing page to democratize access to capable AI for institutions with limited compute budgets.
- Broad multilingual coverage extends the reach of Future Medicine and safety-literacy content into languages and regions that Western-trained models often serve poorly.
- A credible, if unverified, promise of open weights at true flagship scale would be a meaningful escalation of openness relative to every other provider in this series, none of which has open-sourced its actual current-generation flagship.
Risks
- Qwen3.7-Max’s benchmark claims, like Qwen3.8-Max’s, arrive largely through vendor-reported figures; this movement has seen a predecessor model — Qwen3.7-Max’s own precursor — look flagship-tier on vendor numbers and land mid-pack once independently tested, a caution that applies equally to claims from every lab in this series, not Alibaba alone.
- Data-residency and cross-border governance questions apply to Alibaba Cloud exactly as they would to any single nation’s infrastructure hosting a closed frontier model — a structural risk of concentrated closed AI, not a China-specific one.
- An unfulfilled promise of open weights carries no more weight than any other lab’s stated intentions until the license file actually exists; LIWARSE’s own reporting elsewhere in this series has shown labs reverse open-source commitments inside a single product cycle.
The LIWARSE Assessment
Alibaba closes this series on the right note for LIWARSE’s founding purpose: the guarded-versus-unguarded standard, and the 3 Absolute Laws that sit beneath it, do not carry a passport. A model is not safer for being trained in California, and it is not more dangerous for being trained in Hangzhou — what matters, in every case this series has examined, is whether real safety infrastructure was built in before release, and whether independent verification exists to confirm it. Qwen 3.6 stands as genuinely useful, globally accessible open infrastructure; Qwen3.7-Max stands as a capable but unaudited closed system reachable only through one nation’s cloud. Neither claim should be taken purely on faith — from Alibaba, or from any of the four providers examined earlier in this series.
Across all five providers in this series, the same pattern holds: every lab pairs a closed flagship with an open counterpart, and in every case the honest verdict is the same — real progress toward accessible, capable AI, alongside safety commitments that remain unverified, reversible, or simply not yet built. Guarded versus unguarded, accountable versus anonymous — not open versus closed, and not one country versus another — remains the only fault line that actually predicts harm. LIWARSE will keep reviewing this landscape as it changes, because it will keep changing.
— The LIWARSE Movement | liwarse.org
Safety of Life · Advancement of Life · Together.