China's Second Wave: Qwen, Kimi, and the Labs You've Underestimated
DeepSeek got the headlines, but Alibaba's Qwen and Moonshot's Kimi are shipping frontier-class models at a startling pace. A look at the summer of 2026 — and what the West still misunderstands.
The West's mental map of Chinese AI is out of date. For most of the last two years, "Chinese AI" has meant one thing: DeepSeek — the quant-fund spinoff that crashed the AI cost curve in early 2025. But the second half of 2026 has belonged to a different cast. Alibaba's Qwen team, Moonshot AI, and Zhipu AI (Z.ai) have compressed what used to be a quarterly release cycle into something closer to a monthly one, and the models they're shipping are no longer "good for China." They're among the best in the world, on independent benchmarks, with downloadable weights.
What the West misses isn't the capability. It's the strategy.
The thirty-day window#
Between July 27 and August 25, 2026, four frontier models shipped from four different Chinese labs. Kimi K3 on July 16 (weights July 27), DeepSeek V4-Flash on July 31, Qwen3.8-Max weights on August 12, and GLM-5.3-Flash on August 25. Each from a different lab, each with a different architecture, each with a different licensing strategy. As Forkast's analysis of the period put it, the pattern isn't accidental — it reveals a deliberate bifurcation: cheap, permissively licensed "Flash" models for adoption, revenue-gated "Max" models for enterprise capture.
That's worth pausing on. The Western narrative treats Chinese labs as a single undifferentiated force. The reality is closer to a competitive market where labs compete with each other for dominance — Chinese social media greeted DeepSeek V4 Pro's benchmark results with a characteristic verdict: "DeepSeek finally lost its throne as the open-source king, but the successor still comes from China."
Kimi K3: the startup that outran the giants#
Moonshot AI is the lab most Western observers would have placed last. Founded in Beijing as a consumer-assistant startup, it doesn't have Alibaba's cloud, ByteDance's scale, or DeepSeek's mystique. It just shipped the largest open-weight model ever released.
Kimi K3, released July 16, 2026 (weights July 27), is a 2.8-trillion-parameter sparse mixture-of-experts model with only 16 of 896 experts active per token — roughly 104 billion active parameters per forward pass — a 1-million-token context window, and native vision. Moonshot introduced two architectural innovations it calls Kimi Delta Attention and Attention Residuals, credited with faster decoding and better training efficiency.
The independent numbers are what matter: on the Artificial Analysis Intelligence Index (v4.1), K3 scored 57.1 — fourth of all models tested, behind only Claude Fable 5 and GPT-5.6 Sol, and ahead of Claude Opus 4.8. On Arena's Frontend Code Arena, K3 scored 1,679, ranking first globally — ahead of Claude Fable 5 — and Moonshot reported 88.3 on Terminal-Bench 2.1. (Note: Artificial Analysis re-based its index in September 2026, pulling every score down; relative rankings stayed intact, with K3 level with the Western mid-frontier.)
The honest caveat: K3 is priced like a Western frontier model — $3/$15 per million input/output tokens — triple its predecessor. And its weights ship under the Kimi K3 License, a bespoke license that permits commercial use with conditions. Open, yes; unconditional, no.
Alibaba's Qwen: the most aggressive release machine in the industry#
Alibaba's Qwen team has become the highest-velocity lab in AI, full stop. Consider the cadence: Qwen3.7-Max flagship at the Apsara Conference in May 2026; the 2.4-trillion-parameter Qwen3.8-Max in August; the Qwen3.8-Flash series later that month; and Qwen3.8-Omni-Flash, an omni-modal model built around agentic audio-video understanding, in September.
The specs are staggering on their face. Qwen3.8-Max is a sparse MoE with 2.4 trillion total parameters but only ~95 billion active per token, a 1-million-token context window, and full multimodality. Qwen3.8-Flash-Next — released as an open-weight preview of the Qwen4 architecture — uses Gated DeltaNet to compress long context into a fixed recurrent state, with Alibaba claiming one-ninth the training cost of Qwen3.7-Plus at comparable coding and reasoning performance.
But Qwen deserves scrutiny alongside admiration. Alibaba's marketing claims — "second only to Fable 5" — have shipped with thin independent verification; some announcements included no benchmarks or model cards at all. The benchmark figures that are public (80.4 on SWE-Verified, 92.4 on GPQA Diamond for Qwen3.7-Max) are strong, but they're Alibaba's numbers, and independent replication matters. The pattern at Alibaba is clear: announce the flagship, then stagger the evidence.
What is verifiable is the strategy. Qwen is a licensing ladder: Apache 2.0 for the small models (Qwen3.8-27B is the one to actually download), bespoke revenue-sharing licenses for the big Max weights, closed API-only for the newest flagships. The smaller Qwen models remain among the most widely downloaded open models on Hugging Face — ecosystem dominance through the bottom of the lineup, margin capture at the top.
The underestimated rest: GLM, DeepSeek, and the long tail#
Zhipu's GLM series may be the most underrated models in the world. GLM-5.2 shipped MIT-licensed in June 2026; GLM-5.3 followed in August under a custom license, with its Flash variant (321B total, 18B active, MIT, $0.15/$0.50 pricing) offering arguably the best license-to-capability ratio of the entire summer cycle. On the re-based Artificial Analysis index, GLM-5.3 sits above Kimi K3. Yet it gets a fraction of the Western press.
DeepSeek, meanwhile, has settled into its permanent role: the price-killer. V4 Pro supports the Responses API and Code Interpreter as a direct competitor to OpenAI's agentic offerings, with prices an order of magnitude below Western equivalents. Its weights are MIT — genuinely open, no conditions. The company is reportedly raising ~50B yuan at a ~$74B pre-money valuation ahead of a STAR Market listing expected in 2027. DeepSeek isn't a disruption story anymore; it's a company story.
And don't sleep on the long tail: Xiaomi's MiMo V2.5, MiniMax M3, Tencent's Hunyuan Hy4, and an unconfirmed ByteDance 10-trillion-parameter pre-training run. The bench is deep.
What the West gets wrong#
Three misconceptions, corrected.
1. "It's just distillation of Western models." This one won't die. In 2026, Anthropic publicly accused the largest Chinese labs of distilling Claude. Whether or not specific cases are true, the accusation no longer fits the trajectory: Kimi K3 leads on frontend coding, GLM-5.3 leads the open-weight index, and Chinese labs are publishing genuinely novel architecture work — Kimi Delta Attention, Gated DeltaNet, Qwen Sparse Attention. The gap narrative has to be updated: on independent measures, the open-weight frontier is now Chinese, roughly 8 points behind the closed US frontier on the current index.
2. "Sanctions will stop them." Cut off from top-tier chips, Chinese labs went open. Releasing weights is a strategy, not just a philosophy: it harvests external feedback, captures the Global South (Malaysia trending DeepSeek, Singapore trending Qwen), and turns NVIDIA H20-class hardware — the chips they can get — into a served ecosystem. Open weights carried 62% of Vercel's AI Gateway tokens in August 2026, with DeepSeek V4-Flash first. The sanctions shaped the ecosystem; they didn't strangle it.
3. "Open weights mean free-for-all." The most important development of the summer is the quiet end of unconditional openness. Kimi K3's bespoke license and Qwen's revenue-sharing tiers gate the biggest models: fine for you, not fine for a competing model-as-a-service. Only DeepSeek and GLM-Flash held the pure MIT line. "Open" is becoming a licensing spectrum, and the biggest weights now come with commercial strings.
The takeaway#
DeepSeek was the warning shot; the second wave is the campaign. Moonshot proved a startup can beat tech giants to the frontier. Alibaba proved release velocity is a strategy in itself — ship so fast the benchmark regime can't keep up. Zhipu proved you don't need the headlines to win the license-to-capability game.
For builders, the practical implication is blunt: the best open-weight models you can actually download are now mostly Chinese, and choosing among them requires reading licenses the way you read model cards. For everyone else, the implication is blunter still — the assumption that the frontier is an American property is no longer a fact. It's a habit.