Qwen 3.8 Max vs GLM 5.2 vs Kimi K3 vs DeepSeek V4 Flash (2026): The Complete Frontier Model Comparison

🎯 Key Takeaways (TL;DR)

  • Kimi K3 (Moonshot, July 16, 2026) is the only one of the four with fully verified third-party scores: 57 on the Artificial Analysis Intelligence Index, #4 of 189 models — the highest ever for an open-weight model. It is also #1 on Arena's Frontend Coding leaderboard (1,679 Elo, 483,895 blind votes).
  • GLM 5.2 (Zhipu, June 13–17, 2026, ~744B params) is the verified coding-and-agent specialist: #1 globally on Code Arena and Design Arena, with MIT weights already on Hugging Face and a budget-friendly $0.29/M token rate on OpenRouter.
  • Qwen 3.8 Max (Alibaba, July 19, 2026, 2.4T params, 95B active) is the newest flagship — the first Max-scale model with open-weight plans, officially claiming to trail "only Fable 5". Its official announcement (Alibaba Qwen WeChat blog) includes benchmark tables, but no third party has scored it yet, so treat all headline numbers as vendor claims for now.
  • DeepSeek V4 Flash 0731 (July 31, 2026, 284B total / 13B active) is the value king: it beats GLM 5.2 on nearly every agentic benchmark on DeepSeek's own Hugging Face table (Terminal-Bench 2.1: 82.7 vs 81.0; DeepSWE: 54.4 vs 46.2) at just $0.14 / $0.28 per million tokens — about 1.1% of Claude Opus 4.8's output price.

⚠️ Accuracy note: This article clearly separates officially claimed numbers (vendor-published) from independently verified numbers (Artificial Analysis, Arena, third-party evals). As of August 3, 2026, Qwen 3.8 Max has official but not yet independently verified benchmarks; Kimi K3, GLM 5.2, and DeepSeek V4 Flash 0731 have both.

📑 Table of Contents

  1. Why This Four-Way Comparison Matters in 2026
  2. At a Glance: Side-by-Side Specification Table
  3. Release Timeline: A Two-Month Race
  4. Benchmarks: Official Claims vs Independent Verification
  5. Agentic & Coding Benchmarks Head-to-Head
  6. Pricing Breakdown: API, OpenRouter, and Cache
  7. Context Window & Multimodal Capabilities
  8. Open Weights, Licensing & Deployment
  9. How to Choose Between the Four Models
  10. Frequently Asked Questions
  11. Final Verdict & Next Steps

Why This Four-Way Comparison Matters in 2026

The summer of 2026 is the first time four Chinese labs shipped frontier-scale models within six weeks of each other. If you are building a product on LLMs today, this Qwen 3.8 Max vs GLM 5.2 vs Kimi K3 vs DeepSeek V4 Flash decision determines your benchmark ceiling, your per-token cost, and whether you can self-host at all. Three reasons this matchup matters:

  1. The open-weight gap closed. Kimi K3 (2.8T), GLM 5.2 (~744B), and Qwen 3.8 Max (2.4T) all ship or promise open weights under permissive licenses, while DeepSeek V4 Flash 0731's 284B weights are expected within weeks. For the first time, frontier-adjacent performance is available off-cloud.
  2. Scale is diverging, not converging. You can now choose between a 2.8-trillion-parameter monster (Kimi K3), a 2.4T multimodal flagship (Qwen 3.8 Max), a 744B coding specialist (GLM 5.2), and a 284B-parameter / 13B-active value model (DeepSeek V4 Flash 0731) that outperforms models many times its size.
  3. Verified data is finally available for three of four. Artificial Analysis now scores Kimi K3 (57), GLM 5.2 (51), and DeepSeek V4 Flash 0731 (50) on its Intelligence Index. Qwen 3.8 Max remains the open question.

💡 Pro Tip: The single most important accuracy rule in 2026: check whether a benchmark number is vendor-published or third-party-verified before you quote it in a blog or a slide deck.

At a Glance: Side-by-Side Specification Table

SpecificationQwen 3.8 Max (Alibaba)GLM 5.2 (Zhipu)Kimi K3 (Moonshot)DeepSeek V4 Flash 0731
Release dateJuly 19, 2026 (preview)June 13–17, 2026July 16, 2026July 31, 2026 (0731 checkpoint)
Total parameters2.4T (official)~744B2.8T284B
Active parameters95B (official)~40B~50B (16 of 896 experts)13B (6 of 256 experts per token)
ArchitectureMoE (based on Qwen 3.5)MoE + upgraded DeepSeek Sparse AttentionStable LatentMoE + KDA hybrid attentionMoE, 1 shared + 256 routed experts, hash routing
Context window1M tokens (official)1M tokens1M tokens1M tokens (1,048,576)
Native modalitiesText, image, video, documentsText-firstText, image, audio, videoText only
Open weightsPromised "next week" (announcement ~Jul 30)MIT, on Hugging Face nowMIT, full weights July 27, 2026Expected "in coming weeks"
AA Intelligence IndexNot yet scored (official claims only)51 (verified)57 (verified, #4 of 189)50 (verified)
OpenRouter IDn/a (Token Plan only)z-ai/glm-5.2moonshotai/kimi-k3deepseek/deepseek-v4-flash-0731

Release Timeline: A Two-Month Race

  • June 13–17, 2026 — Zhipu releases GLM 5.2 (~744B), topping Code Arena and Design Arena within weeks, trained on Huawei Ascend hardware.
  • July 16, 2026 — Moonshot launches Kimi K3 (2.8T) at WAIC Shanghai, the largest open-weight model ever released.
  • July 19, 2026 — Alibaba previews Qwen 3.8 Max (2.4T) at WAIC, claiming "second only to Fable 5".
  • July 27, 2026 — Kimi K3 full MIT weights land on Hugging Face.
  • July 31, 2026 — DeepSeek ships the V4 Flash 0731 checkpoint (same 284B/13B architecture, fully re-post-trained), producing a huge agentic benchmark jump.

Benchmarks: Official Claims vs Independent Verification

The Qwen 3.8 Max vs GLM 5.2 vs Kimi K3 vs DeepSeek V4 Flash benchmark picture has two layers. Here is the verified layer first.

Independently Verified: Artificial Analysis Intelligence Index

ModelIntelligence IndexGlobal RankNotes
Kimi K3 (max)57#4 of 189Highest ever for an open-weight model; ~on par with Claude Opus 4.8 (~56)
GLM 5.2 (max)51Top 10Verified across full eval suite
DeepSeek V4 Flash 0731 (max)50Within 1 point of GLM 5.2; on par with Gemini 3.6 Flash (50); 1 point behind Muse Spark 1.1 (51)
Qwen 3.8 MaxNot scoredNo third-party evaluation published as of Aug 3, 2026

Reference points: Claude Fable 5 ~60 (#1), GPT-5.6 Sol ~59 (#2).

Officially Claimed: Qwen 3.8 Max (Alibaba's Announcement)

Alibaba's official Qwen announcement (published on the Qwen WeChat account, ~July 30, 2026) is the source the community has been asking for. Confirmed specs from that post: 2.4T total parameters, 95B active parameters, 1M-token context, and first Max-scale model with open-weight plans ("weights next week"). The announcement's performance tables are published as images, but the textual results include:

  • E-Commerce Bench (365-day simulated e-commerce, ¥100,000 start): finished with ¥416,252 total cash (4.16× return)38% higher than runner-up GLM 5.2 and 152% better than Qwen 3.7-Max; year-end-promotion net profit was ~2.4× GLM 5.2's.
  • Autonomous chip design (GCD/RSA crypto accelerator): from first working design at 8,298 gates to 678 gates after ~500 interaction rounds — stated as best among all models evaluated.
  • Autonomous research loop: reproduced a paper's full pipeline in ~37 hours, then independently proposed 18 improvements and beat the original method by +2.7 points on AIME24.
  • Long-horizon coding: 265 commits, 127 PRs, 151 issues across ~16 autonomous days on the oh-my-cli project (GitHub: qwen-code-dev-bot/oh-my-cli).

Verdict on Qwen 3.8 Max numbers: officially published and detailed, but not yet independently verified — Artificial Analysis and LMArena have not scored it. Qwen 3.7-Max's verified baseline (AA Index 56.6 at launch; GPQA-Diamond 92.4, SWE-bench Verified 80.4, Terminal-Bench 2.0 69.7) remains the only family-level third-party reference.

⚠️ Note: The only independent data point on Qwen 3.8 Max so far is Trilogy AI's single blind StackPerf run: Qwen3.8-Max-Preview 80 vs Kimi K3 83. One run, one task — a data point, not a verdict.

Agentic & Coding Benchmarks Head-to-Head

DeepSeek's Hugging Face card for V4 Flash 0731 publishes a direct four-way comparison on agentic benchmarks (model-reported, DeepSeek's own harness). It is the cleanest apples-to-apples table currently public:

BenchmarkDS V4 Flash 0731DS V4 Flash (prev)DS V4 Pro (prev)GLM 5.2Opus 4.8
Terminal-Bench 2.182.761.872.181.085.0
NL2Repo54.239.438.548.969.7
Cybergym76.738.752.783.1
DeepSWE54.47.312.846.258.0
Toolathlon-Verified70.349.755.959.976.2
Agents' Last Exam25.215.816.523.825.7
DSBench-FullStack68.737.041.861.871.6
DSBench-Hard59.625.831.154.571.7

Read carefully: DeepSeek V4 Flash 0731 (13B active) beats GLM 5.2 (~40B active) on every row where both appear, and comes within a few points of Claude Opus 4.8 — while its 0731 checkpoint's DeepSWE jumped from 7.3 to 54.4 versus the preview. The Medium/community assessment that the 0731 build sits at "Claude Opus 4.6 level on agentic benchmarks" is consistent with these tables.

For coding specifically, the verified standings are:

Coding BenchmarkWinner
Arena Frontend CodingKimi K3 — 1,679 Elo, #1 (483,895 blind votes; led 6 of 7 domains)
Code Arena (Z.ai blind eval)GLM 5.2 — #1 globally
Design ArenaGLM 5.2 — #1, Elo 1360
SWE-bench VerifiedKimi K3 ~78% / GLM 5.2 77.8% — near-tie
HumanEvalKimi K3 88.3 (family-reported)
Qwen 3.7-Max baseline (SWE-bench Verified)80.4 — the family reference until 3.8 is verified

Pricing Breakdown: API, OpenRouter, and Cache

ModelInput / 1MOutput / 1MCached Input / 1MNotes
DeepSeek V4 Flash 0731$0.14$0.28$0.0028 (98% off)~1.1% of Opus 4.8's output price; 2,500 concurrent requests; max output 384K
GLM 5.2$0.29 (OpenRouter) / ¥8 (Zhipu API)$0.29Free tier via NVIDIA NIM
Kimi K3Premium tier (Moonshot API)Premium tierFree at kimi.com and Kimi Code
Qwen 3.8 MaxPreview credits only (Token Plan / Qoder)Preview credits onlyPreview at 1/10th rate, 1/50th overnight; no public per-token price yet

Cost-per-task reality check: DeepSeek V4 Flash 0731 is dramatically cheaper per token, and it is 12% more token-efficient than its predecessor (206M vs 234M output tokens for the same Intelligence Index run). Kimi K3's reasoning depth often solves tasks in fewer total tokens, which can offset its premium rate. Always benchmark end-to-end cost per task, not just the rate card.

Context Window & Multimodal Capabilities

All four models support a 1M-token context window — one of the quiet convergences of 2026. The multimodal story differs sharply:

ModelMultimodal?Details
Kimi K3✅ NativeText, image, audio, video in one prompt — the broadest single-model modality support
Qwen 3.8 Max✅ NativeFirst Qwen multimodal above 1T params; official demos cover 200+ page reports, 100+ hour videos ("Video Graph memory"), Vlog editing, screenshot-to-frontend, Blender 3D; RecreationBench for app replication; Qwen-MM-Plugins toolkit
GLM 5.2⚠️ Text-firstVision via companion GLM-4.5V family, not a unified architecture
DeepSeek V4 Flash 0731❌ Text onlyConfirmed by Artificial Analysis: text input/output only

If your workload needs see + reason + act in one model, the choice narrows to Kimi K3 or Qwen 3.8 Max. If you run pure text agents, DeepSeek V4 Flash 0731's narrow modality surface is a feature: smaller deployment, lower cost, no vision tax.

Open Weights, Licensing & Deployment

ModelLicenseWeights StatusSelf-Host Footprint
Kimi K3MITLive since July 27, 2026 (moonshotai/Kimi-K3)Multi-node H100/MI300 (≥8 GPU); vLLM / SGLang
GLM 5.2MITLive now (zai-org/GLM-5.2)Single-node 8×H100 or 4×MI300; vLLM, SGLang, llama.cpp
Qwen 3.8 MaxTBDPromised "next week" (announcement ~Jul 30); no license stated yetUnclear until weights drop; 95B active suggests 8-GPU-class nodes
DeepSeek V4 Flash 0731Expected permissive"Coming weeks" per Artificial Analysis13B active — runs on small clusters; community reports 2×DGX Spark setups

Best Practice: If you need open weights today, pick GLM 5.2 (744B) or Kimi K3 (2.8T). If you want the smallest self-hostable frontier-adjacent model, wait for DeepSeek V4 Flash 0731 weights — 13B active is a dramatically cheaper serving story. Qwen 3.8 Max's weight release date and license were not yet confirmed as of August 3, 2026.

How to Choose Between the Four Models

Recommendations by workload:

  • Production coding agent, today, with verified evalsKimi K3 (AA 57, Frontend Coding #1) or GLM 5.2 (Code Arena #1, cheaper per token).
  • Agentic loops on a tight budgetDeepSeek V4 Flash 0731 — beats GLM 5.2 on DeepSeek's agentic table at ~2% of the premium cost.
  • Autonomous long-horizon work (research, e-commerce, chip design) with vendor evidenceQwen 3.8 Max — the official case studies (4.16× E-Commerce return, 678-gate chip) are the strongest published demos; just verify before you scale.
  • Native multimodal single modelKimi K3 (4 modalities) or Qwen 3.8 Max (vision + video + GUI agents).
  • Self-host on commodity hardwareDeepSeek V4 Flash 0731 (13B active) or GLM 5.2 (single 8-GPU node).

🤔 Frequently Asked Questions

Q: Which of the four models has the best verified benchmarks?

A: Kimi K3 — 57 on the Artificial Analysis Intelligence Index, #4 of 189 models, the highest ever recorded by an open-weight model, plus #1 on Arena Frontend Coding (1,679 Elo). GLM 5.2 (51) and DeepSeek V4 Flash 0731 (50) are close to each other on the Index.

Q: Is Qwen 3.8 Max really "second only to Fable 5"?

A: That is Alibaba's official claim, with detailed internal benchmark tables and case studies published in its announcement (E-Commerce Bench 4.16× return, 38% above GLM 5.2; 152% above Qwen 3.7-Max; chip design to 678 gates). No third party has verified it yet as of August 3, 2026. Treat it as a strong vendor claim pending independent scores.

Q: How can DeepSeek V4 Flash 0731 (13B active) beat GLM 5.2 (40B active)?

A: The 0731 checkpoint is the same 284B/13B architecture as V4 Flash but fully re-post-trained, which DeepSeek's Hugging Face table shows delivering large agentic gains (DeepSWE 7.3 → 54.4; Terminal-Bench 2.1 61.8 → 82.7). Post-training quality, not raw scale, drives its agentic performance — that is the 2026 lesson this model teaches.

Q: What is the cheapest model here?

A: DeepSeek V4 Flash 0731: $0.14/M input, $0.28/M output, and $0.0028/M cached input (98% discount) — roughly 1.1% of Claude Opus 4.8's output price, with a 2,500-concurrency limit. GLM 5.2 at $0.29/M on OpenRouter is second.

Q: Which models can I run locally?

A: GLM 5.2 (MIT weights live) and Kimi K3 (MIT weights live since July 27) today. DeepSeek V4 Flash 0731 weights are expected in coming weeks and are the easiest to serve (13B active). Qwen 3.8 Max promised weights "next week" in its announcement, but no date or license was confirmed as of August 3, 2026.

Q: Do all four support 1M-token context?

A: Yes — Qwen 3.8 Max (officially), GLM 5.2, Kimi K3, and DeepSeek V4 Flash 0731 (1,048,576 tokens) all ship 1M-token context windows.

Q: Which is best for multimodal tasks?

A: Kimi K3 (native text/image/audio/video) and Qwen 3.8 Max (native image/video/documents plus GUI-agent capabilities like RecreationBench). GLM 5.2 is text-first (vision via companion models), and DeepSeek V4 Flash 0731 is text-only.

Final Verdict & Next Steps

The Qwen 3.8 Max vs GLM 5.2 vs Kimi K3 vs DeepSeek V4 Flash comparison in August 2026 resolves into four distinct bets:

  • Bet on verification: Kimi K3 — the only model here with frontier-level third-party scores.
  • Bet on coding + open weights today: GLM 5.2 — Code Arena #1, MIT weights live, cheap tokens.
  • Bet on the newest flagship: Qwen 3.8 Max — biggest official benchmark package of the year, but verify before scaling.
  • Bet on cost per task: DeepSeek V4 Flash 0731 — agentic performance near Opus 4.8 at 1/100th the price.

Three-step action plan

  1. This week: Route agentic coding traffic through DeepSeek V4 Flash 0731 (deepseek/deepseek-v4-flash-0731 on OpenRouter, or DeepSeek API at $0.14/$0.28) and benchmark against your current stack — the 98% cache discount makes repo-context loops almost free.
  2. This month: Evaluate Kimi K3 (AA 57) for your hardest reasoning/multimodal workloads, and watch for Qwen 3.8 Max open weights and first third-party scores — those two events will settle the biggest open question in the comparison.
  3. Re-check in 2–4 weeks: DeepSeek's open-weight release and Qwen 3.8 Max's independent evals will both land; re-run this comparison then, because the gap between official claims and verified reality is exactly where this race will be won.

Sources

  • Alibaba Qwen official announcement (WeChat, ~July 30, 2026): Qwen3.8-Max specs (2.4T params, 95B active, 1M context), E-Commerce Bench, chip design, oh-my-cli case study — mp.weixin.qq.com/s/9S8VZETppDj_AidiUFrZlw
  • DeepSeek-V4-Flash-0731 model card (Hugging Face) — agentic benchmark table vs V4 Pro, GLM-5.2, Opus-4.8: huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731
  • Artificial Analysis — Intelligence Index scores: DeepSeek V4 Flash 0731 (50), GLM-5.2 (51), Kimi K3 (57), Fable 5 (~60), GPT-5.6 Sol (~59)
  • Arena (LMArena) Frontend Coding leaderboard — Kimi K3 1,679 Elo; Z.ai Code Arena / Design Arena — GLM 5.2 #1
  • Moonshot AI Kimi K3 release coverage; OpenRouter model pages; DeepSeek official pricing page ($0.14/$0.28, $0.0028 cached); Trilogy AI StackPerf run (Qwen3.8-Max 80 vs Kimi K3 83)

Last updated: 2026-08-03. All vendor-reported numbers are labeled as claims where third-party verification is pending.

PSL Scale: Curious what AI thinks of your face?Try For Free