The strongest Chinese AI models in October 2026 are Xiaomi's MiMo-V2.6-Pro, Alibaba's Qwen3.8 Max, Z.ai's GLM-5.3 and Moonshot AI's Kimi K3, at 44 to 46 on the Artificial Analysis Intelligence Index. DeepSeek V4.1 Flash scores 39 but is by far the fastest. The top American models score 52 to 58.
I posted on September 17 that Chinese models had caught up to GPT-5.6, and named four of them. The post drew 218 reactions and 134 comments. Eighteen days later, the honest update: three of the four caught up to GPT-5.6 Terra at 42, and all four still trail GPT-5.6 Sol at 47. A newer model now beats all of them. And per task, the cheap ones are not always cheap.
- Xiaomi's MiMo-V2.6-Pro leads at 46, with open weights and a cost of $0.13 per task.
- Qwen3.8 Max (45) is strongest on images and charts, GLM-5.3 (45) on agentic coding and Kimi K3 (44) on long documents.
- DeepSeek V4.1 Flash (39) is the fastest, at 209 tokens per second.
- Per task, GPT-6.1 Sol at $0.72 costs less than GLM-5.3, Kimi K3 or Qwen3.8 Max.
- The top American models: Claude Opus 5.5 at 58, Claude Sonnet 5.5 at 56, GPT-6.1 Sol at 52.
in
“Chinese AI models caught up to GPT-5.6. The 4 best models and when to use each:”
The Best Chinese AI Models Compared
| Model | Maker | Index | Input / output price | Cost per task | Open weights |
|---|---|---|---|---|---|
| MiMo-V2.6-Pro | Xiaomi | 46 | $0.435 / $0.87 | $0.13 | Yes, MIT license |
| Qwen3.8 Max | Alibaba | 45 | $2 / $6 | $5.41 | API only |
| GLM-5.3 | Z.ai (Zhipu) | 45 | $1.40 / $4.40 | $2.01 | Yes, custom license |
| Step 5 Preview | StepFun | 44 | $1 / $2.70 | $0.72 | No |
| Kimi K3 | Moonshot AI | 44 | $3 / $15 | $2.00 | Yes, custom license |
| DeepSeek V4.1 Flash | DeepSeek | 39 | $0.30 / $1.20 | $0.27 | Yes, MIT license |
| Claude Opus 5.5 | Anthropic | 58 | $4 / $20 | $5.98 | No |
| GPT-6.1 Sol | OpenAI | 52 | $2 / $10 | $0.72 | No |
| GPT-5.6 Sol | OpenAI | 47 | $4 / $20 | $1.99 | No |
Two things jump out. The gap at the top is real: the best Chinese score is 46, against 58 for Opus 5.5. And price per token can mislead. Qwen3.8 Max lists at $2 per million input tokens, yet costs $5.41 per task, more than GPT-6 Astra at $3.26, because it spends more tokens to finish. Only MiMo-V2.6-Pro and DeepSeek V4.1 Flash clearly undercut GPT-6.1 Sol per task.
The scores and per-task costs come from Artificial Analysis, an independent benchmark that runs every model through the same tests. The leaderboard moves often, so check it again before you commit a workflow.
MiMo-V2.6-Pro: The New Leader
Xiaomi released MiMo-V2.6-Pro with open weights under the MIT license on September 21, four days after my post. It scores 46, the highest Chinese score on the index, at $0.435 per million input tokens and $0.87 output. Its cost per task, $0.13, is a fraction of every model near its score. For general reasoning work on a budget, it is the one I would test first.
GLM-5.3: The Engineer
Z.ai released GLM-5.3 on August 14 and published the weights on August 28. It is the strongest Chinese model on agentic coding, with 41.9% on Terminal-Bench 4.0, ahead of GPT-5.6 Sol at 39.9%. It reads up to a million tokens of text, but no images. Use it for long coding jobs where a bug hides across many files and the fix takes hours.
Kimi K3: The Analyst
Moonshot AI released Kimi K3 on July 16. It has 2.8 trillion parameters, 104 billion of them active, and reads text and images across a million tokens. Its standout result is long-context reasoning: 88.7% on AA-LCR, the highest in the comparison above. Give it documents and spreadsheets and ask for a report. Moonshot itself says K3 still trails Fable 5 and GPT-5.6 Sol.
Kimi K3 is also the most expensive of these models per token, at $3 for input and $15 for output, more than Sonnet 5.5 or GPT-6.1 Sol. Its cost per task, $2.00, is nearly three times GPT-6.1 Sol's. Keep it for the long-document work where that reasoning score earns the price.
Qwen3.8 Max: Mixed Work
Alibaba launched Qwen3.8 Max on August 3 and updated it on September 2; that update scores 45. The API version reads text, images and video, and it beats the other Chinese leaders on image understanding with 82.8% on MMMU-Pro. Use it when the job mixes screenshots, charts and text, such as a long PDF full of tables. The open weights are the same base model without image input, and score 40.
DeepSeek V4.1 Flash: Speed
DeepSeek released V4.1 Flash on September 10 under the MIT license. It is the fastest model here, at 209 tokens per second, and costs $0.30 per million input tokens and $1.20 output. Off-peak prices are half that, and off-peak covers the whole US working day. It scores 39, so use it for scripts, quick fixes and high-volume jobs where speed and cost matter more than depth.
What Changed Since My Post
Anthropic launched Claude Opus 5.5 on September 22 and Sonnet 5.5 on September 28, and both now score above Fable 5.1 and GPT-6 Astra. So my line about keeping Fable and Astra for the hardest problems is out of date: Opus 5.5 scores higher than both at a lower price per token. And Xiaomi's MiMo-V2.6-Pro arrived on September 21 and took the top Chinese spot.
If you are pricing the American side of that comparison, my guide to Claude API pricing covers every Claude model per million tokens, with caching and batch discounts.
Other Chinese AI Labs to Watch
The list goes beyond the big four names. StepFun's Step 5 Preview scores 44 and matches GPT-6.1 Sol's cost per task. Further down the same index sit MiniMax-M3 at 29, Tencent's Hy3 at 25, ByteDance's Doubao Seed Code at 17 and Baidu's ERNIE 5.0 Thinking Preview at 14. Xiaomi's smaller MiMo-V2.6-Flash scores 38 at just $0.06 per task.
Risks Before You Send Company Data
Check where the data goes. DeepSeek says it stores personal data in China. Kimi stores API data on servers in Singapore, and its policy allows using data to improve its models. Z.ai says its API does not store your content. Alibaba lets you pick a region, including US Virginia, though Qwen3.8 Max may run outside it. If residency matters, self-host the open weights or use a US-hosted provider.
Check the rules that apply to you. Section 1532 of the US defense budget law for 2026 bars Defense Department contractors from using DeepSeek on that work, however it is hosted. NIST's Center for AI Standards and Innovation found GLM-5.3 the most cyber-capable open-weight model, about four months behind the US frontier on cyber tasks. On AA-Omniscience, which scores knowledge and penalizes hallucination, the Chinese leaders score 20 or lower against 46 for Opus 5.5.
Before you move a real workflow, run ten of your own jobs through the AI model evaluation framework and compare accepted outputs, review time and cost per finished job.
Which Chinese AI Model Should Your Team Try First?
My order would be simple. Start with MiMo-V2.6-Pro for general work on a budget and DeepSeek V4.1 Flash for scripts and high-volume jobs where speed decides. Try GLM-5.3 for long coding work, ideally self-hosted. Use Kimi K3 for long documents and Qwen3.8 Max for screenshots and charts. Keep Opus 5.5 for the hardest work and for anything with sensitive data.
Most of a team's day is reports, scripts, fixes and reading files, and these models already handle that work. Measure the cost per finished job before you switch, because a low price per token does not always mean a low bill.
Every Thursday, AI Frontier gives B2B operators one verified signal, my read on it and one practical AI revenue play in under five minutes. The original post and discussion are on LinkedIn. Tell me which of these models your team would try first, and on what job.
Frequently asked questions
What are the best Chinese AI models?
In October 2026 the strongest on the Artificial Analysis Intelligence Index are Xiaomi's MiMo-V2.6-Pro at 46, Alibaba's Qwen3.8 Max and Z.ai's GLM-5.3 at 45, and StepFun's Step 5 Preview and Moonshot AI's Kimi K3 at 44. DeepSeek V4.1 Flash scores 39 and is the fastest.
Are Chinese AI models as good as GPT and Claude?
Close, but behind at the top. The best Chinese models score 44 to 46 on the Artificial Analysis index, against 58 for Claude Opus 5.5, 56 for Claude Sonnet 5.5 and 52 for GPT-6.1 Sol. Several of them beat older American models such as GPT-5.6 Terra at 42.
Which Chinese AI model is the cheapest?
Per token, DeepSeek V4.1 Flash is among the cheapest at $0.30 per million input tokens and $1.20 output, with half price off-peak. Per task on the Artificial Analysis index, Xiaomi's models cost least: MiMo-V2.6-Flash at $0.06 and MiMo-V2.6-Pro at $0.13.
Are Chinese AI models open source?
Most release open weights. MiMo-V2.6-Pro and DeepSeek V4.1 Flash use the MIT license. GLM-5.3 and Kimi K3 use custom licenses with extra terms for companies that resell them as an API above certain revenue levels. Qwen3.8 Max is API only, while its base model is published as open weights.
Is it safe to use DeepSeek for business?
It depends on your data and your rules. DeepSeek says it stores personal data in China. Self-hosting the open weights keeps data under your control, but US Defense Department contractors are barred from using DeepSeek on that work however it is hosted.
Sources
Artificial Analysis, LLM leaderboard, Intelligence Index v4.3.2, read October 5, 2026.
Z.ai, GLM-5.3 announcement, API pricing and license, read October 5, 2026.
Moonshot AI, Kimi K3 announcement, pricing and privacy policy, read October 5, 2026.
Alibaba Cloud, Alibaba unveils Qwen3.8 Max, and Model Studio pricing, read October 5, 2026.
DeepSeek, API pricing and off-peak hours, release notes and privacy policy, read October 5, 2026.
NIST, CAISI assessment of GLM-5.3 cyber capabilities, September 2026.
King & Spalding, FY 2026 NDAA: artificial intelligence provisions, read October 5, 2026.

