Growth Cab Apply to GC
Blog/AI STRATEGY
AI STRATEGY · October 5, 2026 · 8 MIN READ

Best Chinese AI Models in 2026: Scores, Prices and When to Use Each

MiMo-V2.6-Pro, Qwen3.8 Max, GLM-5.3, Kimi K3 and DeepSeek V4.1 Flash on one independent index: scores, cost per task, licenses, data risks and the job each one fits.

Federico DonatoneBy · Founder & CEO, Growth Cab
Best Chinese AI Models in 2026: Scores, Prices and When to Use Each

The strongest Chinese AI models in October 2026 are Xiaomi's MiMo-V2.6-Pro, Alibaba's Qwen3.8 Max, Z.ai's GLM-5.3 and Moonshot AI's Kimi K3, at 44 to 46 on the Artificial Analysis Intelligence Index. DeepSeek V4.1 Flash scores 39 but is by far the fastest. The top American models score 52 to 58.

I posted on September 17 that Chinese models had caught up to GPT-5.6, and named four of them. The post drew 218 reactions and 134 comments. Eighteen days later, the honest update: three of the four caught up to GPT-5.6 Terra at 42, and all four still trail GPT-5.6 Sol at 47. A newer model now beats all of them. And per task, the cheap ones are not always cheap.

Federico Donatonein
Federico Donatone
Founder & CEO, Growth Cab · This article started as a LinkedIn post

“Chinese AI models caught up to GPT-5.6. The 4 best models and when to use each:”

218REACTIONS
134COMMENTS
Read the original post →

The Best Chinese AI Models Compared

Artificial Analysis Intelligence Index v4.3.2 at max effort, API price per million tokens and Artificial Analysis cost per task, read October 5, 2026
ModelMakerIndexInput / output priceCost per taskOpen weights
MiMo-V2.6-ProXiaomi46$0.435 / $0.87$0.13Yes, MIT license
Qwen3.8 MaxAlibaba45$2 / $6$5.41API only
GLM-5.3Z.ai (Zhipu)45$1.40 / $4.40$2.01Yes, custom license
Step 5 PreviewStepFun44$1 / $2.70$0.72No
Kimi K3Moonshot AI44$3 / $15$2.00Yes, custom license
DeepSeek V4.1 FlashDeepSeek39$0.30 / $1.20$0.27Yes, MIT license
Claude Opus 5.5Anthropic58$4 / $20$5.98No
GPT-6.1 SolOpenAI52$2 / $10$0.72No
GPT-5.6 SolOpenAI47$4 / $20$1.99No

Two things jump out. The gap at the top is real: the best Chinese score is 46, against 58 for Opus 5.5. And price per token can mislead. Qwen3.8 Max lists at $2 per million input tokens, yet costs $5.41 per task, more than GPT-6 Astra at $3.26, because it spends more tokens to finish. Only MiMo-V2.6-Pro and DeepSeek V4.1 Flash clearly undercut GPT-6.1 Sol per task.

The scores and per-task costs come from Artificial Analysis, an independent benchmark that runs every model through the same tests. The leaderboard moves often, so check it again before you commit a workflow.

MiMo-V2.6-Pro: The New Leader

Xiaomi released MiMo-V2.6-Pro with open weights under the MIT license on September 21, four days after my post. It scores 46, the highest Chinese score on the index, at $0.435 per million input tokens and $0.87 output. Its cost per task, $0.13, is a fraction of every model near its score. For general reasoning work on a budget, it is the one I would test first.

GLM-5.3: The Engineer

Z.ai released GLM-5.3 on August 14 and published the weights on August 28. It is the strongest Chinese model on agentic coding, with 41.9% on Terminal-Bench 4.0, ahead of GPT-5.6 Sol at 39.9%. It reads up to a million tokens of text, but no images. Use it for long coding jobs where a bug hides across many files and the fix takes hours.

Kimi K3: The Analyst

Moonshot AI released Kimi K3 on July 16. It has 2.8 trillion parameters, 104 billion of them active, and reads text and images across a million tokens. Its standout result is long-context reasoning: 88.7% on AA-LCR, the highest in the comparison above. Give it documents and spreadsheets and ask for a report. Moonshot itself says K3 still trails Fable 5 and GPT-5.6 Sol.

Kimi K3 is also the most expensive of these models per token, at $3 for input and $15 for output, more than Sonnet 5.5 or GPT-6.1 Sol. Its cost per task, $2.00, is nearly three times GPT-6.1 Sol's. Keep it for the long-document work where that reasoning score earns the price.

Qwen3.8 Max: Mixed Work

Alibaba launched Qwen3.8 Max on August 3 and updated it on September 2; that update scores 45. The API version reads text, images and video, and it beats the other Chinese leaders on image understanding with 82.8% on MMMU-Pro. Use it when the job mixes screenshots, charts and text, such as a long PDF full of tables. The open weights are the same base model without image input, and score 40.

DeepSeek V4.1 Flash: Speed

DeepSeek released V4.1 Flash on September 10 under the MIT license. It is the fastest model here, at 209 tokens per second, and costs $0.30 per million input tokens and $1.20 output. Off-peak prices are half that, and off-peak covers the whole US working day. It scores 39, so use it for scripts, quick fixes and high-volume jobs where speed and cost matter more than depth.

What Changed Since My Post

Anthropic launched Claude Opus 5.5 on September 22 and Sonnet 5.5 on September 28, and both now score above Fable 5.1 and GPT-6 Astra. So my line about keeping Fable and Astra for the hardest problems is out of date: Opus 5.5 scores higher than both at a lower price per token. And Xiaomi's MiMo-V2.6-Pro arrived on September 21 and took the top Chinese spot.

If you are pricing the American side of that comparison, my guide to Claude API pricing covers every Claude model per million tokens, with caching and batch discounts.

Other Chinese AI Labs to Watch

The list goes beyond the big four names. StepFun's Step 5 Preview scores 44 and matches GPT-6.1 Sol's cost per task. Further down the same index sit MiniMax-M3 at 29, Tencent's Hy3 at 25, ByteDance's Doubao Seed Code at 17 and Baidu's ERNIE 5.0 Thinking Preview at 14. Xiaomi's smaller MiMo-V2.6-Flash scores 38 at just $0.06 per task.

Risks Before You Send Company Data

Check where the data goes. DeepSeek says it stores personal data in China. Kimi stores API data on servers in Singapore, and its policy allows using data to improve its models. Z.ai says its API does not store your content. Alibaba lets you pick a region, including US Virginia, though Qwen3.8 Max may run outside it. If residency matters, self-host the open weights or use a US-hosted provider.

Check the rules that apply to you. Section 1532 of the US defense budget law for 2026 bars Defense Department contractors from using DeepSeek on that work, however it is hosted. NIST's Center for AI Standards and Innovation found GLM-5.3 the most cyber-capable open-weight model, about four months behind the US frontier on cyber tasks. On AA-Omniscience, which scores knowledge and penalizes hallucination, the Chinese leaders score 20 or lower against 46 for Opus 5.5.

Before you move a real workflow, run ten of your own jobs through the AI model evaluation framework and compare accepted outputs, review time and cost per finished job.

Which Chinese AI Model Should Your Team Try First?

My order would be simple. Start with MiMo-V2.6-Pro for general work on a budget and DeepSeek V4.1 Flash for scripts and high-volume jobs where speed decides. Try GLM-5.3 for long coding work, ideally self-hosted. Use Kimi K3 for long documents and Qwen3.8 Max for screenshots and charts. Keep Opus 5.5 for the hardest work and for anything with sensitive data.

AI FRONTIER
Get one useful AI play every Thursday
The AI changes that matter, Federico's direct read and one practical play, plus the free 10-page Operator Pack.
Free · under five minutes · unsubscribe anytime

Most of a team's day is reports, scripts, fixes and reading files, and these models already handle that work. Measure the cost per finished job before you switch, because a low price per token does not always mean a low bill.

Every Thursday, AI Frontier gives B2B operators one verified signal, my read on it and one practical AI revenue play in under five minutes. The original post and discussion are on LinkedIn. Tell me which of these models your team would try first, and on what job.

Frequently asked questions

What are the best Chinese AI models?

In October 2026 the strongest on the Artificial Analysis Intelligence Index are Xiaomi's MiMo-V2.6-Pro at 46, Alibaba's Qwen3.8 Max and Z.ai's GLM-5.3 at 45, and StepFun's Step 5 Preview and Moonshot AI's Kimi K3 at 44. DeepSeek V4.1 Flash scores 39 and is the fastest.

Are Chinese AI models as good as GPT and Claude?

Close, but behind at the top. The best Chinese models score 44 to 46 on the Artificial Analysis index, against 58 for Claude Opus 5.5, 56 for Claude Sonnet 5.5 and 52 for GPT-6.1 Sol. Several of them beat older American models such as GPT-5.6 Terra at 42.

Which Chinese AI model is the cheapest?

Per token, DeepSeek V4.1 Flash is among the cheapest at $0.30 per million input tokens and $1.20 output, with half price off-peak. Per task on the Artificial Analysis index, Xiaomi's models cost least: MiMo-V2.6-Flash at $0.06 and MiMo-V2.6-Pro at $0.13.

Are Chinese AI models open source?

Most release open weights. MiMo-V2.6-Pro and DeepSeek V4.1 Flash use the MIT license. GLM-5.3 and Kimi K3 use custom licenses with extra terms for companies that resell them as an API above certain revenue levels. Qwen3.8 Max is API only, while its base model is published as open weights.

Is it safe to use DeepSeek for business?

It depends on your data and your rules. DeepSeek says it stores personal data in China. Self-hosting the open weights keeps data under your control, but US Defense Department contractors are barred from using DeepSeek on that work however it is hosted.

Sources

Artificial Analysis, LLM leaderboard, Intelligence Index v4.3.2, read October 5, 2026.

Z.ai, GLM-5.3 announcement, API pricing and license, read October 5, 2026.

Moonshot AI, Kimi K3 announcement, pricing and privacy policy, read October 5, 2026.

Alibaba Cloud, Alibaba unveils Qwen3.8 Max, and Model Studio pricing, read October 5, 2026.

DeepSeek, API pricing and off-peak hours, release notes and privacy policy, read October 5, 2026.

NIST, CAISI assessment of GLM-5.3 cyber capabilities, September 2026.

King & Spalding, FY 2026 NDAA: artificial intelligence provisions, read October 5, 2026.

Want a GTM engine that runs like this?

Growth Cab is the #1 GTM & sales advisory in the US & Europe. We build the outbound, LinkedIn, and closing systems behind these playbooks for founders selling high-ACV deals.

Apply to GC ← All articles
AI FRONTIER

Turn this week's AI noise
into one useful move

Every Thursday: the signal, Federico's direct view and one practical play. Join free and get the 10-page AI Frontier Operator Pack.

Free · under five minutes · unsubscribe anytime