Alibaba, DeepSeek push China’s AI model race towards lower costs


Alibaba has launched Qwen3.8-Max, its largest AI model to date, as DeepSeek’s latest V4-Flash model draws attention for inference pricing that is lower than several competing systems.

Qwen3.8-Max has 2.4 trillion parameters and uses a mixture-of-experts architecture, which activates only part of the model for each request. Alibaba said around 95 billion parameters are active at a time, reducing costs and response delays compared with activating the full model.

DeepSeek uses a similar sparse architecture at a smaller scale. Artificial Analysis lists V4-Flash at 284 billion total parameters, with 13 billion active during inference, while Moonshot AI’s Kimi K3 has 2.8 trillion total parameters and about 104 billion active.

Qwen3.8-Max can process text, images, and video and supports up to one million tokens of context. Alibaba also said the model completed a software engineering project over 16 days.

Its size places it close to Kimi K3, which Moonshot AI released in July. The two companies are also competing on price, with Qwen3.8-Max costing $2 per million input tokens and $6 per million output tokens, compared with $3 and $15, respectively, for Kimi K3.

Model size alone does not determine inference cost. Architecture, active parameter count, token consumption, and the number of calls required to complete a task also affect how much a model costs to run.

Qwen3.8-Max moved to the top position among Chinese text models on crowdsourced comparison platform Arena.AI following its release, although it remained behind several Anthropic models in the overall rankings. It also ranked second on Arena.AI’s leaderboard for models that analyse images and other visual material, behind an Anthropic Claude Fable 5 variant.

DeepSeek pushes down inference pricing

DeepSeek has taken a different approach with V4-Flash. Rather than matching the overall scale of Alibaba’s and Moonshot AI’s latest models, it has priced the model below several widely used AI systems.

V4-Flash costs $0.14 per million input tokens and $0.28 per million output tokens, according to Artificial Analysis. The research firm lists the model with a one-million-token context window and 284 billion total parameters, of which 13 billion are active during inference.

Artificial Analysis lists cache-hit pricing of $0.003 per million tokens for the Max Effort version of V4-Flash, 98% below its standard input rate. Cached input covers previously processed context that can be reused across subsequent requests.

DeepSeek’s lower token rates also carried through to Artificial Analysis’ benchmark testing. Reuters reported that the research firm estimated V4-Flash’s average cost at three cents per test, compared with 86 cents for Kimi K3, $1.86 for OpenAI’s GPT-5.6 Sol, and $3.15 for Anthropic’s Claude Fable 5.

The comparison accounts for the amount of input and output each model uses to complete the benchmark. A lower per-token rate does not necessarily result in a lower task cost if a model generates more output or requires additional interactions.

Artificial Analysis gave the Max Effort reasoning version of DeepSeek V4-Flash a score of 40 on its Intelligence Index. The research firm also recorded an output rate of about 118 tokens per second during testing.

Token prices tell only part of the cost story

Moonshot AI’s Kimi K3 provides another example of how advertised API prices can differ from the cost of completing longer workloads. Artificial Analysis lists the model at $3 per million input tokens and $15 per million output tokens, with cached input priced at $0.30 per million tokens.

On Artificial Analysis’ AA-Briefcase benchmark for agentic knowledge work, Kimi K3 averaged $10.57 per task. It generated around 120,000 output tokens and used an average of 83 turns per task.

Artificial Analysis said the cost reflected Kimi K3’s token pricing, output volume, and number of model interactions. Repeated model calls and larger outputs can therefore raise the total cost of completing a workload beyond what the headline API rate suggests.

Kimi K3 recorded the second-highest overall score on the AA-Briefcase evaluation at the time of testing, behind Claude Fable 5. It also scored 57 on Artificial Analysis’ broader Intelligence Index.

The comparison with DeepSeek shows why cost-per-task measurements add useful context to standard API pricing. Models with different architectures and usage patterns can consume substantially different amounts of compute and tokens while working through the same type of task.

Open weights add another deployment option

Cost is also being shaped by how Chinese developers distribute their models. Alibaba, DeepSeek, and Moonshot AI have continued to support open-weight releases alongside hosted API access, giving developers more options for how the models are deployed.

Artificial Analysis lists DeepSeek V4-Flash as an open-weight model under an MIT licence, with weights available through Hugging Face. Kimi K3 is also available as an open-weight model under Moonshot AI’s own licence.

Open weights allow developers to run models on their own infrastructure or through third-party providers instead of relying solely on a developer-hosted inference service. Deployment costs still depend on the hardware and infrastructure used, but access to the model is not tied to a single hosted API.

The approach differs from the main models offered by OpenAI, Anthropic, and Google, which generally keep their model weights closed.

Lian Jye Su, chief analyst at Omdia, said model selection for many business workloads does not depend solely on having access to the highest-performing system.

“Many business workflows do not need the industry’s very best model,” Su said. “They need models that are good enough, affordable, transparent and accessible, and open-weight models help meet that demand.”

(Photo by Solen Feyissa)

See also: Alibaba is designing AI chips around agents, and that changes what the race is actually about

Banner for AI & Big Data Expo by TechEx events.

Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and is co-located with other leading technology events including the Cyber Security & Cloud Expo. Click here for more information.

AI News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here.

Leave a Reply

Your email address will not be published. Required fields are marked *