
Chinese startup DeepSeek officially released its V4-Flash model on Friday, positioning itself as the world's most cost-effective AI solution. According to research firm Artificial Analysis, the model charges $0.14 per million input tokens and $0.28 per million output tokens, making it more than 100 times cheaper to run than Anthropic's Claude Fable 5. The startup's R1 model had become a global sensation in early 2025, triggering a selloff in global technology stocks and raising questions about large AI spending by U.S. companies. As per multiple reports, DeepSeek is also preparing for a potential IPO as it seeks to regain momentum in the competitive Chinese AI market.
San Francisco-based Artificial Analysis estimated V4-Flash's average cost at 3 cents per test, significantly outperforming competitors. The comparison shows V4-Flash at 3 cents per test versus 86 cents for Kimi K3 from Chinese rival Moonshot AI, $1.86 for OpenAI's GPT-5.6 Sol, and $3.15 for Claude Fable 5. The research firm noted that this cost comparison provides a more realistic measure of value than pricing alone, accounting for the amount of data required to complete tasks. A model with low headline price can still prove expensive if it requires significantly more steps to produce an answer, making the cost-per-test metric more meaningful for practical AI deployment. The gap is significant because inference — the process through which a trained model answers questions or performs tasks — is becoming one of the largest operating expenses for companies deploying AI at scale.
Artificial Analysis reported that DeepSeek's V4-Flash scored 50 out of 100 on its Intelligence Index, which combines results from nine benchmarks spanning coding, reasoning and workplace-style assignments. This score matches Google's Gemini 3.6 Flash and trails Meta's Muse Spark 1.1 and GLM-5.2 from Z.AI by one point. Moonshot's Kimi K3 scored 57 points, while Anthropic's Claude Opus 5, Fable 5, and OpenAI GPT-5.6 scored nine or more points higher. The model's pricing strength comes from its sparse mixture-of-experts architecture, which activates only part of its overall network for each token rather than using every parameter for every request. DeepSeek has also implemented compressed memory techniques and specialized attention mechanisms to lower the cost of processing long sequences of text. The model's cache-hit price is substantially lower at $0.0028 per million input tokens, making it particularly attractive for applications that repeatedly process similar documents, instructions or conversation histories.
DeepSeek faces intense competition from domestic rivals including Moonshot, MiniMax, and Z.AI as well as tech giants like ByteDance and Alibaba. The startup is also preparing a more powerful version called V4-Pro, though no official release date has been announced. Separately, Alibaba has expanded its Qwen range with higher-capacity models, while Moonshot has promoted Kimi K3 as a cost-efficient alternative for developers requiring large context windows and agentic features. All companies are vying with U.S. tech firms for global adoption, targeting businesses seeking cheaper ways to deploy AI at scale. The widening price divide is encouraging companies to use multiple models rather than relying on one provider, with AI routing platforms directing basic tasks to cheaper systems while reserving costly frontier models for difficult reasoning and research tasks. Lower prices could expand AI adoption among smaller businesses that have been unable to justify high inference bills, potentially benefiting call centers, online retailers, financial technology companies and software developers.