<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Cost Structure on Korea Invest Insights</title><link>https://koreainvestinsights.com/tags/cost-structure/</link><description>Recent content in Cost Structure on Korea Invest Insights</description><generator>Hugo -- gohugo.io</generator><language>en</language><copyright>koreainvestinsights.com · @korea_invest_insights</copyright><lastBuildDate>Fri, 17 Jul 2026 21:11:57 +0900</lastBuildDate><atom:link href="https://koreainvestinsights.com/tags/cost-structure/feed.xml" rel="self" type="application/rss+xml"/><item><title>Is China's AI API Pricing Sustainable? Verifying the Cost Structure Through Disclosures</title><link>https://koreainvestinsights.com/post/china-ai-api-pricing-sustainability-cost-structure-2026-07-18/</link><pubDate>Sat, 18 Jul 2026 11:00:00 +0900</pubDate><guid>https://koreainvestinsights.com/post/china-ai-api-pricing-sustainability-cost-structure-2026-07-18/</guid><description>
 &lt;blockquote&gt;
 &lt;p&gt;Context
This post is a follow-up to &lt;a class="link" href="https://koreainvestinsights.com/post/china-open-model-convergence-value-chain-redistribution-2026-07-17/" &gt;What China&amp;rsquo;s Open-Model Convergence Actually Changes&lt;/a&gt;. Where that piece covered the &lt;strong&gt;value-chain implications&lt;/strong&gt; of performance convergence, this one verifies through Hong Kong disclosures whether that low-cost API pricing is &lt;strong&gt;sustainable at the cost-structure level&lt;/strong&gt;. Best read alongside &lt;a class="link" href="https://koreainvestinsights.com/post/kimi-k3-linear-api-pricing-semiconductor-big-tech-impact-2026-07-17/" &gt;Kimi K3 Resets the AI Price Curve&lt;/a&gt; and &lt;a class="link" href="https://koreainvestinsights.com/post/semiconductor-bull-bear-four-clocks-capital-intensity-cycle-2026-07-17/" &gt;The Real Debate in Semiconductors&lt;/a&gt;. Related hubs are the &lt;a class="link" href="https://koreainvestinsights.com/page/korea-semiconductor-hbm-kospi-hub/" &gt;AI HBM Hub&lt;/a&gt; and the &lt;a class="link" href="https://koreainvestinsights.com/page/exclusive-analysis-hub/" &gt;Exclusive Analysis Hub&lt;/a&gt;.&lt;/p&gt;

 &lt;/blockquote&gt;
&lt;h2 id="tldr"&gt;TL;DR
&lt;/h2&gt;&lt;ul&gt;
&lt;li&gt;Current Chinese AI API pricing is &lt;strong&gt;a blend of structurally lower cost and strategic loss-leading competition&lt;/strong&gt;. It is unlikely that every ultra-low price holds, and equally unlikely that prices fully revert to US frontier levels.&lt;/li&gt;
&lt;li&gt;Hong Kong disclosures split the answer. Zhipu&amp;rsquo;s API gross margin improved to &lt;strong&gt;18.9%&lt;/strong&gt; (from 3.3% a year earlier), meaning it has &lt;strong&gt;started clearing marginal inference cost&lt;/strong&gt;. But its total gross profit covers only &lt;strong&gt;9.3%&lt;/strong&gt; of R&amp;amp;D, so &lt;strong&gt;fully loaded cost remains unrecovered&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;China&amp;rsquo;s cost edge is real, but it is not Huawei silicon or cheap electricity. The order is &lt;strong&gt;model architecture and reduced active parameters &amp;gt; batching, caching and quantization &amp;gt; cloud GPU utilization &amp;gt; low target margins &amp;gt; power and domestic NPUs&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;Even generously sized, the power effect amounts to only &lt;strong&gt;about 2%&lt;/strong&gt; of compute price. The fact that Qwen 3.5 Flash is priced identically in Beijing, Frankfurt and Virginia backs this up.&lt;/li&gt;
&lt;li&gt;Unlike EVs, AI models have &lt;strong&gt;low switching costs&lt;/strong&gt;. If winner-take-all emerges, it will be in the cloud and distribution layer, not the model layer. The most likely path is a &lt;strong&gt;low-price oligopoly (60%)&lt;/strong&gt;. [Analysis scope]&lt;/li&gt;
&lt;/ul&gt;
&lt;hr&gt;
&lt;div class="thesis-callout"&gt;
 &lt;div class="thesis-callout__label"&gt;Key Framing&lt;/div&gt;
 &lt;div class="thesis-callout__body"&gt;
 Chinese API prices are low not because they are already self-sustaining at a normal price level, but because improving inference efficiency and massive capital raising are propping them up together. That is a burden on the API margins of Western model vendors, but because the capital raised flows back into training and computing infrastructure, it does not mean total semiconductor demand is falling.
 &lt;/div&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;h2 id="1-how-cheap-is-it"&gt;1. How Cheap Is It
&lt;/h2&gt;&lt;p&gt;Prices per 1M tokens as of July 2026. [Fact: official pricing pages]&lt;/p&gt;
&lt;table&gt;
 &lt;thead&gt;
 &lt;tr&gt;
 &lt;th&gt;Model&lt;/th&gt;
 &lt;th style="text-align: right"&gt;Input&lt;/th&gt;
 &lt;th style="text-align: right"&gt;Output&lt;/th&gt;
 &lt;/tr&gt;
 &lt;/thead&gt;
 &lt;tbody&gt;
 &lt;tr&gt;
 &lt;td&gt;DeepSeek V4 Flash&lt;/td&gt;
 &lt;td style="text-align: right"&gt;$0.14&lt;/td&gt;
 &lt;td style="text-align: right"&gt;$0.28&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;DeepSeek V4 Pro&lt;/td&gt;
 &lt;td style="text-align: right"&gt;$0.435&lt;/td&gt;
 &lt;td style="text-align: right"&gt;$0.87&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Gemini 3.5 Flash&lt;/td&gt;
 &lt;td style="text-align: right"&gt;$1.50&lt;/td&gt;
 &lt;td style="text-align: right"&gt;$9&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;GPT-5.6 Sol&lt;/td&gt;
 &lt;td style="text-align: right"&gt;$5&lt;/td&gt;
 &lt;td style="text-align: right"&gt;$30&lt;/td&gt;
 &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;DeepSeek V4-Pro is about &lt;strong&gt;11 times&lt;/strong&gt; cheaper than GPT-5.6 Sol on input and about &lt;strong&gt;34 times&lt;/strong&gt; cheaper on output. That does not, of course, mean performance, latency, SLA, security and tooling are equivalent.&lt;/p&gt;
&lt;h3 id="the-price-gap-is-large-even-within-china"&gt;The Price Gap Is Large Even Within China
&lt;/h3&gt;&lt;ul&gt;
&lt;li&gt;Qwen 3.7 Max: RMB 12 input, RMB 36 output, currently 50% off&lt;/li&gt;
&lt;li&gt;Doubao Seed 2.1 Pro: RMB 6 input, RMB 30 output&lt;/li&gt;
&lt;li&gt;DeepSeek V4-Pro: far more aggressive than even its Chinese rivals&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;So &lt;strong&gt;DeepSeek&amp;rsquo;s pricing should not be read as the normal price for China as a whole&lt;/strong&gt;. [Inference: price dispersion reading]&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id="2-the-sustainable-part-marginal-cost"&gt;2. The Sustainable Part: Marginal Cost
&lt;/h2&gt;&lt;p&gt;DeepSeek V4-Pro is an MoE architecture that &lt;strong&gt;activates only 49 billion&lt;/strong&gt; of its 1.6 trillion total parameters. Combining high utilization with batch processing and caching can sharply lower the actual compute cost per token. The cache-hit price for repeated input is just &lt;strong&gt;$0.003625&lt;/strong&gt; per 1M tokens. [Fact: DeepSeek official materials]&lt;/p&gt;
&lt;p&gt;So for the following traffic, even current pricing can be sustainable on a marginal-cost basis.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Batch processing&lt;/li&gt;
&lt;li&gt;Cache hits&lt;/li&gt;
&lt;li&gt;Traffic without a latency guarantee&lt;/li&gt;
&lt;li&gt;High-volume use that can sustain high utilization&lt;/li&gt;
&lt;/ul&gt;
&lt;hr&gt;
&lt;h2 id="3-the-hard-to-sustain-part-fully-loaded-cost"&gt;3. The Hard-to-Sustain Part: Fully Loaded Cost
&lt;/h2&gt;&lt;p&gt;Published pricing alone struggles to recover model training cost, R&amp;amp;D, failed experiments, idle capacity, and security, sales and enterprise SLA costs.&lt;/p&gt;
&lt;p&gt;Price normalization has already begun. [Fact: company actions and reporting]&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;DeepSeek cut V4 pricing 75%, then &lt;strong&gt;doubled peak-hour pricing&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;Alibaba also raised prices on some AI services by up to &lt;strong&gt;34%&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;Dedicated enterprise throughput and low-latency service are priced above the public API&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This is evidence that ultra-low prices do not hold across every time slot and customer segment.&lt;/p&gt;
&lt;table&gt;
 &lt;thead&gt;
 &lt;tr&gt;
 &lt;th&gt;Service&lt;/th&gt;
 &lt;th&gt;Sustainability&lt;/th&gt;
 &lt;/tr&gt;
 &lt;/thead&gt;
 &lt;tbody&gt;
 &lt;tr&gt;
 &lt;td&gt;Cache, batch and non-priority traffic&lt;/td&gt;
 &lt;td&gt;High&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Peak hours, low latency, high concurrency&lt;/td&gt;
 &lt;td&gt;Low&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Enterprise security and SLA&lt;/td&gt;
 &lt;td&gt;Settling into separate, higher pricing&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Full cost recovery for a standalone API business&lt;/td&gt;
 &lt;td&gt;Difficult at current prices&lt;/td&gt;
 &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;hr&gt;
&lt;h2 id="4-verifying-through-disclosures-zhipu-and-minimax"&gt;4. Verifying Through Disclosures: Zhipu and MiniMax
&lt;/h2&gt;&lt;p&gt;This is where the real analysis begins. Rather than speculate, we can check this against the &lt;strong&gt;disclosures of Hong Kong-listed companies&lt;/strong&gt;.&lt;/p&gt;
&lt;table&gt;
 &lt;thead&gt;
 &lt;tr&gt;
 &lt;th&gt;2025&lt;/th&gt;
 &lt;th style="text-align: right"&gt;Zhipu (02513)&lt;/th&gt;
 &lt;th style="text-align: right"&gt;MiniMax (00100)&lt;/th&gt;
 &lt;/tr&gt;
 &lt;/thead&gt;
 &lt;tbody&gt;
 &lt;tr&gt;
 &lt;td&gt;Revenue&lt;/td&gt;
 &lt;td style="text-align: right"&gt;RMB 724m&lt;/td&gt;
 &lt;td style="text-align: right"&gt;$79.04m&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;API/platform revenue share&lt;/td&gt;
 &lt;td style="text-align: right"&gt;26.3%&lt;/td&gt;
 &lt;td style="text-align: right"&gt;32.8%&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;API or company-wide gross margin&lt;/td&gt;
 &lt;td style="text-align: right"&gt;API 18.9%&lt;/td&gt;
 &lt;td style="text-align: right"&gt;Company 25.4%&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Prior-year gross margin&lt;/td&gt;
 &lt;td style="text-align: right"&gt;API 3.3%&lt;/td&gt;
 &lt;td style="text-align: right"&gt;Company 12.2%&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;R&amp;amp;D / revenue&lt;/td&gt;
 &lt;td style="text-align: right"&gt;439%&lt;/td&gt;
 &lt;td style="text-align: right"&gt;320%&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Adjusted net loss / revenue&lt;/td&gt;
 &lt;td style="text-align: right"&gt;439%&lt;/td&gt;
 &lt;td style="text-align: right"&gt;317%&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Gross profit / R&amp;amp;D&lt;/td&gt;
 &lt;td style="text-align: right"&gt;&lt;strong&gt;9.3%&lt;/strong&gt;&lt;/td&gt;
 &lt;td style="text-align: right"&gt;&lt;strong&gt;7.9%&lt;/strong&gt;&lt;/td&gt;
 &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;[Fact: Zhipu 2025 results, MiniMax 2025 annual report]&lt;/p&gt;
&lt;h3 id="zhipu-marginal-cost-recovery-is-confirmed"&gt;Zhipu: Marginal-Cost Recovery Is Confirmed
&lt;/h3&gt;&lt;p&gt;API and open-platform revenue rose &lt;strong&gt;293%, from RMB 48.48m to RMB 190.4m&lt;/strong&gt;, and gross margin climbed from 3.3% to 18.9%. The company attributed this to inference efficiency, economies of scale and price increases.&lt;/p&gt;
&lt;p&gt;There is a more important fact. As of March 2026, the company had raised API prices &lt;strong&gt;83% versus the end of the prior year, and demand still exceeded supply&lt;/strong&gt;. GLM-5&amp;rsquo;s current official price is RMB 4 input and RMB 18 output per 1M tokens. The top-tier API appears to have moved past a structure that dumps below direct inference cost. [Inference: implication of the price hike with excess demand]&lt;/p&gt;
&lt;h3 id="minimax-plausible-but-unconfirmed"&gt;MiniMax: Plausible but Unconfirmed
&lt;/h3&gt;&lt;p&gt;The company has not disclosed an API-only gross margin. Company-wide gross margin improved from 12.2% to &lt;strong&gt;25.4%&lt;/strong&gt;, which it attributed to model and system efficiency and optimized infrastructure allocation.&lt;/p&gt;
&lt;p&gt;M2.5 is currently priced at $0.30 input and $1.20 output per 1M tokens, with a fast tier at $0.60/$2.40. Because M2.5 launched in 2026 while the disclosed financials run only through 2025, however, &lt;strong&gt;the standalone API cost ratio at current pricing has not yet been verified&lt;/strong&gt;. [Blocked]&lt;/p&gt;
&lt;h3 id="fully-loaded-cost-remains-unrecovered-at-both-companies"&gt;Fully Loaded Cost Remains Unrecovered at Both Companies
&lt;/h3&gt;&lt;p&gt;Zhipu&amp;rsquo;s total gross profit of RMB 297m covers only &lt;strong&gt;9.3%&lt;/strong&gt; of its RMB 3.18bn in R&amp;amp;D. MiniMax fared similarly: gross profit of $20.08m covered only &lt;strong&gt;7.9%&lt;/strong&gt; of its $253m in R&amp;amp;D, and operating cash outflow was $280m.&lt;/p&gt;
&lt;p&gt;Services MiniMax purchased from the Alibaba Group in 2025 totaled &lt;strong&gt;$75.88m, or 96% of revenue&lt;/strong&gt;. That figure blends training and serving cloud spend, so it cannot be read directly as API cost, but it is evidence of heavy reliance on external computing infrastructure. Alibaba holds a 17.06% stake in MiniMax.&lt;/p&gt;
&lt;p&gt;Importantly, the disclosure states that related-party transactions were conducted &lt;strong&gt;on the same public pricing and terms offered to general customers&lt;/strong&gt;. Capacity assurance and integration convenience from the strategic relationship are plausible, but &lt;strong&gt;there is no basis to conclude that a hidden price subsidy existed&lt;/strong&gt;. [Inference: reading of the disclosure language]&lt;/p&gt;
&lt;p&gt;Even after listing, Zhipu raised a net HK$31.375bn, and MiniMax raised a combined HK$15.957bn through new shares and convertible bonds. Most of the proceeds go into R&amp;amp;D and computing infrastructure. The fundraising itself is not evidence of an API loss, but &lt;strong&gt;frontier-model competition is not, at this stage, sustained by internal cash alone&lt;/strong&gt;. [Fact: disclosure]&lt;/p&gt;
&lt;h3 id="verdict"&gt;Verdict
&lt;/h3&gt;&lt;ol&gt;
&lt;li&gt;Marginal-cost sustainability of paid, top-tier APIs: &lt;strong&gt;confirmed&lt;/strong&gt; for Zhipu, plausible but unconfirmed for MiniMax&lt;/li&gt;
&lt;li&gt;Sustainability of fully loaded cost including training: &lt;strong&gt;unmet&lt;/strong&gt; at both companies&lt;/li&gt;
&lt;li&gt;Sustainability of the price war: &lt;strong&gt;high&lt;/strong&gt;. Capital-market fundraising can keep loss-making prices in place for a long time&lt;/li&gt;
&lt;li&gt;Market structure: rather than winner-take-all, more likely &lt;strong&gt;segment-by-segment oligopoly&lt;/strong&gt;, for example Zhipu in Chinese enterprise and on-premises deployment, MiniMax in overseas, consumer and multimodal&lt;/li&gt;
&lt;/ol&gt;
&lt;hr&gt;
&lt;h2 id="5-decomposing-the-cost-advantage-what-is-the-real-reason"&gt;5. Decomposing the Cost Advantage: What Is the Real Reason
&lt;/h2&gt;&lt;p&gt;China&amp;rsquo;s cost advantage in API serving is real. But the core driver is not cheap electricity or Huawei silicon alone, it is the following combination.&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;Small active-parameter count and quantization
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;+ High batch-processing rate and cache utilization
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;+ GPU utilization at large-scale clouds like Alibaba
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;+ Low target margins
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;+ Cheap power in western China
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;table&gt;
 &lt;thead&gt;
 &lt;tr&gt;
 &lt;th&gt;Factor&lt;/th&gt;
 &lt;th style="text-align: right"&gt;Cost advantage&lt;/th&gt;
 &lt;th&gt;Judgment&lt;/th&gt;
 &lt;/tr&gt;
 &lt;/thead&gt;
 &lt;tbody&gt;
 &lt;tr&gt;
 &lt;td&gt;Model architecture, quantization, caching&lt;/td&gt;
 &lt;td style="text-align: right"&gt;Very large&lt;/td&gt;
 &lt;td&gt;Core of China&amp;rsquo;s API price competitiveness&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Cloud scale and utilization&lt;/td&gt;
 &lt;td style="text-align: right"&gt;Large&lt;/td&gt;
 &lt;td&gt;Alibaba&amp;rsquo;s most defensible moat&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Low target margins&lt;/td&gt;
 &lt;td style="text-align: right"&gt;Large&lt;/td&gt;
 &lt;td&gt;Why prices are cheap, but a weakness for long-run profitability&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Huawei Ascend&lt;/td&gt;
 &lt;td style="text-align: right"&gt;Moderate, conditional&lt;/td&gt;
 &lt;td&gt;Advantageous on optimized workloads, weak on general-purpose use&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Western China power&lt;/td&gt;
 &lt;td style="text-align: right"&gt;Moderate&lt;/td&gt;
 &lt;td&gt;Effective for batch inference, limited for global real-time APIs&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Government subsidies and policy finance&lt;/td&gt;
 &lt;td style="text-align: right"&gt;Present&lt;/td&gt;
 &lt;td&gt;Favorable for industrial buildout, but per-service amounts are opaque&lt;/td&gt;
 &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;hr&gt;
&lt;h2 id="6-is-power-really-the-deciding-factor"&gt;6. Is Power Really the Deciding Factor
&lt;/h2&gt;&lt;p&gt;It is true that Chinese power is cheap. The landed power price at the Ningxia Zhongwei data center is &lt;strong&gt;RMB 0.36/kWh&lt;/strong&gt; as of 2026, about 45% of the level in eastern China. The 2025 US industrial average is &lt;strong&gt;8.62 cents/kWh&lt;/strong&gt;, which works out to about RMB 0.62/kWh at an assumed exchange rate of 7.2 RMB/USD. Zhongwei is about &lt;strong&gt;42% cheaper&lt;/strong&gt;. [Fact: Chinese local government data, US EIA]&lt;/p&gt;
&lt;p&gt;But power alone does not make API prices 5 or 10 times cheaper. &lt;strong&gt;Generously sizing&lt;/strong&gt; the power effect, assuming a 1kW IT load per card and a PUE of 1.1, gives this.&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;Zhongwei: 1kW × 1.1 × RMB 0.36 = RMB 0.40/hour
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;US average: 1kW × 1.1 × RMB 0.62 = RMB 0.68/hour
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;Difference: about RMB 0.29/hour
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Alibaba&amp;rsquo;s L20 on-demand price is RMB 14.4/hour, so even under this upper-bound assumption, the power difference is only &lt;strong&gt;about 2%&lt;/strong&gt; of compute price. [Inference: own calculation]&lt;/p&gt;
&lt;p&gt;Power matters, but it is not the main driver of the API price gap.&lt;/p&gt;
&lt;h3 id="the-decisive-counter-evidence"&gt;The Decisive Counter-Evidence
&lt;/h3&gt;&lt;p&gt;The western-power advantage suits batch inference and training well, but low-latency APIs must be deployed in eastern or overseas regions close to users. And Alibaba&amp;rsquo;s Qwen 3.5 Flash is priced &lt;strong&gt;identically at $0.029 input and $0.287 output per 1M tokens in Beijing as well as in Frankfurt and Virginia&lt;/strong&gt;. [Fact: Alibaba Model Studio pricing]&lt;/p&gt;
&lt;p&gt;That is the most direct evidence that low-cost APIs are not simply a result of Chinese power. [Inference: implication of identical regional pricing]&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id="7-alibabas-real-moat-is-utilization"&gt;7. Alibaba&amp;rsquo;s Real Moat Is Utilization
&lt;/h2&gt;&lt;p&gt;Alibaba Cloud&amp;rsquo;s real strength lies less in power than in its &lt;strong&gt;ability to keep GPUs from sitting idle&lt;/strong&gt;. [Fact: Alibaba official documentation]&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Batch inference: &lt;strong&gt;50% off&lt;/strong&gt; list price&lt;/li&gt;
&lt;li&gt;L20 spot instances: typically around RMB 2.88, versus RMB 14.4 on demand, a saving of roughly &lt;strong&gt;80%&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;Idle GPUs: $0.000007 per CU idle rate versus $0.000018 per CU active rate&lt;/li&gt;
&lt;li&gt;Mixing on-demand, spot and autoscaling to handle peak traffic&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Spot capacity carries reclaim risk, however, so a real-time API with a strict SLA cannot run entirely on spot. The structure keeps base demand on on-demand or dedicated resources and routes only elastic demand to spot. [Inference: operational constraint]&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id="8-is-huawei-ascend-unconditionally-cheap"&gt;8. Is Huawei Ascend Unconditionally Cheap
&lt;/h2&gt;&lt;p&gt;Strictly speaking, Ascend is not a GPU but an &lt;strong&gt;NPU&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;Huawei makes the following claims for CloudMatrix 384. [Fact: Huawei official materials]&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Single-card-equivalent inference throughput of 2,300 tokens/s&lt;/li&gt;
&lt;li&gt;About 4x versus a non-supernode configuration&lt;/li&gt;
&lt;li&gt;MFU improved by more than 50%&lt;/li&gt;
&lt;li&gt;3-4x average per-card inference performance versus H20&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;It is a system-design approach that pools resources and uses MoE expert parallelism to compensate for the weaknesses of lower-spec individual chips.&lt;/p&gt;
&lt;h3 id="but-it-should-not-be-read-as-low-cost-right-away"&gt;But It Should Not Be Read as Low Cost Right Away
&lt;/h3&gt;&lt;ul&gt;
&lt;li&gt;Huawei has &lt;strong&gt;not published an externally verified TCO on a like-for-like basis of model, latency and power&lt;/strong&gt;. [Blocked]&lt;/li&gt;
&lt;li&gt;Because CloudMatrix raises throughput by binding together many NPUs and network fabric, &lt;strong&gt;per-card throughput and total cluster cost are different things&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;A 2026 field study of Ascend 910 needed 12 source patches, disabling of some high-throughput features, and repeated fault-handling measures to achieve stable inference. Even if the hardware is cheap, &lt;strong&gt;engineering and operating costs can rise&lt;/strong&gt;. [Fact: Ascend field study]&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="conditional-verdict"&gt;Conditional Verdict
&lt;/h3&gt;&lt;table&gt;
 &lt;thead&gt;
 &lt;tr&gt;
 &lt;th&gt;Condition&lt;/th&gt;
 &lt;th&gt;Verdict&lt;/th&gt;
 &lt;/tr&gt;
 &lt;/thead&gt;
 &lt;tbody&gt;
 &lt;tr&gt;
 &lt;td&gt;Models deeply optimized for Ascend, such as Qwen, DeepSeek and GLM&lt;/td&gt;
 &lt;td&gt;Cost competitiveness possible&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;CUDA-based models ported as-is&lt;/td&gt;
 &lt;td&gt;Cost advantage uncertain&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Large-scale batch inference within China&lt;/td&gt;
 &lt;td&gt;Favorable&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Global enterprise real-time service&lt;/td&gt;
 &lt;td&gt;Unconfirmed once software and SLA costs are counted&lt;/td&gt;
 &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;hr&gt;
&lt;h2 id="9-how-is-this-different-from-evs"&gt;9. How Is This Different From EVs
&lt;/h2&gt;&lt;p&gt;The similarities are a strategic industry, large-scale investment, large-enterprise subsidies, and capturing share through price cuts.&lt;/p&gt;
&lt;p&gt;The difference is that &lt;strong&gt;switching costs for AI models are far lower&lt;/strong&gt;. API specifications are similar, and customers can route across multiple models simultaneously or self-host open-source models. If one vendor raises prices, competing models and self-hosting create a price ceiling.&lt;/p&gt;
&lt;p&gt;Platforms that combine cloud, data, security, payments, advertising and productivity tools, by contrast, have high switching costs. &lt;strong&gt;If winner-take-all appears, it will be in the distribution and cloud layer rather than the model layer&lt;/strong&gt;. [Inference: switching-cost structure]&lt;/p&gt;
&lt;p&gt;The most likely structure looks like this.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Alibaba: cloud, e-commerce, enterprise customers&lt;/li&gt;
&lt;li&gt;ByteDance: consumer traffic, advertising, content&lt;/li&gt;
&lt;li&gt;Tencent: WeChat, enterprise services&lt;/li&gt;
&lt;li&gt;Huawei and Baidu: domestic infrastructure plus public-sector and enterprise customers&lt;/li&gt;
&lt;li&gt;One or two independent model companies such as DeepSeek: setting the technical reference price&lt;/li&gt;
&lt;/ul&gt;
&lt;hr&gt;
&lt;h2 id="10-the-expected-path"&gt;10. The Expected Path
&lt;/h2&gt;&lt;table&gt;
 &lt;thead&gt;
 &lt;tr&gt;
 &lt;th&gt;Scenario&lt;/th&gt;
 &lt;th style="text-align: right"&gt;Assessed probability&lt;/th&gt;
 &lt;th&gt;Outcome&lt;/th&gt;
 &lt;/tr&gt;
 &lt;/thead&gt;
 &lt;tbody&gt;
 &lt;tr&gt;
 &lt;td&gt;Low-price oligopoly&lt;/td&gt;
 &lt;td style="text-align: right"&gt;&lt;strong&gt;60%&lt;/strong&gt;&lt;/td&gt;
 &lt;td&gt;Consolidates to 3-5 players, base API stays cheap, peak and SLA pricing rises&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Full commoditization&lt;/td&gt;
 &lt;td style="text-align: right"&gt;25%&lt;/td&gt;
 &lt;td&gt;Open source, MoE and in-house chips drive further price declines&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Large-scale normalization&lt;/td&gt;
 &lt;td style="text-align: right"&gt;15%&lt;/td&gt;
 &lt;td&gt;Chip, power and funding burdens push prices up 2-5x&lt;/td&gt;
 &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;[Inference: scenario estimate]&lt;/p&gt;
&lt;p&gt;Even a 5x price increase would put DeepSeek V4-Pro&amp;rsquo;s output price at about $4.35, still below US frontier models. &lt;strong&gt;Current pricing may not be the floor, but the direction toward lower prices is hard to reverse.&lt;/strong&gt;&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id="11-semiconductor-implications"&gt;11. Semiconductor Implications
&lt;/h2&gt;&lt;p&gt;Cheap APIs raise inference usage, which favors demand for server DRAM, NAND, networking and power. At the same time, they lower compute per token and HBM intensity, which pressures the per-unit economics of premium GPUs and HBM.&lt;/p&gt;
&lt;table&gt;
 &lt;thead&gt;
 &lt;tr&gt;
 &lt;th&gt;Target&lt;/th&gt;
 &lt;th&gt;Judgment&lt;/th&gt;
 &lt;/tr&gt;
 &lt;/thead&gt;
 &lt;tbody&gt;
 &lt;tr&gt;
 &lt;td&gt;Samsung Electronics&lt;/td&gt;
 &lt;td&gt;Relatively cushioned by its large general-purpose DRAM and NAND exposure&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;SK Hynix&lt;/td&gt;
 &lt;td&gt;A long-run warning sign for the HBM scarcity premium&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;NVIDIA&lt;/td&gt;
 &lt;td&gt;Pressure on premium GPU unit pricing, offset by rising total inference volume&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Alibaba, Tencent&lt;/td&gt;
 &lt;td&gt;Benefit from cloud and distribution ecosystem expansion rather than API margin&lt;/td&gt;
 &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;The key variable is not the rate of API price decline but the &lt;strong&gt;growth rate of total token usage&lt;/strong&gt;. If usage grows faster than prices fall, the Jevons effect wins; if not, the multiple compresses first for premium GPUs and HBM. [Inference: directional judgment]&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id="12-the-decisive-test-is-the-first-half-2026-disclosures"&gt;12. The Decisive Test Is the First-Half 2026 Disclosures
&lt;/h2&gt;&lt;ul&gt;
&lt;li&gt;Does &lt;strong&gt;Zhipu&amp;rsquo;s API gross margin&lt;/strong&gt; climb above 25-30% following the 83% price increase&lt;/li&gt;
&lt;li&gt;Does &lt;strong&gt;MiniMax disclose an API cost ratio&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;Does &lt;strong&gt;revenue growth keep outpacing computing-cost growth&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;Does the gap between peak/SLA pricing and public API pricing become entrenched&lt;/li&gt;
&lt;li&gt;Does Huawei publish a like-for-like, externally verified TCO&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;These five questions separate &amp;ldquo;sustainable cost&amp;rdquo; from &amp;ldquo;prices propped up by capital.&amp;rdquo;&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id="closing"&gt;Closing
&lt;/h2&gt;&lt;p&gt;Chinese API prices are low &lt;strong&gt;not because they are already a self-sustaining, normal price level&lt;/strong&gt;. It is because improving inference efficiency and massive capital raising are propping them up together.&lt;/p&gt;
&lt;p&gt;China has a structural cost advantage in domestic, batch-oriented, optimized model serving. But the extremely low API prices observed today cannot be explained by Huawei silicon and cheap power alone. The largest cost advantages, in order, are model size and reduced active parameters, batching/caching/quantization, cloud GPU utilization, and low target margins, with power and domestic NPUs coming after those.&lt;/p&gt;
&lt;p&gt;From an investment standpoint, China&amp;rsquo;s cheap APIs do not immediately mean that &amp;ldquo;Chinese AI has secured an overwhelming cost moat.&amp;rdquo; The more accurate reading is that &lt;strong&gt;as inference service rapidly commoditizes, model companies&amp;rsquo; margins fall, and value is increasingly likely to shift toward cloud, power and semiconductor infrastructure&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;That is a burden on the API margins of Western model vendors, but because the capital raised flows back into training and computing infrastructure, it &lt;strong&gt;does not mean total semiconductor demand is declining&lt;/strong&gt;.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;&lt;em&gt;This post synthesizes Hong Kong listing disclosures (Zhipu 02513&amp;rsquo;s 2025 results and share placement, MiniMax 00100&amp;rsquo;s 2025 annual report and fundraising), official company pricing pages (DeepSeek, OpenAI, Google, Alibaba Model Studio, Volcano Engine, Zhipu BigModel, MiniMax Platform), Huawei official materials and CloudMatrix-related research, Chinese local government power data, and US EIA statistics. MiniMax&amp;rsquo;s standalone API cost ratio and Huawei&amp;rsquo;s like-for-like, externally verified TCO remain undisclosed, and the scenario probabilities and power calculations are author estimates based on assumptions at the time of writing. The stocks mentioned are examples used to illustrate cost structure and are not a recommendation to buy or sell any specific security. Prices and disclosed figures are as of the announcement date and may change afterward. Investment decisions and responsibility for them rest with the individual investor.&lt;/em&gt;&lt;/p&gt;
&lt;hr&gt;
&lt;h3 id="related-posts"&gt;Related Posts
&lt;/h3&gt;&lt;ul&gt;
&lt;li&gt;&lt;a class="link" href="https://koreainvestinsights.com/post/china-open-model-convergence-value-chain-redistribution-2026-07-17/" &gt;What China&amp;rsquo;s Open-Model Convergence Actually Changes: Value-Chain Redistribution, Not Demand Collapse&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a class="link" href="https://koreainvestinsights.com/post/kimi-k3-linear-api-pricing-semiconductor-big-tech-impact-2026-07-17/" &gt;Kimi K3 Resets the AI Price Curve: From Kimi Linear to HBM and Big Tech Strategy&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a class="link" href="https://koreainvestinsights.com/post/semiconductor-bull-bear-four-clocks-capital-intensity-cycle-2026-07-17/" &gt;The Real Debate in Semiconductors: Four Physical Clocks and One Stock-Price Clock&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a class="link" href="https://koreainvestinsights.com/post/memory-fair-value-fcfe-terminal-samsung-hynix-micron-2026-07-17/" &gt;Are Semiconductors Cyclical, and What Is Fair Value? Pricing Samsung, SK Hynix and Micron with FCFE and Normalized Earnings&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a class="link" href="https://koreainvestinsights.com/post/ai-token-value-memory-value-added-2026-07-09/" &gt;AI Token Value Today and Tomorrow: Value Added for Memory Companies&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;</description></item></channel></rss>