What China's Open-Model Convergence Actually Changes: Value-Chain Redistribution, Not Demand Collapse

Chinese open models are converging on US frontier performance while being served on domestic Chinese inference infrastructure. This is value-chain redistribution rather than an AI demand collapse. The monopoly value of model APIs and leading-edge GPUs falls, while cheaper inference lifts token volume and can raise the value of memory, storage, networking, power and cloud distribution. This post maps how US AI policy and US-China tension reshape that reading, and whether Chinese model APIs can be adopted by Western enterprises.

Context This piece is a follow-up to Kimi K3 Resets the AI Price Curve. Where that piece verified the pricing and architecture of a single model, this one expands the lens to how the entire Chinese open-model ecosystem redistributes the semiconductor and Big Tech value chain, and how US policy and US-China tension reshape that reading. It pairs well with The Real Debate in Semiconductors, Are Semiconductors Cyclical, and What Is Fair Value?, and CXMT IPO And Memory Price Risk. Related hubs are the AI HBM Hub and the Exclusive Analysis Hub.

TL;DR

  • Chinese open models converging in performance looks more like value-chain redistribution than an AI demand collapse. The monopoly value of model APIs and leading-edge GPUs falls, while cheaper inference lifts usage and can raise the value of memory, storage, networking, power, and cloud distribution.
  • Relative preference splits this way. Within Korean memory, Samsung Electronics > SK Hynix; within US semiconductors, Micron and SanDisk > NVIDIA; within Big Tech, Meta and Amazon > Google and Microsoft > pure model vendors. [Inference: relative judgment]
  • Chinese open models have proven falling compute and memory cost per unit of intelligence. They have not proven a decline in total silicon spend. TSMC, if anything, raised its 2026 capex from $52 billion-$56 billion to $60 billion-$64 billion.
  • In HBM, CXMT trails the leading three by 1.5 to 2 product generations and by a 2- to 3-year commercialization gap. The 2028 threat is therefore a conditional option, not a present-day supply shock.
  • The technical diffusion of Chinese models and the revenue diffusion of Chinese API vendors are two different things. The most realistic path is rehosting Chinese models on AWS, Azure, or enterprise VPCs and selling them through Western security and contracting frameworks. [Analysis scope]

Key Framing
Chinese models can shake the pricing of US models, but they do not immediately collapse demand for AWS, Azure, US chips, or Korean memory. The direct casualty is the margin of closed-model APIs. What survives longest is usage-based infrastructure: the cloud distribution and security layer, along with memory, networking, and power.

1. Getting the Facts Straight First

The ecosystem direction is real, but not every new model has disclosed its training hardware. That distinction matters.

DeepSeek-V3 was trained on 2,048 NVIDIA H800 GPUs using 2.788 million GPU-hours. [Fact: DeepSeek V3 technical report] The H800 is export-restricted, but it is not a cheap consumer GPU. Kimi K3, by contrast, has disclosed 2.8 trillion parameters, a 1M-token context window, and 2.5x higher scaling efficiency versus K2, but the full technical report and training hardware are scheduled for release on July 27. [Fact: Kimi official announcement]

So the claim that “Kimi K3 was also trained on cheap NVIDIA GPUs” cannot yet be confirmed. [Blocked]

The Efficiency Gains Are Real

DeepSeek V4-Pro activates only 49 billion of its 1.6 trillion parameters. At the 1M-token range, per-token compute is 27% and KV cache is 10% of V3.2’s levels. [Fact: DeepSeek V4 model card] That is direct evidence that GPU and HBM use per token can fall sharply while performance holds.

But So Is the Other Side

Huawei’s CloudMatrix384 pools 384 Ascend 910C chips to serve DeepSeek-R1. By JPMAM’s tally, CloudMatrix uses 49TB of HBM and 599kW of power, versus 21TB and 145kW for the comparison system, GB300 NVL72. [Fact: CloudMatrix paper, JPMAM comparison]

It is not an apples-to-apples performance comparison, but the direction is clear: China compensates for a weaker single chip with more chips, more memory, and more power. A cheaper individual chip does not necessarily mean less total silicon, networking, or power. [Inference: structural reading]


2. The Core Equation: What to Actually Watch

The answer to this debate is not in benchmark scores. It is in two equations.

Total compute demand = total tokens × compute per token

Total memory demand = total tokens × memory use per token
                     + model, KV, and retrieval data storage

If open models cut token prices 70% and usage rises 5x, total demand increases even after the efficiency gain. Conversely, if usage merely doubles while per-token HBM falls 70%, HBM demand declines.

So the indicator to watch going forward compresses into one question: by how much does the token growth rate outpace the decline rate in memory per token?

The Verdict So Far

Total spend is determined by the following equation.

Total silicon spend = number of training runs × cost per run
                     + inference tokens × cost per token

On its 2Q26 call, TSMC raised its 2026 capex from $52 billion-$56 billion to $60 billion-$64 billion. Microsoft, too, raised inference throughput 40%, yet large-customer token use still rose 30% quarter-over-quarter, and it held roughly $190 billion in 2026 capex. [Fact: TSMC 2Q26 call, Microsoft FY26 Q3 call]

So far, usage growth is beating efficiency gains. [Inference: data synthesis]


3. Semiconductor Impact: It Diverges by Stock

Stock / GroupStock-Price ImpactKey Interpretation
Samsung ElectronicsPositiveBroadest beneficiary, capturing not just HBM but server DRAM, NAND/eSSD, and general-purpose memory
SK HynixMixed-to-positiveToken growth helps, but MoE and KV compression cut HBM per token, pressuring the scarcity multiple
MicronPositiveCombined DRAM, HBM, and NAND exposure across the US supply chain; broad-memory upside similar to Samsung
SanDiskPositiveLocal deployment, model weights, and growing RAG/cache use flow into eSSD and NAND demand
NVIDIANegative near-term, mixed long-termChina share and top-tier GPU monopoly value decline; offset by rising total tokens and H200 shipments
AMDRelatively positiveOpen models make it easier to migrate across heterogeneous hardware; ROCm and actual serving share are the key variables
Broadcom / Marvell / AristaPositive medium-termRising demand for Western custom ASICs, Ethernet, optical, and SerDes
TSMCNeutral-to-positiveTraining-chip demand from NVIDIA, AMD, and ASICs holds, but Chinese inference shifts to Ascend and SMIC

Why Samsung Has the Relative Edge Over Hynix

Among the same three memory makers, Samsung Electronics and SK Hynix diverge in direction. Cheap inference and open-model diffusion do not just lift HBM; they raise the total volume of server DRAM, NAND, and general-purpose memory. Samsung Electronics, with its broader exposure, captures more of that diffusion.

MoE and KV compression, conversely, reduce HBM per token. Hynix, with its concentrated HBM exposure, receives both the benefit of rising tokens and the burden of falling HBM per token at the same time. It is a structure where the scarcity multiple, not absolute demand, comes under pressure first. [Inference: exposure structure analysis]

Samsung Electronics does carry an offsetting risk, however: it is the first to be exposed to CXMT’s ramp-up of general-purpose DRAM output.

NVIDIA and H200

The US recently began shipping limited volumes of H200 to China, but the quantity is still small. It is positive for NVIDIA’s near-term revenue, but not enough to reverse the shift toward Huawei-centered domestic infrastructure. [Fact: July 2026 Reuters reporting] [Inference: impact judgment]


4. Big Tech Impact

CompanyVerdictReason
MetaMost positiveOpen-ecosystem strategy is validated, and AI returns are captured through advertising and recommendation rather than APIs
AmazonPositiveSells AWS inference, storage, and networking regardless of which model wins
GoogleMixed-to-positiveTPU and cloud benefit, but Gemini API price premium comes under pressure
MicrosoftMixedAzure usage rises, but OpenAI model rent and capex payback are pressured
ApplePositiveCheaper small and open models lower on-device and Private Cloud costs
OpenAI / AnthropicNegativeShrinking performance gap and API price premium pressure high valuations and fundraising

Contrary to a common misconception, Meta is not the biggest casualty. It makes money from advertising and recommendation rather than model sales, and it uses open models to cut costs, so it stands to relatively benefit. The burden falls instead on capex and depreciation. [Inference: business model analysis]

The Lesson from the 2025 DeepSeek Shock

During the 2025 DeepSeek shock, NVIDIA lost 17% and roughly $593 billion in market capitalization in a single day. The initial reaction was a broad sell-off across GPUs, power, and data-center infrastructure, but it partly rebounded soon after. [Fact: 2025 market reporting]

This time, too, near-term stock prices are likely to react to efficiency fears, while medium-term earnings react to rising usage. [Inference: historical pattern]


5. Verdicts on the Chinese Open-Model Thesis

Judging the claims coming from both the bull and bear sides, one by one, looks like this.

ClaimVerdict
Frontier performance requires top-tier GPUsWeakened, but not disproven
AI intelligence is a scarce resourceLargely collapsed at the level of raw model and token pricing
Capital scale is the moatThe pretraining moat weakens; the moat shifts to deployment, data, power, and distribution
Open source has zero marginal costOnly the weight price approaches zero; inference cost is ongoing
A $1 billion training run has overtaken $100 billion of investmentAn inaccurate claim comparing two different cost categories

The last item is especially often misused. Training cost and total infrastructure investment are not comparable categories.


6. How Far Has CXMT’s HBM Actually Gotten

This is the part of the China-threat discussion that gets exaggerated most often. Breaking it down gate by gate reveals the reality.

GateCurrent VerdictBasis for Judgment
DRAM cell processCommercialized, improving fastSelling DDR5 and LPDDR5X; Apple is also testing DRAM for China-market products
TSV stackingSmall-volume HBM2, early HBM3HBM2 production has been reported, but volume and yield are undisclosed
Packaging and base dieUnder constructionThe domestic ecosystem is forming, but mass-production yield, thermal, and reliability data are absent
Customer qualificationHBM unconfirmedThe Tencent contract and Apple testing are evidence of general DRAM capability, not HBM
Distance from the leaders1.5-2 generationsThe leading three are ramping HBM4; CXMT has not even verified HBM3 mass production

As of July 17, 2026, CXMT’s official published product lineup includes only DDR5, LPDDR5/5X, DDR4, and LPDDR4X, and no HBM. TrendForce likewise classifies CXMT’s HBM3 as still in early verification, and assesses that technical barriers and domestic-equipment requirements are delaying mass production. [Fact: CXMT official materials, TrendForce]

Capacity estimates also diverge. A plateau near 240,000 wafers per month conflicts with a year-end forecast of 350,000, and because the 350,000 figure comes from a private model rather than company guidance, it is hard to treat as a confirmed number. [Blocked]

How the Legacy-DRAM Hypothesis Needs to Be Revised

  • 2026-2027: CXMT’s HBM investment delays the easing of Chinese legacy-DRAM supply.
  • 2028: If HBM3/3E secures Chinese accelerator customer qualification and meaningful volume, it could begin displacing the older-generation HBM market within China first.
  • 2029 and beyond: Only once HBM4, base die, and packaging yield catch up does it directly pressure the global leading market and Hynix’s margins.

In other words, rather than “CXMT is making legacy supply tighter right now,” the more valid framing is “CXMT is not freeing up as much legacy supply as expected." [Inference: stage-by-stage judgment]


7. How US Policy Reshapes This Reading

The US is treating AI less as a commercial technology and more as core infrastructure for the allied bloc. That is why a purely technical analysis cannot supply the full answer.

The US AI Action Plan and Executive Order 14320 state explicitly that the US intends to export a full-stack American AI system, bundling hardware, cloud, networking, models, and applications, to allied nations, while reducing technological dependence on adversary nations. Even if a Chinese model is superior on performance and price, the US stack retains a policy advantage in allied-nation government procurement and critical industries. [Fact: White House AI Action Plan, EO 14320]

Yet This Is Not a Full Decoupling

In January 2026, the US BIS moved to review H200 and MI325X exports to China case by case, under approved-customer and security conditions. It is a compromise that keeps China partly inside the US chip ecosystem while retaining control. [Fact: BIS 2026-01-13]

Nor is the use of Chinese models by US private companies fully banned. The No DeepSeek on Government Devices Act is still only at the bill-introduction stage. Australia’s government, however, has ordered DeepSeek removed from government systems, and Italy’s privacy authority has restricted processing of user data. [Fact: official actions by country]

De facto bans from procurement and security departments are likely to take effect before legislation does. [Inference: policy sequencing]


8. Can Western Enterprises Actually Use Chinese Model APIs?

Diffusion potential differs completely by pathway. This table is the answer to the question.

Adoption MethodDiffusion PotentialJudgment
Direct calls to mainland China-hosted APIsLowData, jurisdiction, and procurement risk
Qwen’s US, EU, Japan, and Singapore APIsMediumData localization is possible, but Chinese-vendor risk remains
Chinese models rehosted by AWS or AzureMedium-high to highWestern cloud providers hold the contracting, security, and data control
Enterprise VPC or on-premises open weightsHighNo need to send data to Chinese servers
Government, defense, and critical infrastructureVery lowProcurement restrictions can extend even to model lineage

The most realistic pathway looks like this.

Chinese model developed
→ Rehosted on AWS, Azure, or enterprise VPC
→ Sold through Western security, contracting, and audit frameworks

In other words, the technical diffusion of Chinese models and the revenue diffusion of Chinese API vendors are separate matters. [Inference: pathway analysis]

The Limits of Direct APIs

DeepSeek’s Privacy Policy states that it collects, processes, and stores personal data, including input data, directly in China, and may use it to improve its service and models. [Fact: DeepSeek Privacy Policy] That is a disqualifying condition for companies handling source code, customer information, healthcare or financial data, or export-controlled technology.

Qwen is somewhat different. Alibaba Cloud Model Studio offers US, German, Japanese, and Singapore regions, can restrict data and inference to a specific region, and states explicitly that it does not use customer data for model training. [Fact: Alibaba Cloud region and security policy] Direct API adoption could therefore grow in non-regulated industries and in Southeast Asia, the Middle East, and Latin America.

Data localization and SOC 2, however, do not eliminate the risk of service disruption, sanctions, or procurement exposure arising from US-China tension. [Inference: residual risk]


9. Revising the Thesis: What Changes and What Stays

First, the case for a US hyperscaler collapse weakens. AWS and Azure directly host DeepSeek, providing data isolation, SLAs, and security assessments. US clouds can absorb the cost innovation of Chinese models and turn it into revenue. [Fact: AWS and Azure official announcements]

Second, pricing pressure hits closed-model APIs first. It burdens the token prices and margins of OpenAI, Anthropic, and Google, but inference volume and AWS/Azure usage can still rise. It is relatively favorable for Meta’s open-model strategy.

Third, the US-China split supports total infrastructure investment. Both blocs build out duplicate accelerators, memory, networking, and power grids. That lowers the likelihood that efficiency gains translate directly into a decline in global silicon spend.

Fourth, SK Hynix’s HBM premium can persist longer within the allied bloc. CXMT’s HBM is more likely to penetrate Chinese accelerators and Chinese data centers first. Adoption by Western CSPs requires policy and supply-chain certification on top of technical qualification. That said, export controls accelerate China’s push for self-sufficiency, which is a risk to Hynix’s China-market share from 2028 onward.

Fifth, Samsung Electronics is a relative hedge. Cheap inference and open-model diffusion can raise the total volume of server DRAM, NAND, and general-purpose memory, not just HBM. On the other side is the risk of being first exposed to CXMT’s ramp-up of general-purpose DRAM output.


10. The Order of Casualties and Beneficiaries

  1. Greatest casualty: the excess profit of closed-model APIs, and highly leveraged GPU-leasing operators
  2. Intermediate risk: heavily front-invested, customer-concentrated operators such as Oracle
  3. Meta: not the biggest casualty. It uses open models to cut costs, so relative benefit is possible. The burden falls on capex and depreciation
  4. NVIDIA: the multiple and product mix come under pressure before near-term EPS does. Displacement inside China is a risk, but the CUDA, networking, power-efficiency, and development-timeline moats remain
  5. Memory: the base case is not a decline in absolute HBM demand, but a narrowing HBM scarcity premium plus tier expansion into DDR/CXL/eSSD

11. Scenario Update

It makes sense to provisionally revise the probabilities used in Are Semiconductors Cyclical, and What Is Fair Value?

ScenarioPriorRevised
Excess demand persists30%30%
Asset re-concentration40%40%
Supply/efficiency normalization20%25%
System demand short-circuit10%5%

Chinese model performance lowers the probability that AI demand itself disappears (the short-circuit scenario, from 10% to 5%). In exchange, efficiency gains and Chinese-made hardware lower the scarcity premium of NVIDIA and HBM (the normalization scenario, from 20% to 25%).

The timeline scenarios also hold: usage beats efficiency through 2027 at 70%, mix and pricing normalize from 2028 onward at 25%, and total capex contracts at 5%. However, the driver of the 70% scenario shifts from a single global ecosystem to dual investment across the US bloc and the China bloc. [Inference: scenario recalibration]


12. What Would Change This Judgment

  • The Kimi K3 technical report (July 27): once training hardware is disclosed, the truth of the “frontier training on cheap GPUs” claim will be settled
  • Token growth rate versus the decline rate in memory per token: the gap between these two values is the real answer to this debate
  • CXMT’s HBM3E mass-production qualification and actual packaging volume: if confirmed, both Hynix’s 2028 EPS and its multiple need to come down together
  • Slowing HBM pricing and content growth: the first signal of a narrowing scarcity premium
  • Expanding rehosting of Chinese models by Western CSPs: evidence that cloud-distribution value is rising
  • Actual enforcement of US procurement and security regulation: the scope of the de facto ban that operates ahead of legislation

This is not the stage to conflate terminal risk with current-period earnings. Hynix’s 2028-and-beyond outlook can be adjusted once CXMT’s HBM3E mass-production qualification is confirmed. [Inference: sequencing of judgment]


Closing

Returning to the question, the answer is this.

It is a fact that Chinese open models are converging on the US frontier, and it is a fact that cost per unit of intelligence has fallen sharply. But that does not mean total silicon spend is declining. TSMC’s capex increase and Microsoft’s 30% rise in tokens are the answer so far.

US policy substantially reshapes this reading. The US-China split produces duplicate investment across both blocs, supporting total infrastructure demand, and within the allied bloc, it protects the position of the US stack and Korean memory by policy.

Western enterprise adoption of Chinese model APIs must be viewed by separating the model from the vendor. The technology spreads as open weight, but the party selling it is more likely to be AWS, Azure, and enterprise VPCs than the Chinese API vendors themselves.

So the conclusion is redistribution. What collapses is the excess profit of closed-model APIs. What remains is usage-based infrastructure.


This post synthesizes public papers (DeepSeek V3/V4, Huawei CloudMatrix384), official company announcements (Kimi, the TSMC 2Q26 call, the Microsoft FY26 Q3 call, CXMT, Alibaba Cloud, DeepSeek’s Privacy Policy), US government materials (the AI Action Plan, EO 14320, BIS), regulatory actions by various countries, and market research (TrendForce, JPMAM). Kimi K3’s training hardware remains unconfirmed until its technical report is released on July 27, and CXMT’s capacity estimates and the scenario probabilities are author estimates as of the time of writing, not company guidance. The stocks mentioned are examples used to illustrate value-chain structure and are not a recommendation to buy or sell any specific security. Investment decisions and responsibility for them rest with the individual investor.


Built with Hugo
Theme Stack designed by Jimmy