<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>AI Agents on Korea Invest Insights</title><link>https://koreainvestinsights.com/tags/ai-agents/</link><description>Recent content in AI Agents on Korea Invest Insights</description><generator>Hugo -- gohugo.io</generator><language>en</language><copyright>koreainvestinsights.com · @korea_invest_insights</copyright><lastBuildDate>Thu, 27 Aug 2026 17:45:13 +0900</lastBuildDate><atom:link href="https://koreainvestinsights.com/tags/ai-agents/feed.xml" rel="self" type="application/rss+xml"/><item><title>NVIDIA's Earnings Changed the Question: Watch Memory Bottlenecks and Supplier Financing, Not Just Peak AI Capex</title><link>https://koreainvestinsights.com/post/nvidia-q2-fy2027-memory-bottleneck-supplier-financing-2026-08-27/</link><pubDate>Thu, 27 Aug 2026 17:30:00 +0900</pubDate><guid>https://koreainvestinsights.com/post/nvidia-q2-fy2027-memory-bottleneck-supplier-financing-2026-08-27/</guid><description>&lt;p&gt;NVIDIA reported fiscal Q2 2027 revenue of $96.2 billion, up 106% from a year earlier. Data Center revenue alone reached $89.0 billion.&lt;/p&gt;
&lt;p&gt;The stock initially weakened after hours. Investors priced the coming gross-margin decline before rewarding the revenue beat. The direction changed only after management provided its longer-term growth outlook on the earnings call.&lt;/p&gt;
&lt;p&gt;The reversal came on the earnings call, when management said fiscal 2028 revenue could grow by approximately 70%. The pre-earnings analyst average was roughly 44%. NVIDIA said customer demand could support growth closer to 100%, but memory, components, packaging, power, and data-center capacity limit what the company can actually ship.&lt;/p&gt;
&lt;p&gt;The quarter weakened the claim that AI capital spending will peak in 2026. It also created a more uncomfortable question. NVIDIA is no longer only selling semiconductors. It is increasingly coordinating land, power, data centers, customer credit, and financing around its AI infrastructure. Investors now need to assess not only the size of growth, but also how much NVIDIA capital and credit support that growth requires.&lt;/p&gt;
&lt;h2 id="tldr"&gt;TL;DR
&lt;/h2&gt;&lt;ul&gt;
&lt;li&gt;Revenue of $96.2 billion and adjusted EPS of $2.22 exceeded expectations. Fiscal Q3 revenue guidance of $108 billion was also above the LSEG average of $104.19 billion. The outlook assumes no Data Center compute revenue from China.&lt;/li&gt;
&lt;li&gt;Hyperscale revenue rose 13% sequentially to $48.7 billion. AI Clouds, Industrial, and Enterprise, or ACIE, rose 25% to $40.3 billion. ACIE generated approximately 59% of the sequential increase in Data Center revenue.&lt;/li&gt;
&lt;li&gt;Supply and capacity commitments rose from $119 billion to $279 billion in one quarter, primarily for memory procurement. Rising memory prices reduce NVIDIA&amp;rsquo;s gross margin but directly confirm stronger pricing power for HBM and DRAM suppliers.&lt;/li&gt;
&lt;li&gt;The risk is visible in receivables and credit support. Days sales outstanding rose from 45 to 60, while five direct customers represented 70% of receivables. NVIDIA also disclosed $366 billion of multi-year commitments, another $56 billion tied to AI clouds and third-party leases, and maximum gross guarantee exposure of $108.5 billion.&lt;/li&gt;
&lt;li&gt;Recalculating the source report&amp;rsquo;s scenario framework suggests that a stock price around $219 offers little margin of safety if fiscal 2028 growth is 60%. There is upside if 70% is delivered, but more than 20% downside if growth returns to 45%. This is a sensitivity analysis, not a price target.&lt;/li&gt;
&lt;/ul&gt;
&lt;div class="thesis-callout"&gt;
&lt;div class="thesis-callout__label"&gt;Core view&lt;/div&gt;
&lt;p&gt;NVIDIA&amp;rsquo;s quarter showed that the AI constraint is moving. The bottleneck is shifting from GPU demand to memory, power, land, packaging, and credit. NVIDIA is beginning to procure and guarantee those constraints directly. The growth outlook improved, but so did the number of balance-sheet items required to judge growth quality.&lt;/p&gt;
&lt;/div&gt;
&lt;h2 id="1-incremental-profit-was-stronger-than-headline-revenue"&gt;1. Incremental profit was stronger than headline revenue
&lt;/h2&gt;&lt;table&gt;
 &lt;thead&gt;
 &lt;tr&gt;
 &lt;th&gt;Metric&lt;/th&gt;
 &lt;th style="text-align: right"&gt;Fiscal Q2 2027&lt;/th&gt;
 &lt;th style="text-align: right"&gt;Comparison&lt;/th&gt;
 &lt;th&gt;Reading&lt;/th&gt;
 &lt;/tr&gt;
 &lt;/thead&gt;
 &lt;tbody&gt;
 &lt;tr&gt;
 &lt;td&gt;Revenue&lt;/td&gt;
 &lt;td style="text-align: right"&gt;$96.2B&lt;/td&gt;
 &lt;td style="text-align: right"&gt;$92.27B market average&lt;/td&gt;
 &lt;td&gt;4.3% beat&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Data Center revenue&lt;/td&gt;
 &lt;td style="text-align: right"&gt;$89.0B&lt;/td&gt;
 &lt;td style="text-align: right"&gt;+117% year over year&lt;/td&gt;
 &lt;td&gt;92.5% of total revenue&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Adjusted EPS&lt;/td&gt;
 &lt;td style="text-align: right"&gt;$2.22&lt;/td&gt;
 &lt;td style="text-align: right"&gt;$2.09-$2.10 market average&lt;/td&gt;
 &lt;td&gt;Beat&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Non-GAAP operating income&lt;/td&gt;
 &lt;td style="text-align: right"&gt;$64.0B&lt;/td&gt;
 &lt;td style="text-align: right"&gt;+19% sequentially&lt;/td&gt;
 &lt;td&gt;Operating leverage continued&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Fiscal Q3 revenue guide&lt;/td&gt;
 &lt;td style="text-align: right"&gt;$108B, plus or minus 2%&lt;/td&gt;
 &lt;td style="text-align: right"&gt;$104.19B LSEG average&lt;/td&gt;
 &lt;td&gt;3.7% above consensus&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Fiscal Q3 gross-margin guide&lt;/td&gt;
 &lt;td style="text-align: right"&gt;74.0%, plus or minus 50 bps&lt;/td&gt;
 &lt;td style="text-align: right"&gt;75.0% in Q2&lt;/td&gt;
 &lt;td&gt;Memory cost pressure&lt;/td&gt;
 &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;Revenue increased by $14.6 billion sequentially, while non-GAAP operating income increased by $10.2 billion. The implied incremental operating margin was approximately 69.6%. Gross margin was almost unchanged at 75%, yet most of the additional revenue converted into operating profit.&lt;/p&gt;
&lt;p&gt;GAAP net income requires more caution. GAAP EPS of $2.46 included $7.8 billion of net gains from equity securities. Those gains are not recurring product economics. Adjusted EPS of $2.22 and operating income provide a cleaner view of the quarter&amp;rsquo;s underlying business performance.&lt;/p&gt;
&lt;h2 id="2-the-70-fiscal-2028-outlook-is-a-supply-number-more-than-a-demand-number"&gt;2. The 70% fiscal 2028 outlook is a supply number more than a demand number
&lt;/h2&gt;&lt;p&gt;NVIDIA expects fiscal 2028 revenue to grow by approximately 70%. Providing an annual growth outlook one year ahead was unusual for the company. The analyst average before earnings was about 44%.&lt;/p&gt;
&lt;p&gt;The key is how management framed the outlook. Customer demand forecasts could support growth closer to 100%, but actual revenue is limited by memory, components, advanced packaging, power, and data-center capacity. The 70% figure is therefore closer to the upper bound of currently securable supply than a pure demand forecast.&lt;/p&gt;
&lt;p&gt;If that framing is correct, investors need a different dashboard. Big Tech capex is no longer sufficient. HBM contracts, Rubin yields, advanced-packaging capacity, gigawatt-scale power connections, land availability, and financing terms for AI clouds become equally important.&lt;/p&gt;
&lt;p&gt;There is upside if one of those constraints eases. The fiscal Q3 guide includes no Data Center compute revenue from China. A regulatory opening or faster supply expansion could create upside. Continued export restrictions combined with supply delays would put even the 70% outlook at risk.&lt;/p&gt;
&lt;h2 id="3-incremental-growth-moved-beyond-the-largest-cloud-platforms"&gt;3. Incremental growth moved beyond the largest cloud platforms
&lt;/h2&gt;&lt;p&gt;NVIDIA split fiscal Q2 Data Center revenue into two customer groups.&lt;/p&gt;
&lt;table&gt;
 &lt;thead&gt;
 &lt;tr&gt;
 &lt;th&gt;Customer group&lt;/th&gt;
 &lt;th style="text-align: right"&gt;Fiscal Q2 2027 revenue&lt;/th&gt;
 &lt;th style="text-align: right"&gt;Sequential growth&lt;/th&gt;
 &lt;th style="text-align: right"&gt;Year-over-year growth&lt;/th&gt;
 &lt;/tr&gt;
 &lt;/thead&gt;
 &lt;tbody&gt;
 &lt;tr&gt;
 &lt;td&gt;Hyperscale&lt;/td&gt;
 &lt;td style="text-align: right"&gt;$48.7B&lt;/td&gt;
 &lt;td style="text-align: right"&gt;+13%&lt;/td&gt;
 &lt;td style="text-align: right"&gt;+102%&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;ACIE&lt;/td&gt;
 &lt;td style="text-align: right"&gt;$40.3B&lt;/td&gt;
 &lt;td style="text-align: right"&gt;+25%&lt;/td&gt;
 &lt;td style="text-align: right"&gt;+138%&lt;/td&gt;
 &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;ACIE includes AI clouds, industrial, and enterprise customers. Of the $13.8 billion sequential increase in Data Center revenue, ACIE contributed $8.1 billion, or approximately 59%. Hyperscale remains larger, but the center of marginal growth moved outside the largest cloud platforms.&lt;/p&gt;
&lt;p&gt;That finding challenges the common view that NVIDIA is simply a capex derivative of Microsoft, Amazon, Google, and Meta. Neoclouds, sovereign AI, and enterprise customers are creating incremental demand. Yet many of these customers have weaker balance sheets than hyperscalers. Demand broadened, but credit quality became more uneven.&lt;/p&gt;
&lt;p&gt;AWS&amp;rsquo;s plan to deploy two million additional GPUs provides the clearest hard order signal. AWS and NVIDIA said the deployment will span Blackwell Ultra, Rubin, and Rubin Ultra GPUs in 2027 and 2028, alongside Vera CPUs, networking, models, and robotics software. AWS&amp;rsquo;s Trainium and NVIDIA GPUs increasingly look like workload-specific complements rather than a simple substitution story.&lt;/p&gt;
&lt;h2 id="4-nvidia-is-selling-the-entire-ai-factory-not-only-the-gpu"&gt;4. NVIDIA is selling the entire AI factory, not only the GPU
&lt;/h2&gt;&lt;p&gt;Vera Rubin is not a single-chip transition. It combines the Vera CPU, Rubin GPUs, NVLink, Spectrum networking, BlueField, Groq 3 LPX, and software into one system. NVIDIA began production shipments in fiscal Q3 and said racks were running at CoreWeave, Google Cloud, Microsoft Azure, Oracle Cloud Infrastructure, and Nebius.&lt;/p&gt;
&lt;p&gt;Management estimated NVIDIA revenue opportunity per gigawatt at roughly $18 billion for Hopper, $25 billion for Blackwell, and $40 billion for Vera Rubin. Those are company estimates, not independently verified market figures. The strategic direction is still clear. NVIDIA is expanding the portion of each fixed gigawatt of infrastructure that it monetizes.&lt;/p&gt;
&lt;p&gt;The CPU push follows the same logic. Trailing-12-month Grace CPU revenue exceeded $5 billion. Vera is positioned as a controller for agent planning, tool calls, memory management, and GPU work allocation. If NVIDIA internalizes the CPU alongside GPUs and networking, Intel and AMD face a longer-term competitive pressure.&lt;/p&gt;
&lt;h2 id="5-lower-gross-margin-confirms-memory-suppliers-pricing-power"&gt;5. Lower gross margin confirms memory suppliers&amp;rsquo; pricing power
&lt;/h2&gt;&lt;p&gt;NVIDIA said memory prices rose more than it had expected. Management&amp;rsquo;s call commentary suggested a gross-margin path of 75.0% in fiscal Q2, 74.0% in Q3, 71%-72% in Q4, and 72%-73% in fiscal 2028 after product price increases begin to offset input costs.&lt;/p&gt;
&lt;p&gt;The margin rate falls, but gross profit still rises. At the midpoint of the fiscal Q3 guide, gross profit would increase from $72.1 billion in Q2 to $79.9 billion in Q3, or about 10.8%. This is not an earnings contraction. It is a period in which part of NVIDIA&amp;rsquo;s excess economics shifts to HBM and DRAM suppliers.&lt;/p&gt;
&lt;p&gt;The strongest official evidence is the commitment schedule. Supply and capacity commitments increased from $119 billion to $279 billion, primarily for memory procurement. NVIDIA did not disclose the supplier allocation, so the exact shares of SK hynix, Micron, and Samsung Electronics remain unknown.&lt;/p&gt;
&lt;p&gt;The August 27 Korean close pointed in the same direction. KIS data showed Samsung Electronics up 1.72%, SK hynix up 2.49%, and Hanmi Semiconductor up 1.14%. The market recognized the memory signal but did not reprice the entire equipment chain equally.&lt;/p&gt;
&lt;table&gt;
 &lt;thead&gt;
 &lt;tr&gt;
 &lt;th&gt;Value-chain segment&lt;/th&gt;
 &lt;th&gt;Impact&lt;/th&gt;
 &lt;th&gt;Next metric to verify&lt;/th&gt;
 &lt;/tr&gt;
 &lt;/thead&gt;
 &lt;tbody&gt;
 &lt;tr&gt;
 &lt;td&gt;HBM and DRAM&lt;/td&gt;
 &lt;td&gt;Most direct positive&lt;/td&gt;
 &lt;td&gt;Contract pricing, bit supply growth, supplier allocation&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Foundry&lt;/td&gt;
 &lt;td&gt;Positive volume effect&lt;/td&gt;
 &lt;td&gt;Rubin yields and wafer starts&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Advanced packaging&lt;/td&gt;
 &lt;td&gt;Capacity-expansion beneficiary&lt;/td&gt;
 &lt;td&gt;CoWoS and bonding-equipment orders, utilization&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Networking and optics&lt;/td&gt;
 &lt;td&gt;Traffic upside and NVIDIA integration risk&lt;/td&gt;
 &lt;td&gt;Spectrum-X, custom ASIC, optical-module revenue&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Competing GPUs&lt;/td&gt;
 &lt;td&gt;Larger market, relative share pressure&lt;/td&gt;
 &lt;td&gt;AMD accelerator revenue and customer breadth&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Server CPUs&lt;/td&gt;
 &lt;td&gt;Negative from NVIDIA&amp;rsquo;s direct entry&lt;/td&gt;
 &lt;td&gt;Vera adoption and Grace revenue&lt;/td&gt;
 &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;h2 id="6-the-changes-already-incurred-are-visible-in-cash-flow-and-collection-periods"&gt;6. The changes already incurred are visible in cash flow and collection periods
&lt;/h2&gt;&lt;p&gt;Fiscal Q2 free cash flow fell to $21.3 billion from $48.6 billion in the prior quarter. Operating cash flow declined to $24.1 billion from $50.3 billion. Working capital and cash taxes explain much of the change.&lt;/p&gt;
&lt;p&gt;Accounts receivable increased to $63.1 billion and days sales outstanding rose from 45 to 60. NVIDIA disclosed that it may offer certain investment-grade customers payment terms of 90 days to as long as one year for large, multi-quarter agreements. Five direct customers represented 22%, 14%, 13%, 11%, and 10% of receivables, or 70% in total.&lt;/p&gt;
&lt;p&gt;Direct customers include distributors, ODMs, OEMs, cloud providers, AI model makers, and system integrators. The concentration cannot be read as end-customer concentration. It still shows that collection periods lengthened while receivables became concentrated among a small number of direct counterparties.&lt;/p&gt;
&lt;p&gt;Inventory also rose from $25.8 billion to $31.6 billion as NVIDIA prepared for Rubin. That is reasonable launch inventory if demand forecasts prove correct. If Rubin ramps late or memory prices reverse, the same inventory and supply commitments could create provisions or losses.&lt;/p&gt;
&lt;h2 id="7-commitments-and-guarantees-are-not-yet-losses-but-they-require-separate-scrutiny"&gt;7. Commitments and guarantees are not yet losses, but they require separate scrutiny
&lt;/h2&gt;&lt;p&gt;NVIDIA&amp;rsquo;s commitments need to be separated into three layers.&lt;/p&gt;
&lt;table&gt;
 &lt;thead&gt;
 &lt;tr&gt;
 &lt;th&gt;Category&lt;/th&gt;
 &lt;th style="text-align: right"&gt;Maximum amount&lt;/th&gt;
 &lt;th&gt;Economic nature&lt;/th&gt;
 &lt;/tr&gt;
 &lt;/thead&gt;
 &lt;tbody&gt;
 &lt;tr&gt;
 &lt;td&gt;Supply, cloud, lease, investment, and capital commitments&lt;/td&gt;
 &lt;td style="text-align: right"&gt;$366B&lt;/td&gt;
 &lt;td&gt;Multi-year purchasing and investment plans&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Separate AI cloud and third-party lease commitments&lt;/td&gt;
 &lt;td style="text-align: right"&gt;$56B&lt;/td&gt;
 &lt;td&gt;Customer infrastructure support&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Land, power, and shell guarantees&lt;/td&gt;
 &lt;td style="text-align: right"&gt;$108.5B&lt;/td&gt;
 &lt;td&gt;Exposure only if specified conditions occur&lt;/td&gt;
 &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;Adding these amounts as current debt or expected losses would be wrong. The commitments span multiple years, and guarantees require a trigger. It would also be wrong to ignore them. In a downturn, equipment orders, investment marks, guarantee losses, and delayed receivable collections could deteriorate together.&lt;/p&gt;
&lt;p&gt;The largest item is the Ohio SB Energy site to be used by OpenAI. NVIDIA may provide up to $105 billion of guarantees for land, power, and shell leases covering approximately 4.25 gigawatts. The first guarantee is expected to become effective in fiscal 2029, exposure declines as OpenAI makes lease payments, and the guarantee covers defined portions of lease and power payments rather than all site costs or all customer obligations.&lt;/p&gt;
&lt;p&gt;This structure is not automatically circular revenue. Installed equipment serving third-party users has real economic substance. But the model is changing from customers buying GPUs solely with their own cash flow and credit to NVIDIA contributing capital and credit to an ecosystem that purchases NVIDIA equipment. Credit cost is becoming an internal variable of the business model.&lt;/p&gt;
&lt;h2 id="8-separate-observed-agent-demand-from-management-narrative"&gt;8. Separate observed agent demand from management narrative
&lt;/h2&gt;&lt;p&gt;The hard data are strong.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;ACIE revenue increased 25% sequentially and 138% year over year.&lt;/li&gt;
&lt;li&gt;AWS plans to deploy two million additional NVIDIA GPUs in 2027 and 2028.&lt;/li&gt;
&lt;li&gt;Rubin production shipments began and racks are running at major cloud partners.&lt;/li&gt;
&lt;li&gt;Trailing-12-month Grace CPU revenue exceeded $5 billion.&lt;/li&gt;
&lt;li&gt;Networking, CPUs, storage, and software are being sold as one agent infrastructure stack.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Other claims remain closer to management hypotheses. Jensen Huang said agent workloads can require 15 to 100 times more computation than a single conventional query. NVIDIA&amp;rsquo;s Rubin cost and performance claims and its vision of many more always-on agents than employees are not independent industry statistics.&lt;/p&gt;
&lt;p&gt;The demand mechanism is plausible. Agents repeatedly plan, search, run code, call external tools, evaluate results, and re-plan. A single task can create many model calls and memory accesses. Even if token cost falls, total compute can rise if the number of automatable tasks grows faster.&lt;/p&gt;
&lt;p&gt;Efficiency could still win. Smaller models, routing, KV-cache optimization, and custom ASICs can reduce GPU use per task. Security, reliability, and liability concerns can delay enterprise autonomy. Strong infrastructure demand does not guarantee attractive economics for every agent software company.&lt;/p&gt;
&lt;h2 id="9-the-current-price-is-highly-sensitive-to-execution-of-the-70-outlook"&gt;9. The current price is highly sensitive to execution of the 70% outlook
&lt;/h2&gt;&lt;p&gt;The table below recalculates the source report&amp;rsquo;s assumptions. It is not a consensus target price.&lt;/p&gt;
&lt;p&gt;The assumptions are fiscal 2027 revenue of $405.8 billion, 24.1 billion diluted shares, a 12% discount rate, and a 17-month discount period from the end of fiscal 2028. Fiscal 2027 revenue combines first-half actual revenue of $177.8 billion, fiscal Q3 guidance of $108 billion, and a $120 billion fiscal Q4 assumption.&lt;/p&gt;
&lt;table&gt;
 &lt;thead&gt;
 &lt;tr&gt;
 &lt;th&gt;Scenario&lt;/th&gt;
 &lt;th style="text-align: right"&gt;Fiscal 2028 growth&lt;/th&gt;
 &lt;th style="text-align: right"&gt;Net margin&lt;/th&gt;
 &lt;th style="text-align: right"&gt;EPS&lt;/th&gt;
 &lt;th style="text-align: right"&gt;P/E&lt;/th&gt;
 &lt;th style="text-align: right"&gt;Present value&lt;/th&gt;
 &lt;th style="text-align: right"&gt;Versus $219.5&lt;/th&gt;
 &lt;/tr&gt;
 &lt;/thead&gt;
 &lt;tbody&gt;
 &lt;tr&gt;
 &lt;td&gt;Bear&lt;/td&gt;
 &lt;td style="text-align: right"&gt;45%&lt;/td&gt;
 &lt;td style="text-align: right"&gt;45%&lt;/td&gt;
 &lt;td style="text-align: right"&gt;$10.99&lt;/td&gt;
 &lt;td style="text-align: right"&gt;18x&lt;/td&gt;
 &lt;td style="text-align: right"&gt;About $168&lt;/td&gt;
 &lt;td style="text-align: right"&gt;-23%&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Base&lt;/td&gt;
 &lt;td style="text-align: right"&gt;60%&lt;/td&gt;
 &lt;td style="text-align: right"&gt;49%&lt;/td&gt;
 &lt;td style="text-align: right"&gt;$13.20&lt;/td&gt;
 &lt;td style="text-align: right"&gt;20x&lt;/td&gt;
 &lt;td style="text-align: right"&gt;About $225&lt;/td&gt;
 &lt;td style="text-align: right"&gt;+2%&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Bull&lt;/td&gt;
 &lt;td style="text-align: right"&gt;70%&lt;/td&gt;
 &lt;td style="text-align: right"&gt;52%&lt;/td&gt;
 &lt;td style="text-align: right"&gt;$14.89&lt;/td&gt;
 &lt;td style="text-align: right"&gt;22x&lt;/td&gt;
 &lt;td style="text-align: right"&gt;About $279&lt;/td&gt;
 &lt;td style="text-align: right"&gt;+27%&lt;/td&gt;
 &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;If NVIDIA delivers its 70% outlook, upside remains. If growth is discounted to 60%, the after-hours price near $219 is close to fair value. If growth returns to roughly the pre-earnings 45% expectation, downside exceeds 20%.&lt;/p&gt;
&lt;p&gt;The weaknesses are significant. Fiscal Q4 revenue, net margins, multiples, and discount rates are all fixed assumptions. Equity gains, buybacks, product price increases, and China revenue are simplified. The table is a tool for identifying which assumptions support the current price, not a forecast.&lt;/p&gt;
&lt;h2 id="10-what-would-invalidate-this-reading"&gt;10. What would invalidate this reading
&lt;/h2&gt;&lt;p&gt;The positive interpretation weakens if:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Fiscal 2028 revenue growth guidance falls below 50%.&lt;/li&gt;
&lt;li&gt;Gross margin remains below 70% after product price increases.&lt;/li&gt;
&lt;li&gt;DSO stays above 65 days for two or more quarters.&lt;/li&gt;
&lt;li&gt;Rubin yields, HBM procurement, advanced packaging, or power connections run behind plan.&lt;/li&gt;
&lt;li&gt;Financing costs for AI clouds and model makers rise enough to increase NVIDIA&amp;rsquo;s actual guarantee exposure.&lt;/li&gt;
&lt;li&gt;HBM contract prices fall while bit-supply growth accelerates.&lt;/li&gt;
&lt;li&gt;Customer adoption of custom ASICs reduces NVIDIA&amp;rsquo;s wallet share per system.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The 70% outlook becomes more credible if ACIE outgrows hyperscale for another two quarters, Rubin scales on schedule, DSO moves back toward 50 days, and memory commitments convert into delivered systems without large provisions.&lt;/p&gt;
&lt;h2 id="11-conclusion-the-owner-of-the-bottleneck-matters-more-than-another-earnings-beat"&gt;11. Conclusion: the owner of the bottleneck matters more than another earnings beat
&lt;/h2&gt;&lt;p&gt;The quarter weakened the argument that AI infrastructure investment will peak in 2026. Demand broadened beyond Big Tech, while Rubin and AWS&amp;rsquo;s two-million-GPU plan attached hard orders to the next growth phase.&lt;/p&gt;
&lt;p&gt;The clearest second-order effect is memory. Once NVIDIA disclosed $279 billion of supply and capacity commitments primarily related to memory, HBM and DRAM pricing power was no longer merely an industry forecast. Lower NVIDIA gross margin is closer to a transfer of incremental economics to memory suppliers than a decline in absolute NVIDIA profit.&lt;/p&gt;
&lt;p&gt;NVIDIA&amp;rsquo;s role is changing at the same time. A company that sold GPUs is now coordinating land, power, long-term leases, guarantees, and equity investments. That can extend the duration of growth, but it also layers credit and capital losses onto the ordinary risks of a product cycle.&lt;/p&gt;
&lt;p&gt;The peak-AI-capex debate is not over. Its center has moved from GPU orders to bottlenecks in memory, power, land, and credit.&lt;/p&gt;
&lt;p&gt;The next quarter should be judged on whether the 70% outlook converts into installed supply, whether receivables and guarantee exposure stabilize, and whether memory suppliers retain higher pricing in their own earnings.&lt;/p&gt;
&lt;h2 id="sources-and-timing"&gt;Sources and timing
&lt;/h2&gt;&lt;ul&gt;
&lt;li&gt;&lt;a class="link" href="https://investor.nvidia.com/news/press-release-details/2026/NVIDIA-Announces-Financial-Results-for-Second-Quarter-Fiscal-2027/default.aspx" target="_blank" rel="noopener"
 &gt;NVIDIA fiscal Q2 2027 earnings release, August 26, 2026&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a class="link" href="https://www.sec.gov/Archives/edgar/data/1045810/000104581026000073/q2fy27cfocommentary.htm" target="_blank" rel="noopener"
 &gt;NVIDIA fiscal Q2 2027 CFO Commentary, SEC Exhibit 99.2&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a class="link" href="https://www.sec.gov/Archives/edgar/data/1045810/000104581026000075/nvda-20260726.htm" target="_blank" rel="noopener"
 &gt;NVIDIA fiscal Q2 2027 Form 10-Q, August 26, 2026&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a class="link" href="https://nvidianews.nvidia.com/news/aws-and-nvidia-to-deliver-2-million-additional-gpus-and-next-generation-infrastructure-for-agentic-and-physical-ai" target="_blank" rel="noopener"
 &gt;AWS and NVIDIA plan to deploy two million additional GPUs, August 26, 2026&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a class="link" href="https://live.euronext.com/en/financial-news/nvidia-forecasts-70-sales-growth-next-year-signals-ai-spending-boom-has-years-left" target="_blank" rel="noopener"
 &gt;Reuters on the fiscal 2028 growth outlook and after-hours reaction, August 27, 2026&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a class="link" href="https://investor.nvidia.com/events-and-presentations/events-and-presentations/event-details/2026/NVIDIA-2nd-Quarter-FY27-Financial-Results/default.aspx" target="_blank" rel="noopener"
 &gt;NVIDIA fiscal Q2 2027 earnings-call replay&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Korean close: Korea Investment &amp;amp; Securities KIS Open API, KRX close collected at 17:08 KST on August 27, 2026. Samsung Electronics KRW266,000, SK hynix KRW1,730,000, and Hanmi Semiconductor KRW222,000.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;small&gt;Data are current through the August 27, 2026 Korean close and the end of the August 26 U.S. after-hours session. Fiscal 2028 growth, the margin path, revenue opportunity per gigawatt, and agent compute intensity are management forecasts or earnings-call statements and are not guaranteed. Scenario values are sensitivities based on disclosed assumptions, not target prices. This article is for research and information purposes only and is not investment advice.&lt;/small&gt;&lt;/p&gt;</description></item><item><title>AI's Bottleneck Really Is HBM, Not the GPU: But Bottlenecks Invite Workarounds</title><link>https://koreainvestinsights.com/post/hbm-bottleneck-verified-and-its-workarounds-2026-08-16/</link><pubDate>Sun, 16 Aug 2026 14:30:00 +0900</pubDate><guid>https://koreainvestinsights.com/post/hbm-bottleneck-verified-and-its-workarounds-2026-08-16/</guid><description>
 &lt;blockquote&gt;
 &lt;p&gt;Series context: over the course of August, this series has covered in turn the deceleration in memory prices, the gap between confirmed demand and the stock price, and the conditions for a rebound. This piece verifies the physical foundation beneath all of that: how far the claim that AI&amp;rsquo;s bottleneck has moved from the GPU to HBM actually holds, what profit that bottleneck guarantees suppliers and for how long, and what countervailing force the bottleneck itself is generating.&lt;/p&gt;

 &lt;/blockquote&gt;
&lt;h2 id="tldr"&gt;TL;DR
&lt;/h2&gt;&lt;ul&gt;
&lt;li&gt;Start with where the claim and its numbers come from. Morgan Stanley&amp;rsquo;s June report offers the frame that &lt;strong&gt;&amp;ldquo;the bottleneck moves along the supply chain: first chips, then power, then memory, then networking, then cooling&amp;rdquo;&lt;/strong&gt;, along with the judgment that &amp;ldquo;memory will be in shortage through the end of 2026.&amp;rdquo; [Fact: Morgan Stanley, June 9] But &lt;strong&gt;the roughly 50 billion Gb of 2027 HBM demand often cited alongside it is a figure absent from that document&lt;/strong&gt;. The real sources are Kiwoom Securities (53.4 billion Gb) and Korean research aggregation (6.20 exabytes, about 49.6 billion Gb converted), and working backward from TrendForce&amp;rsquo;s bit-share forecast lands in the same range. Because three separate lines converge, the number is usable, but its attribution needs to be corrected.&lt;/li&gt;
&lt;li&gt;The first basis for the bottleneck is the gap in growth rates between compute and bandwidth. A single GPU&amp;rsquo;s compute has grown three to four times per generation, but HBM capacity is &lt;strong&gt;the same 288GB for both GB300 and the next-generation Rubin&lt;/strong&gt;. What did grow is bandwidth, from the 8 terabyte class to the 20 terabyte class, and even that falls short of the roughly 3.5x growth in compute. [Fact: NVIDIA specifications, SemiAnalysis] When memory cannot keep up, buying more chips just means compute sits and waits.&lt;/li&gt;
&lt;li&gt;The second basis is a change in the nature of inference. When an agent works for an extended stretch while holding a long context, &lt;strong&gt;the KV cache, an intermediate memory, occupies memory for the whole session&lt;/strong&gt;. Field data shows the processing cost of a cached token versus an uncached one differs by a factor of 10, and NVIDIA thought the problem large enough to launch a dedicated cache storage tier as a separate product line earlier this year. [Fact: industry data] That said, no published quantitative estimate yet exists for how many gigabytes of additional HBM demand agent workloads generate. [Blocked: no public estimate]&lt;/li&gt;
&lt;li&gt;The third basis is the arithmetic of supply. Producing the same capacity in HBM uses &lt;strong&gt;roughly three times the wafers of ordinary DRAM&lt;/strong&gt;, and even if a single die&amp;rsquo;s yield is 95%, cumulative yield falls to 66% at 8 layers of stacking and 44% at 16 layers. A fab takes three to five years from groundbreaking to volume production, and new capacity does not contribute meaningfully until 2028. [Fact: industry data, TrendForce] That is why 2027 tightness looks less like a forecast and more like physics. Micron said it cannot fill even half of data center demand and that &lt;strong&gt;&amp;ldquo;2027 will be tighter than 2026&amp;rdquo;&lt;/strong&gt;, and SK Hynix CEO Kwak Noh-jung described 2027, from a supply standpoint, as the tightest year in the industry&amp;rsquo;s history. [Fact: company statements]&lt;/li&gt;
&lt;li&gt;But the same arithmetic is moving buyers too. Once memory cost reached 29% of system cost, NVIDIA &lt;strong&gt;cut the CPU-side memory module on Vera Rubin in half&lt;/strong&gt;, reports say Rubin Ultra&amp;rsquo;s HBM will be set at 192GB, below the prior generation&amp;rsquo;s 288GB, and the prefill-only chip (Rubin CPX) uses GDDR7 instead of HBM. The High Bandwidth Flash that SanDisk and SK Hynix standardized builds a cheaper capacity tier beneath HBM. [Fact: TrendForce, company announcements] &lt;strong&gt;The deeper the bottleneck, the bigger the reward for working around it.&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;Where the bottleneck&amp;rsquo;s profit ends up is a matter of contracts. Micron has locked in roughly half its revenue through contracts with 16 strategic customers running to around 2030, &amp;ldquo;several-fold&amp;rdquo; increases in 2027 HBM contract prices are being discussed, and custom base dies tailored to individual customers (custom HBM) lock supplier and customer to each other. [Fact: company data, TrendForce] The rent is being fixed in the form of multi-year contracts rather than a spike in spot prices.&lt;/li&gt;
&lt;li&gt;The implication for Korea. Two of the three oligopoly players are Korean, and SK Hynix&amp;rsquo;s HBM share sits in the mid-to-high 50% range. Profit visibility through 2027 is backed by the physics of the bottleneck. What remains open is 2028. New capacity lands that year, and CXMT&amp;rsquo;s HBM also arrives around then, still carrying yield problems. That is also when the workaround technologies described above accumulate enough effect to matter. The bottleneck claim is correct, but it is &lt;strong&gt;a claim that holds through 2027&lt;/strong&gt;, and beyond that it becomes a matter of contract structure and supply discipline.&lt;/li&gt;
&lt;/ul&gt;
&lt;div class="thesis-callout"&gt;
&lt;div class="thesis-callout__label"&gt;Key Framing&lt;/div&gt;
&lt;p&gt;The diagnosis that AI&amp;rsquo;s bottleneck has moved from the GPU to HBM passes verification. The gap between compute and bandwidth, the memory occupied by agent inference, three times the wafers for half the yield, and a three-year lead time are all measured facts. But the bottleneck is not static. Once memory reached three-tenths of system cost, the largest buyer began stripping memory out of its designs, and the owners of the bottleneck are themselves fixing prices through multi-year contracts to narrow the swing of the next downturn. The profit coming out of the bottleneck is real, but its size is set by contract terms rather than market price, and its lifespan runs only until new supply arrives in 2028. What investors should ask is not whether the bottleneck exists, but whose books the bottleneck&amp;rsquo;s rent lands on, and in what form.&lt;/p&gt;
&lt;/div&gt;
&lt;h2 id="1-where-the-claim-and-its-numbers-come-from"&gt;1. Where the Claim and Its Numbers Come From
&lt;/h2&gt;&lt;p&gt;Morgan Stanley Investment Management&amp;rsquo;s June report, &amp;ldquo;Big Picture: Ten Truths About Investing in AI,&amp;rdquo; is the document cited as the origin of this claim. What the document actually argues is this. The bottleneck in AI infrastructure does not stay in one place; it moves along the supply chain. It was chips first, then power, now memory, and next comes networking and cooling. Memory will be in shortage through the end of 2026, and incremental memory demand in 2027 will run 75 to 100 exabytes, doubling again in 2028. Agentic AI requires roughly a million times the compute of the original conversational models, and token consumption grew more than tenfold in 2025 alone. [Fact: Morgan Stanley, June 9, 2026]&lt;/p&gt;
&lt;p&gt;There is one thing to correct here. The figure widely cited alongside this report, &amp;ldquo;roughly 50 billion Gb in 2027 AI-related HBM demand,&amp;rdquo; &lt;strong&gt;does not appear in the document&lt;/strong&gt;. The document contains no HBM-specific figures and does not name any of the three suppliers. Tracing this number to its actual source turns up two lines. Kiwoom Securities estimated 2027 HBM demand at 53.4 billion Gb, a 56% increase from the prior year, and Korean research aggregating TrendForce, McKinsey, and others put it at 6.20 exabytes, roughly 49.6 billion Gb when converted to gigabits. [Fact: Kiwoom Securities, VLSI Research Korea]&lt;/p&gt;
&lt;figure class="kii-figure"&gt;
&lt;div class="kii-figure__frame"&gt;
&lt;svg viewBox="0 0 700 300" xmlns="http://www.w3.org/2000/svg" role="img"&gt;
&lt;line x1="60" y1="236.0" x2="676.0" y2="236.0" stroke="var(--kii-chart-grid)" stroke-width="1"/&gt;
&lt;text x="51" y="240.0" fill="var(--card-text-color-tertiary)" font-size="11" text-anchor="end"&gt;0.0&lt;/text&gt;
&lt;line x1="60" y1="196.8" x2="676.0" y2="196.8" stroke="var(--kii-chart-grid)" stroke-width="1"/&gt;
&lt;text x="51" y="200.8" fill="var(--card-text-color-tertiary)" font-size="11" text-anchor="end"&gt;12.0&lt;/text&gt;
&lt;line x1="60" y1="157.6" x2="676.0" y2="157.6" stroke="var(--kii-chart-grid)" stroke-width="1"/&gt;
&lt;text x="51" y="161.6" fill="var(--card-text-color-tertiary)" font-size="11" text-anchor="end"&gt;23.9&lt;/text&gt;
&lt;line x1="60" y1="118.4" x2="676.0" y2="118.4" stroke="var(--kii-chart-grid)" stroke-width="1"/&gt;
&lt;text x="51" y="122.4" fill="var(--card-text-color-tertiary)" font-size="11" text-anchor="end"&gt;35.9&lt;/text&gt;
&lt;line x1="60" y1="79.2" x2="676.0" y2="79.2" stroke="var(--kii-chart-grid)" stroke-width="1"/&gt;
&lt;text x="51" y="83.2" fill="var(--card-text-color-tertiary)" font-size="11" text-anchor="end"&gt;47.8&lt;/text&gt;
&lt;line x1="60" y1="40.0" x2="676.0" y2="40.0" stroke="var(--kii-chart-grid)" stroke-width="1"/&gt;
&lt;text x="51" y="44.0" fill="var(--card-text-color-tertiary)" font-size="11" text-anchor="end"&gt;59.8&lt;/text&gt;
&lt;rect x="111.3" y="61.0" width="102.7" height="175.0" rx="4" fill="var(--kii-cat-1)"/&gt;
&lt;text x="162.7" y="53.0" fill="var(--card-text-color-main)" font-size="12.5" font-weight="600" text-anchor="middle"&gt;53.4&lt;/text&gt;
&lt;text x="162.7" y="256.0" fill="var(--card-text-color-secondary)" font-size="12" text-anchor="middle"&gt;Kiwoom Securities&lt;/text&gt;
&lt;text x="162.7" y="274.0" fill="var(--card-text-color-tertiary)" font-size="10.5" text-anchor="middle"&gt;+56% YoY&lt;/text&gt;
&lt;rect x="316.7" y="73.5" width="102.7" height="162.5" rx="4" fill="var(--kii-cat-1)"/&gt;
&lt;text x="368.0" y="65.5" fill="var(--card-text-color-main)" font-size="12.5" font-weight="600" text-anchor="middle"&gt;49.6&lt;/text&gt;
&lt;text x="368.0" y="256.0" fill="var(--card-text-color-secondary)" font-size="12" text-anchor="middle"&gt;Research synthesis&lt;/text&gt;
&lt;text x="368.0" y="274.0" fill="var(--card-text-color-tertiary)" font-size="10.5" text-anchor="middle"&gt;6.20EB converted&lt;/text&gt;
&lt;rect x="522.0" y="72.1" width="102.7" height="163.9" rx="4" fill="var(--kii-cat-3)"/&gt;
&lt;text x="573.3" y="64.1" fill="var(--card-text-color-main)" font-size="12.5" font-weight="600" text-anchor="middle"&gt;50.0&lt;/text&gt;
&lt;text x="573.3" y="256.0" fill="var(--card-text-color-secondary)" font-size="12" text-anchor="middle"&gt;Share back-solve (own)&lt;/text&gt;
&lt;text x="573.3" y="274.0" fill="var(--card-text-color-tertiary)" font-size="10.5" text-anchor="middle"&gt;13% bit share&lt;/text&gt;
&lt;line x1="60" y1="236.0" x2="676.0" y2="236.0" stroke="var(--kii-chart-axis)" stroke-width="1"/&gt;
&lt;text x="60" y="24" fill="var(--card-text-color-tertiary)" font-size="11.5"&gt;2027 HBM demand estimates (bn Gb). Three different methods land in the same place&lt;/text&gt;
&lt;/svg&gt;
&lt;/div&gt;
&lt;figcaption&gt;&lt;strong&gt;Three routes to 50 billion Gb.&lt;/strong&gt; Kiwoom's estimate, a research synthesis converting 6.20 exabytes, and our own back-solve from a 13% bit share land in the same place. The figure does not appear in the Morgan Stanley document.&lt;/figcaption&gt;
&lt;details class="kii-figure__table"&gt;&lt;summary&gt;View as table&lt;/summary&gt;
&lt;table&gt;
 &lt;thead&gt;
 &lt;tr&gt;
 &lt;th&gt;Route&lt;/th&gt;
 &lt;th&gt;2027 HBM demand&lt;/th&gt;
 &lt;th&gt;Method&lt;/th&gt;
 &lt;/tr&gt;
 &lt;/thead&gt;
 &lt;tbody&gt;
 &lt;tr&gt;
 &lt;td&gt;Kiwoom Securities&lt;/td&gt;
 &lt;td&gt;53.4bn Gb&lt;/td&gt;
 &lt;td&gt;+56% year on year&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Research synthesis&lt;/td&gt;
 &lt;td&gt;about 49.6bn Gb&lt;/td&gt;
 &lt;td&gt;6.20 exabytes converted&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Own back-solve&lt;/td&gt;
 &lt;td&gt;about 50bn Gb&lt;/td&gt;
 &lt;td&gt;TrendForce 13% bit share&lt;/td&gt;
 &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;/details&gt;
&lt;/figure&gt;
&lt;p&gt;The number itself is solid. Two estimates using different methods converge near 50 billion Gb, and a third path verifies it as well. TrendForce expects HBM to reach about 13% of total DRAM bit supply in 2027; if 50 billion Gb is 13%, total DRAM works out to roughly 380 to 440 billion Gb, which matches the industry&amp;rsquo;s usual scale estimates. [Inference: our own back-calculation] The conclusion is this. The number is usable, but the source is Korean brokerages and research aggregation, not Morgan Stanley, and it should be cited that way.&lt;/p&gt;
&lt;h2 id="2-why-hbm-not-the-gpu-the-gap-in-growth-rates"&gt;2. Why HBM, Not the GPU: The Gap in Growth Rates
&lt;/h2&gt;&lt;p&gt;The first foundation of the bottleneck claim is the gap between how fast compute grows and how fast memory grows. Line up GPU specifications generation by generation and the gap is plain to see.&lt;/p&gt;
&lt;figure class="kii-figure"&gt;
&lt;div class="kii-figure__frame"&gt;
&lt;svg viewBox="0 0 700 300" xmlns="http://www.w3.org/2000/svg" role="img"&gt;
&lt;line x1="60" y1="236.0" x2="676.0" y2="236.0" stroke="var(--kii-chart-grid)" stroke-width="1"/&gt;
&lt;text x="51" y="240.0" fill="var(--card-text-color-tertiary)" font-size="11" text-anchor="end"&gt;0&lt;/text&gt;
&lt;line x1="60" y1="196.8" x2="676.0" y2="196.8" stroke="var(--kii-chart-grid)" stroke-width="1"/&gt;
&lt;text x="51" y="200.8" fill="var(--card-text-color-tertiary)" font-size="11" text-anchor="end"&gt;4.48&lt;/text&gt;
&lt;line x1="60" y1="157.6" x2="676.0" y2="157.6" stroke="var(--kii-chart-grid)" stroke-width="1"/&gt;
&lt;text x="51" y="161.6" fill="var(--card-text-color-tertiary)" font-size="11" text-anchor="end"&gt;8.96&lt;/text&gt;
&lt;line x1="60" y1="118.4" x2="676.0" y2="118.4" stroke="var(--kii-chart-grid)" stroke-width="1"/&gt;
&lt;text x="51" y="122.4" fill="var(--card-text-color-tertiary)" font-size="11" text-anchor="end"&gt;13.44&lt;/text&gt;
&lt;line x1="60" y1="79.2" x2="676.0" y2="79.2" stroke="var(--kii-chart-grid)" stroke-width="1"/&gt;
&lt;text x="51" y="83.2" fill="var(--card-text-color-tertiary)" font-size="11" text-anchor="end"&gt;17.92&lt;/text&gt;
&lt;line x1="60" y1="40.0" x2="676.0" y2="40.0" stroke="var(--kii-chart-grid)" stroke-width="1"/&gt;
&lt;text x="51" y="44.0" fill="var(--card-text-color-tertiary)" font-size="11" text-anchor="end"&gt;22.4&lt;/text&gt;
&lt;rect x="98.5" y="206.7" width="77.0" height="29.3" rx="4" fill="var(--kii-cat-1)"/&gt;
&lt;text x="137.0" y="198.7" fill="var(--card-text-color-main)" font-size="12.5" font-weight="600" text-anchor="middle"&gt;3.35&lt;/text&gt;
&lt;text x="137.0" y="256.0" fill="var(--card-text-color-secondary)" font-size="12" text-anchor="middle"&gt;H100&lt;/text&gt;
&lt;text x="137.0" y="274.0" fill="var(--card-text-color-tertiary)" font-size="10.5" text-anchor="middle"&gt;80GB HBM&lt;/text&gt;
&lt;rect x="252.5" y="166.0" width="77.0" height="70.0" rx="4" fill="var(--kii-cat-1)"/&gt;
&lt;text x="291.0" y="158.0" fill="var(--card-text-color-main)" font-size="12.5" font-weight="600" text-anchor="middle"&gt;8&lt;/text&gt;
&lt;text x="291.0" y="256.0" fill="var(--card-text-color-secondary)" font-size="12" text-anchor="middle"&gt;B200&lt;/text&gt;
&lt;text x="291.0" y="274.0" fill="var(--card-text-color-tertiary)" font-size="10.5" text-anchor="middle"&gt;192GB HBM&lt;/text&gt;
&lt;rect x="406.5" y="166.0" width="77.0" height="70.0" rx="4" fill="var(--kii-cat-1)"/&gt;
&lt;text x="445.0" y="158.0" fill="var(--card-text-color-main)" font-size="12.5" font-weight="600" text-anchor="middle"&gt;8&lt;/text&gt;
&lt;text x="445.0" y="256.0" fill="var(--card-text-color-secondary)" font-size="12" text-anchor="middle"&gt;GB300&lt;/text&gt;
&lt;text x="445.0" y="274.0" fill="var(--card-text-color-tertiary)" font-size="10.5" text-anchor="middle"&gt;288GB HBM&lt;/text&gt;
&lt;rect x="560.5" y="61.0" width="77.0" height="175.0" rx="4" fill="var(--kii-cat-3)"/&gt;
&lt;text x="599.0" y="53.0" fill="var(--card-text-color-main)" font-size="12.5" font-weight="600" text-anchor="middle"&gt;20&lt;/text&gt;
&lt;text x="599.0" y="256.0" fill="var(--card-text-color-secondary)" font-size="12" text-anchor="middle"&gt;Vera Rubin&lt;/text&gt;
&lt;text x="599.0" y="274.0" fill="var(--card-text-color-tertiary)" font-size="10.5" text-anchor="middle"&gt;288GB, capacity flat&lt;/text&gt;
&lt;line x1="60" y1="236.0" x2="676.0" y2="236.0" stroke="var(--kii-chart-axis)" stroke-width="1"/&gt;
&lt;text x="60" y="24" fill="var(--card-text-color-tertiary)" font-size="11.5"&gt;HBM bandwidth per GPU (TB/s). Rubin targets 22, early shipments reported near 20&lt;/text&gt;
&lt;/svg&gt;
&lt;/div&gt;
&lt;figcaption&gt;&lt;strong&gt;Compute jumps, capacity stalls.&lt;/strong&gt; GB300 and Vera Rubin carry the same 288GB of HBM. The generational gain is bandwidth, and even that (about 2.75x) trails the compute gain (about 3.5x) over the same span.&lt;/figcaption&gt;
&lt;details class="kii-figure__table"&gt;&lt;summary&gt;View as table&lt;/summary&gt;
&lt;table&gt;
 &lt;thead&gt;
 &lt;tr&gt;
 &lt;th&gt;GPU&lt;/th&gt;
 &lt;th&gt;HBM capacity&lt;/th&gt;
 &lt;th&gt;Bandwidth&lt;/th&gt;
 &lt;/tr&gt;
 &lt;/thead&gt;
 &lt;tbody&gt;
 &lt;tr&gt;
 &lt;td&gt;H100&lt;/td&gt;
 &lt;td&gt;80GB&lt;/td&gt;
 &lt;td&gt;3.35TB/s&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;B200&lt;/td&gt;
 &lt;td&gt;192GB&lt;/td&gt;
 &lt;td&gt;about 8TB/s&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;GB300&lt;/td&gt;
 &lt;td&gt;288GB&lt;/td&gt;
 &lt;td&gt;about 8TB/s&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Vera Rubin&lt;/td&gt;
 &lt;td&gt;288GB&lt;/td&gt;
 &lt;td&gt;22 targeted, about 20TB/s reported&lt;/td&gt;
 &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;/details&gt;
&lt;/figure&gt;
&lt;p&gt;The H100&amp;rsquo;s HBM was 80GB at 3.35 terabytes per second. B200 moved to 192GB at 8 terabytes, and GB300 raised capacity to 288GB. But the next-generation Vera Rubin&amp;rsquo;s HBM4 &lt;strong&gt;capacity stays at the same 288GB&lt;/strong&gt;. What grew is bandwidth, targeted at 22 terabytes per second, with actual early shipments reported around 20 terabytes, roughly 2.75 times GB300. Over the same span, GPU compute performance grew about 3.5 times. [Fact: NVIDIA specifications, SemiAnalysis, reporting] The pattern of compute outpacing bandwidth keeps repeating generation after generation, and this is the substance of what is called the memory wall.&lt;/p&gt;
&lt;p&gt;This gap shows up as real cost during inference. Inference splits into two stages. The prefill stage, which processes the entire prompt at once, is compute-intensive, so GPU compute performance is the bottleneck there. But the decode stage, which generates tokens one at a time, has to reread the model weights and the entire intermediate memory from memory at every single token, so &lt;strong&gt;the bottleneck becomes memory bandwidth rather than compute&lt;/strong&gt;. During this stage the GPU sits idle for a good share of the time, waiting on memory. That is why buying more expensive chips does not raise throughput when HBM cannot keep up.&lt;/p&gt;
&lt;h2 id="3-what-agents-change-resident-memory"&gt;3. What Agents Change: Resident Memory
&lt;/h2&gt;&lt;p&gt;The second foundation is the change in workload. Unlike a chatbot that answers one question at a time, an agent searches the web, writes code, tests it, and fixes it, over stretches that run from tens of minutes to several hours. What the model maintains to remember everything up to that point is the KV cache, an intermediate memory, and this cache grows larger the longer the context, the longer the session runs, and the more agents run at once. A single agent&amp;rsquo;s task splits into multiple model calls, for routing, search, tool selection, verification, and each call holds its own cache, so the memory resident at any one time grows to a different order of magnitude than in the chatbot era.&lt;/p&gt;
&lt;p&gt;This shift already shows up in pricing. Field data puts the processing cost of a cached input token at one-tenth that of an uncached one, and operational data also shows that a single server&amp;rsquo;s DRAM ceiling (typically 1 to 2 terabytes) makes it hard to hold a cache for more than an hour. The fact that NVIDIA launched a dedicated KV cache storage tier (CMX) as a separate product line this past January is itself evidence that the problem has grown large enough to become a hardware product. [Fact: industry data, NVIDIA announcement] TrendForce says the spread of agents is shifting the CPU-to-GPU ratio in servers, with enterprise agents consuming up to four times the tokens of before, and projects the global memory market at $1.28 trillion in 2027. [Fact: TrendForce, May 29]&lt;/p&gt;
&lt;p&gt;An honest gap is worth noting too. No published quantitative estimate yet exists for exactly how many exabytes agent workloads add to HBM demand. [Blocked: no public estimate] The direction is clear, but the magnitude still belongs to narrative rather than data.&lt;/p&gt;
&lt;h2 id="4-why-only-three-suppliers-the-arithmetic-of-supply"&gt;4. Why Only Three Suppliers: The Arithmetic of Supply
&lt;/h2&gt;&lt;p&gt;The third foundation is a structure in which supply is hard to expand. Three numbers summarize that structure.&lt;/p&gt;
&lt;figure class="kii-figure"&gt;
&lt;div class="kii-figure__frame"&gt;
&lt;svg viewBox="0 0 700 300" xmlns="http://www.w3.org/2000/svg" role="img"&gt;
&lt;line x1="60" y1="236.0" x2="676.0" y2="236.0" stroke="var(--kii-chart-grid)" stroke-width="1"/&gt;
&lt;text x="51" y="240.0" fill="var(--card-text-color-tertiary)" font-size="11" text-anchor="end"&gt;0%&lt;/text&gt;
&lt;line x1="60" y1="196.8" x2="676.0" y2="196.8" stroke="var(--kii-chart-grid)" stroke-width="1"/&gt;
&lt;text x="51" y="200.8" fill="var(--card-text-color-tertiary)" font-size="11" text-anchor="end"&gt;21%&lt;/text&gt;
&lt;line x1="60" y1="157.6" x2="676.0" y2="157.6" stroke="var(--kii-chart-grid)" stroke-width="1"/&gt;
&lt;text x="51" y="161.6" fill="var(--card-text-color-tertiary)" font-size="11" text-anchor="end"&gt;43%&lt;/text&gt;
&lt;line x1="60" y1="118.4" x2="676.0" y2="118.4" stroke="var(--kii-chart-grid)" stroke-width="1"/&gt;
&lt;text x="51" y="122.4" fill="var(--card-text-color-tertiary)" font-size="11" text-anchor="end"&gt;64%&lt;/text&gt;
&lt;line x1="60" y1="79.2" x2="676.0" y2="79.2" stroke="var(--kii-chart-grid)" stroke-width="1"/&gt;
&lt;text x="51" y="83.2" fill="var(--card-text-color-tertiary)" font-size="11" text-anchor="end"&gt;85%&lt;/text&gt;
&lt;line x1="60" y1="40.0" x2="676.0" y2="40.0" stroke="var(--kii-chart-grid)" stroke-width="1"/&gt;
&lt;text x="51" y="44.0" fill="var(--card-text-color-tertiary)" font-size="11" text-anchor="end"&gt;106%&lt;/text&gt;
&lt;rect x="111.3" y="61.0" width="102.7" height="175.0" rx="4" fill="var(--kii-cat-1)"/&gt;
&lt;text x="162.7" y="53.0" fill="var(--card-text-color-main)" font-size="12.5" font-weight="600" text-anchor="middle"&gt;95%&lt;/text&gt;
&lt;text x="162.7" y="256.0" fill="var(--card-text-color-secondary)" font-size="12" text-anchor="middle"&gt;Single die&lt;/text&gt;
&lt;text x="162.7" y="274.0" fill="var(--card-text-color-tertiary)" font-size="10.5" text-anchor="middle"&gt;assume 95%&lt;/text&gt;
&lt;rect x="316.7" y="114.4" width="102.7" height="121.6" rx="4" fill="var(--kii-cat-3)"/&gt;
&lt;text x="368.0" y="106.4" fill="var(--card-text-color-main)" font-size="12.5" font-weight="600" text-anchor="middle"&gt;66%&lt;/text&gt;
&lt;text x="368.0" y="256.0" fill="var(--card-text-color-secondary)" font-size="12" text-anchor="middle"&gt;8-high stack&lt;/text&gt;
&lt;text x="368.0" y="274.0" fill="var(--card-text-color-tertiary)" font-size="10.5" text-anchor="middle"&gt;0.95 to the 8th&lt;/text&gt;
&lt;rect x="522.0" y="154.9" width="102.7" height="81.1" rx="4" fill="var(--kii-cat-4)"/&gt;
&lt;text x="573.3" y="146.9" fill="var(--card-text-color-main)" font-size="12.5" font-weight="600" text-anchor="middle"&gt;44%&lt;/text&gt;
&lt;text x="573.3" y="256.0" fill="var(--card-text-color-secondary)" font-size="12" text-anchor="middle"&gt;16-high stack&lt;/text&gt;
&lt;text x="573.3" y="274.0" fill="var(--card-text-color-tertiary)" font-size="10.5" text-anchor="middle"&gt;0.95 to the 16th&lt;/text&gt;
&lt;line x1="60" y1="236.0" x2="676.0" y2="236.0" stroke="var(--kii-chart-axis)" stroke-width="1"/&gt;
&lt;text x="60" y="24" fill="var(--card-text-color-tertiary)" font-size="11.5"&gt;Cumulative yield (%). Illustrative maths at 95% per-die yield&lt;/text&gt;
&lt;/svg&gt;
&lt;/div&gt;
&lt;figcaption&gt;&lt;strong&gt;The arithmetic of stacking.&lt;/strong&gt; A 95% die yield becomes 66% at 8-high and 44% at 16-high, and one bad via among more than 8,000 kills the stack. This arithmetic is why only three firms produce at scale and why HBM eats three times the wafers per bit.&lt;/figcaption&gt;
&lt;details class="kii-figure__table"&gt;&lt;summary&gt;View as table&lt;/summary&gt;
&lt;table&gt;
 &lt;thead&gt;
 &lt;tr&gt;
 &lt;th&gt;Configuration&lt;/th&gt;
 &lt;th&gt;Cumulative yield&lt;/th&gt;
 &lt;th&gt;Calculation&lt;/th&gt;
 &lt;/tr&gt;
 &lt;/thead&gt;
 &lt;tbody&gt;
 &lt;tr&gt;
 &lt;td&gt;Single die&lt;/td&gt;
 &lt;td&gt;95%&lt;/td&gt;
 &lt;td&gt;assumption&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;8-high stack&lt;/td&gt;
 &lt;td&gt;about 66%&lt;/td&gt;
 &lt;td&gt;0.95 to the 8th power&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;16-high stack&lt;/td&gt;
 &lt;td&gt;about 44%&lt;/td&gt;
 &lt;td&gt;0.95 to the 16th power&lt;/td&gt;
 &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;/details&gt;
&lt;/figure&gt;
&lt;p&gt;To produce the same capacity, HBM consumes roughly three times the wafers of ordinary server DRAM. This is because the dies are large, the through-silicon vias (TSVs) needed for stacking take up area, and yield losses compound. The arithmetic of yield is unforgiving. Even if a single die&amp;rsquo;s yield is 95%, cumulative yield falls to 66% at 8 layers of stacking and 44% at 16 layers. A single bad via, out of the more than 8,000 through-silicon vias on one die, is enough to kill the entire stack. [Fact: industry technical data] This barrier is what turned HBM into a three-way oligopoly, and Omdia concluded that HBM is far more complex than standard DRAM, that only three companies can produce it at volume, and that the bottleneck will last at least through 2027. [Fact: Omdia, July 30]&lt;/p&gt;
&lt;p&gt;A time barrier sits on top of that. A fab takes three to five years from groundbreaking to volume production, and lithography equipment alone has a lead time of a year to a year and a half. Qualification eats time too. It took Samsung Electronics&amp;rsquo;s 12-layer HBM3E about 18 months to pass NVIDIA&amp;rsquo;s qualification, and the Blackwell cycle had already passed by the time it did. [Fact: industry data] That is why capacity now under construction (SK Hynix&amp;rsquo;s Cheongju M15X in mid-2027, Samsung&amp;rsquo;s Pyeongtaek P5 in 2028) cannot contribute meaningful volume before 2028 at the earliest. TrendForce&amp;rsquo;s supply-demand outlook points the same way. The DRAM fulfillment rate is -1% to -2% in 2026, the gap widens further in 2027, and meaningful new supply does not arrive until 2028. [Fact: TrendForce]&lt;/p&gt;
&lt;p&gt;The result is what suppliers are saying right now. Micron&amp;rsquo;s chief business officer said on August 10 that the company cannot fill even half of data center customer demand and that &lt;strong&gt;&amp;ldquo;2027 will be tighter than 2026&amp;rdquo;&lt;/strong&gt;, and SK Hynix CEO Kwak Noh-jung described 2027, from a supply standpoint, as the tightest year in the industry&amp;rsquo;s history. [Fact: TrendForce, press reporting] HBM&amp;rsquo;s share of DRAM wafer input rises from 18% (end of 2025) to 30% (end of 2027), but on a bit basis that is only 8% to 13%. The three-times-the-wafers structure shows up here as well. [Fact: TrendForce]&lt;/p&gt;
&lt;p&gt;That covers the core of the bottleneck claim: the gap in growth rates, resident memory, and the arithmetic of supply. All three are measured facts, and the tightness through 2027 looks less like a forecast and more like physics.&lt;/p&gt;
&lt;h2 id="5-the-workarounds-the-bottleneck-invites"&gt;5. The Workarounds the Bottleneck Invites
&lt;/h2&gt;&lt;p&gt;But this claim has another half. The deeper the bottleneck grows, the bigger the reward for designs that route around it. Over the past month, evidence of that workaround has been piling up.&lt;/p&gt;
&lt;figure class="kii-figure"&gt;
&lt;div class="kii-figure__frame"&gt;
&lt;svg viewBox="0 0 700 326" xmlns="http://www.w3.org/2000/svg" role="img"&gt;
&lt;line x1="132" y1="20" x2="132" y2="282" stroke="var(--kii-chart-axis)" stroke-width="1.5"/&gt;
&lt;text x="116" y="30" fill="var(--card-text-color-secondary)" font-size="12" font-weight="600" text-anchor="end"&gt;January&lt;/text&gt;
&lt;circle cx="132" cy="26" r="6" fill="var(--kii-cat-1)"/&gt;
&lt;circle cx="132" cy="26" r="9" fill="none" stroke="var(--kii-cat-1)" stroke-opacity="0.32" stroke-width="2"/&gt;
&lt;text x="152" y="30" fill="var(--card-text-color-main)" font-size="13" font-weight="600"&gt;NVIDIA announces CMX&lt;/text&gt;
&lt;text x="152" y="48" fill="var(--card-text-color-tertiary)" font-size="11.5"&gt;a dedicated KV-cache storage tier became a product&lt;/text&gt;
&lt;text x="116" y="84" fill="var(--card-text-color-secondary)" font-size="12" font-weight="600" text-anchor="end"&gt;Roadmap&lt;/text&gt;
&lt;circle cx="132" cy="80" r="6" fill="var(--kii-cat-1)"/&gt;
&lt;circle cx="132" cy="80" r="9" fill="none" stroke="var(--kii-cat-1)" stroke-opacity="0.32" stroke-width="2"/&gt;
&lt;text x="152" y="84" fill="var(--card-text-color-main)" font-size="13" font-weight="600"&gt;Rubin CPX uses GDDR7&lt;/text&gt;
&lt;text x="152" y="102" fill="var(--card-text-color-tertiary)" font-size="11.5"&gt;the prefill chip is designed without HBM&lt;/text&gt;
&lt;text x="116" y="138" fill="var(--card-text-color-secondary)" font-size="12" font-weight="600" text-anchor="end"&gt;28 July&lt;/text&gt;
&lt;circle cx="132" cy="134" r="6" fill="var(--kii-cat-3)"/&gt;
&lt;circle cx="132" cy="134" r="9" fill="none" stroke="var(--kii-cat-3)" stroke-opacity="0.32" stroke-width="2"/&gt;
&lt;text x="152" y="138" fill="var(--card-text-color-main)" font-size="13" font-weight="600"&gt;SOCAMM halving reported&lt;/text&gt;
&lt;text x="152" y="156" fill="var(--card-text-color-tertiary)" font-size="11.5"&gt;modules cut in half as memory neared 29% of system cost&lt;/text&gt;
&lt;text x="116" y="192" fill="var(--card-text-color-secondary)" font-size="12" font-weight="600" text-anchor="end"&gt;3 August&lt;/text&gt;
&lt;circle cx="132" cy="188" r="6" fill="var(--kii-cat-3)"/&gt;
&lt;circle cx="132" cy="188" r="9" fill="none" stroke="var(--kii-cat-3)" stroke-opacity="0.32" stroke-width="2"/&gt;
&lt;text x="152" y="192" fill="var(--card-text-color-main)" font-size="13" font-weight="600"&gt;High Bandwidth Flash spec released&lt;/text&gt;
&lt;text x="152" y="210" fill="var(--card-text-color-tertiary)" font-size="11.5"&gt;SanDisk and SK Hynix, with Google in the consortium&lt;/text&gt;
&lt;text x="116" y="246" fill="var(--card-text-color-secondary)" font-size="12" font-weight="600" text-anchor="end"&gt;Early August&lt;/text&gt;
&lt;circle cx="132" cy="242" r="6" fill="var(--kii-cat-4)"/&gt;
&lt;circle cx="132" cy="242" r="9" fill="none" stroke="var(--kii-cat-4)" stroke-opacity="0.32" stroke-width="2"/&gt;
&lt;text x="152" y="246" fill="var(--card-text-color-main)" font-size="13" font-weight="600"&gt;Rubin Ultra reported at 192GB&lt;/text&gt;
&lt;text x="152" y="264" fill="var(--card-text-color-tertiary)" font-size="11.5"&gt;a flagship whose HBM would shrink below its predecessor&lt;/text&gt;
&lt;/svg&gt;
&lt;/div&gt;
&lt;figcaption&gt;&lt;strong&gt;The workarounds accumulate.&lt;/strong&gt; Five confirmed this year: a dedicated cache tier, a prefill chip without HBM, module capacity halved, a flash-based capacity tier, and a flagship whose HBM may shrink. Design changes induced by price do not reverse when price falls.&lt;/figcaption&gt;
&lt;details class="kii-figure__table"&gt;&lt;summary&gt;View as table&lt;/summary&gt;
&lt;table&gt;
 &lt;thead&gt;
 &lt;tr&gt;
 &lt;th&gt;When&lt;/th&gt;
 &lt;th&gt;Event&lt;/th&gt;
 &lt;th&gt;Nature&lt;/th&gt;
 &lt;/tr&gt;
 &lt;/thead&gt;
 &lt;tbody&gt;
 &lt;tr&gt;
 &lt;td&gt;January&lt;/td&gt;
 &lt;td&gt;NVIDIA announces CMX&lt;/td&gt;
 &lt;td&gt;KV cache moved to a tier outside HBM&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Roadmap&lt;/td&gt;
 &lt;td&gt;Rubin CPX adopts GDDR7&lt;/td&gt;
 &lt;td&gt;HBM removed from prefill&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;28 July&lt;/td&gt;
 &lt;td&gt;SOCAMM halving reported&lt;/td&gt;
 &lt;td&gt;CPU-side memory cut&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;3 August&lt;/td&gt;
 &lt;td&gt;High Bandwidth Flash spec&lt;/td&gt;
 &lt;td&gt;a capacity tier below HBM&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Early August&lt;/td&gt;
 &lt;td&gt;Rubin Ultra at 192GB reported&lt;/td&gt;
 &lt;td&gt;flagship HBM under review&lt;/td&gt;
 &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;/details&gt;
&lt;/figure&gt;
&lt;p&gt;The most significant signal comes from NVIDIA itself. After memory cost climbed to 29% of Vera Rubin system cost, above the company&amp;rsquo;s preferred 20% ceiling, reports emerged that the CPU-side memory module (SOCAMM) would be cut in half, from 192GB to 96GB. CPU-side memory in a single rack falls from 55 terabytes to 28 terabytes. Supply allocation is part of the reason too: even combining all three suppliers, NVIDIA can only receive about 60% of estimated demand, so lowering the specification lets the company ship more systems from the same allocation. [Fact: TrendForce, July 28] Rubin Ultra went further. The specification originally announced as 1 terabyte of HBM4E has stepped down in stages, and reports now say the flagship configuration will be &lt;strong&gt;set at 192GB of HBM4, below even the current Rubin&amp;rsquo;s 288GB&lt;/strong&gt;. Memory content falling as the generation rises is a direction never before seen in a flagship. [Fact: reporting, unconfirmed by NVIDIA]&lt;/p&gt;
&lt;p&gt;The workaround is not limited to cutting specifications. Rubin CPX, designed solely for the prefill stage, skips HBM entirely and uses GDDR7 instead. The design splits prefill, a compute bottleneck, from decode, a memory bottleneck, at the hardware level, so it does not pay HBM&amp;rsquo;s price where HBM is not needed. High Bandwidth Flash (HBF), for which SanDisk and SK Hynix published a standard specification in early August, stacks NAND to deliver up to 512GB of capacity with bandwidth at the low end of the HBM class, and its consortium includes Google. It lays a cheaper capacity tier beneath HBM. On the software side, techniques that compress the KV cache itself have entered commercial models. The attention architecture used by the DeepSeek family cuts cache per token to one-third to one-fifth of the conventional approach. It is still concentrated in Chinese-origin models, but the direction is clear. [Fact: company announcements, technical literature]&lt;/p&gt;
&lt;p&gt;The supplier&amp;rsquo;s own words confirm this list. SK Hynix Chairman Chey Tae-won, in a CNBC interview last week, likened demand to a war but also said &lt;strong&gt;prices had risen too fast&lt;/strong&gt;. The reason the largest beneficiary of the bottleneck worries about how deep it runs is exactly the list above. Once a workaround triggered by price gets designed in, it does not reverse even if prices come back down.&lt;/p&gt;
&lt;h2 id="6-where-the-bottlenecks-rent-goes"&gt;6. Where the Bottleneck&amp;rsquo;s Rent Goes
&lt;/h2&gt;&lt;p&gt;If the bottleneck is real and the workarounds keep growing, the question left standing is who captures the bottleneck&amp;rsquo;s profit, and in what form. The answer is increasingly the contract.&lt;/p&gt;
&lt;p&gt;On the demand side, multi-year contracts lock in volume. Micron has tied up roughly half its revenue through contracts with 16 strategic customers running to around 2030, and reports keep describing the three suppliers&amp;rsquo; 2027 capacity as effectively fully allocated already. &amp;ldquo;Several-fold&amp;rdquo; increases in 2027 HBM contract prices are being discussed, though TrendForce&amp;rsquo;s own wording is a qualitative description that does not specify a multiple, so reading more precision into it than that would be an overreach. [Fact: company data, TrendForce] On the pricing side, customization is changing the structure. Starting with HBM4, the base die moves to TSMC logic processes (mainly 12nm, with premium 5nm and next-generation 3nm), and as NVIDIA and Google each demand their own custom specifications, HBM is turning from a general-purpose component into a part incompatible across customers. A custom base die makes it hard for a customer to switch suppliers, while also locking the supplier to that customer. [Fact: industry data]&lt;/p&gt;
&lt;p&gt;What this structure means matches the conclusion this series has been building toward. The bottleneck&amp;rsquo;s rent is being &lt;strong&gt;fixed by the terms of multi-year contracts&lt;/strong&gt; rather than a spike in spot prices. A fixed rent trims the upside in a boom in exchange for limiting the downside in a bust. The wide spread across institutions in HBM market size forecasts for 2027, from $116 billion to as much as 260 trillion won, is itself a feature of this transition, which makes the share of contract coverage a better thing to watch than any single figure. [Fact: various institutions, wide dispersion]&lt;/p&gt;
&lt;h2 id="7-the-implication-for-korea-2027s-physics-2028s-calendar"&gt;7. The Implication for Korea: 2027&amp;rsquo;s Physics, 2028&amp;rsquo;s Calendar
&lt;/h2&gt;&lt;p&gt;Here is what this verification means for Korean memory makers.&lt;/p&gt;
&lt;p&gt;Ownership of the bottleneck is Korea&amp;rsquo;s position. Two of the three companies capable of volume production are Korean, and SK Hynix&amp;rsquo;s HBM share runs from the mid-50s to high-50s percent depending on the source. Some tallies put SK Hynix at roughly 70% of NVIDIA&amp;rsquo;s next-platform initial HBM4 volume. [Fact: reporting, share varies by source] Profit visibility through 2027 is backed by the physics laid out above. Because of lead times, 2027 supply is already determined at this very moment, and a substantial share of that volume is already sold under contract.&lt;/p&gt;
&lt;p&gt;What remains open is 2028. New fabs begin contributing volume that year (Cheongju M15X, Pyeongtaek P5). CXMT&amp;rsquo;s HBM capacity, if it stays on plan, also climbs to a scale of 100,000 wafers that year, though still carrying a roughly 25% yield and a customer base limited to domestic Chinese buyers. That is also when the effects of the workaround technologies described above accumulate enough to matter. [Fact: company data, SemiAnalysis] The bottleneck claim is a matter of physics through 2027, but from 2028 on it reverts to a question of supply discipline and contract terms. The conclusion this series reached in early August, that the investment question is not whether prices rise further but whether they hold, and that holding is most solid when a contract guarantees it, survives intact even after passing through the physics of the bottleneck.&lt;/p&gt;
&lt;p&gt;One last sentence, carried over directly from the Morgan Stanley report: &amp;ldquo;Infrastructure comes first. The applications that will justify that build-out are not yet visible.&amp;rdquo; Even the strongest supporter of the bottleneck claim attached this qualifier. A bottleneck produces rent only when demand is real. This earnings season showed demand measured in fact, and its persistence gets re-verified every quarter.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;The stocks mentioned in this piece are illustrative examples for analysis and do not constitute a recommendation to buy or sell any particular stock. Responsibility for investment decisions and their outcomes rests with the investor. The figure of roughly 50 billion Gb in 2027 HBM demand is an estimate from Kiwoom Securities and Korean research aggregation, not a figure from the Morgan Stanley document, and the back-calculated share of total DRAM is our own computation. GPU specifications and bandwidth combine announced targets with reported figures and may differ from actual shipped specifications. Yield figures are example calculations drawn from industry technical data, and actual yields are not disclosed by company. The Rubin Ultra specification adjustment and the SOCAMM reduction come from research firms and press reporting and have not been confirmed by NVIDIA. Forecasts of HBM market size vary widely across institutions. The contribution of agent workloads to HBM demand has no published quantitative estimate, so only the direction is described. Prices and outlooks are as of mid-August 2026.&lt;/p&gt;
&lt;h3 id="related-posts"&gt;Related Posts
&lt;/h3&gt;&lt;ul&gt;
&lt;li&gt;&lt;a class="link" href="https://koreainvestinsights.com/post/weekly-wrap-wallst-financing-reaction-shift-2026-08-14/" &gt;The Second Week of August in Review: AI Capital Moves to Wall Street, and Good News Starts Lasting Two Days&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a class="link" href="https://koreainvestinsights.com/post/memory-price-deceleration-p-holds-thesis-2026-08-03/" &gt;Memory No Longer Needs Prices to Rise Further: Anatomy of the August 3 Plunge and the Contract Price Deceleration&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a class="link" href="https://koreainvestinsights.com/post/hyperscaler-proof-memory-selling-three-hypotheses-2026-08-04/" &gt;What Hyperscaler Earnings Proved, and What They Didn&amp;rsquo;t: Three Explanations for the Memory Sell-Off&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a class="link" href="https://koreainvestinsights.com/post/july-2026-earnings-two-listings-kimi-four-worries-2026-07-31/" &gt;July Earnings Season Wrap: AI Demand Was Confirmed, and Memory Pricing Became an Industry-Wide Cost&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;</description></item></channel></rss>