<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Astra on Korea Invest Insights</title><link>https://koreainvestinsights.com/tags/astra/</link><description>Recent content in Astra on Korea Invest Insights</description><generator>Hugo -- gohugo.io</generator><language>en</language><copyright>koreainvestinsights.com · @korea_invest_insights</copyright><lastBuildDate>Fri, 11 Sep 2026 12:36:32 +0900</lastBuildDate><atom:link href="https://koreainvestinsights.com/tags/astra/feed.xml" rel="self" type="application/rss+xml"/><item><title>Office Work Could Drive the Next Token Wave: Translating Astra's Gains into Korean Equity Earnings</title><link>https://koreainvestinsights.com/post/astra-office-agent-token-demand-korean-equities-2026-09-11/</link><pubDate>Fri, 11 Sep 2026 12:00:00 +0900</pubDate><guid>https://koreainvestinsights.com/post/astra-office-agent-token-demand-korean-equities-2026-09-11/</guid><description>&lt;p&gt;GPT-6 Astra, introduced on September 3, and the financial-work and data-analysis products announced on September 10 suggest a change in the next competitive question for AI. It is moving from the quality of an answer toward how reliably a system can stay with a delegated assignment until it is finished. OpenAI&amp;rsquo;s financial product connects data, sources and company document templates; its data product connects enterprise information and permissions.&lt;sup id="fnref:1"&gt;&lt;a href="#fn:1" class="footnote-ref" role="doc-noteref"&gt;1&lt;/a&gt;&lt;/sup&gt;&lt;sup id="fnref:2"&gt;&lt;a href="#fn:2" class="footnote-ref" role="doc-noteref"&gt;2&lt;/a&gt;&lt;/sup&gt;&lt;sup id="fnref:3"&gt;&lt;a href="#fn:3" class="footnote-ref" role="doc-noteref"&gt;3&lt;/a&gt;&lt;/sup&gt;&lt;/p&gt;
&lt;p&gt;This could extend intensive usage associated with developers into office work more broadly. A launch, however, is not evidence of a token-demand explosion. Multiplying the office workforce by a developer&amp;rsquo;s token consumption is not a defensible market forecast.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The thesis is conditional: better computer use and long-horizon execution can reduce human intervention, increasing the number of tasks people actually delegate. Korean equity investors must then identify which companies convert that demand into memory shipments and pricing, electrical-equipment orders, or recurring software profit.&lt;/strong&gt;&lt;/p&gt;
&lt;h2 id="coding-offered-an-environment-for-execution-and-verification"&gt;Coding offered an environment for execution and verification
&lt;/h2&gt;&lt;p&gt;Software work provides code repositories to read, tools to execute and tests against which to check results. A system can modify a program, examine what failed and try again. Not every software outcome is easy to verify. The narrower inference is that connecting execution with evaluation was often easier to organize in development than in fragmented office workflows.&lt;/p&gt;
&lt;p&gt;An April 2026 study of coding agents found substantial variation in token usage across models and execution paths, including repeated runs on the same task. Spending more did not consistently improve accuracy. This is another reason not to extrapolate developers&amp;rsquo; usage mechanically to office workers.&lt;sup id="fnref:4"&gt;&lt;a href="#fn:4" class="footnote-ref" role="doc-noteref"&gt;4&lt;/a&gt;&lt;/sup&gt;&lt;/p&gt;
&lt;p&gt;Office work faces different frictions. A promise in an email, a customer record, an Excel calculation and an internal policy may reside in separate systems. A correct answer does not compensate for an incorrectly updated record. Permissions and approvals must be connected before an answer becomes completed work.&lt;/p&gt;
&lt;p&gt;“Office work after coding” therefore does not mean one market opens only after the other finishes. The markets already overlap. The claim is that the intensity of repeat usage in non-developer roles could become the next source of growth.&lt;/p&gt;
&lt;h2 id="computer-use-expands-reach-long-horizon-capability-changes-delegation-economics"&gt;Computer use expands reach; long-horizon capability changes delegation economics
&lt;/h2&gt;&lt;p&gt;Computer use means reading screens, clicking controls and entering information. Long-horizon capability means preserving objectives and evidence across steps, recovering from failures and checking the result. The former broadens accessible workflows; the latter reduces the need for a person to supervise continuously.&lt;/p&gt;
&lt;p&gt;The two can reinforce each other. Reliable clicking is insufficient if the agent loses the objective at step ten. Extended reasoning is insufficient if a person must transfer every result into the operating system of the business. This is an economic interpretation, not a verified mathematical property of the model architecture.&lt;/p&gt;
&lt;p&gt;Published evaluations should be separated from evidence of production adoption.&lt;/p&gt;
&lt;table&gt;
 &lt;thead&gt;
 &lt;tr&gt;
 &lt;th&gt;Published measure&lt;/th&gt;
 &lt;th style="text-align: right"&gt;Astra&lt;/th&gt;
 &lt;th style="text-align: right"&gt;GPT-5.6 Sol&lt;/th&gt;
 &lt;th&gt;Appropriate interpretation&lt;/th&gt;
 &lt;/tr&gt;
 &lt;/thead&gt;
 &lt;tbody&gt;
 &lt;tr&gt;
 &lt;td&gt;AutomationBench&lt;/td&gt;
 &lt;td style="text-align: right"&gt;41.4%&lt;/td&gt;
 &lt;td style="text-align: right"&gt;18.1%&lt;/td&gt;
 &lt;td&gt;Improvement on an automation evaluation, not the share of all office work automated&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;OSWorld 2.0 partial score&lt;/td&gt;
 &lt;td style="text-align: right"&gt;72.6%&lt;/td&gt;
 &lt;td style="text-align: right"&gt;65.7%&lt;/td&gt;
 &lt;td&gt;Partial credit on an offline subset, not fully autonomous completion probability&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;OSWorld latency simulation&lt;/td&gt;
 &lt;td style="text-align: right"&gt;40 minutes&lt;/td&gt;
 &lt;td style="text-align: right"&gt;75 minutes&lt;/td&gt;
 &lt;td&gt;Time under particular evaluation conditions, not measured labor savings in companies&lt;/td&gt;
 &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;Source: OpenAI. OSWorld uses the v2026.08.08 offline subset. Evaluation settings and tool environments can differ from production; improvements in the execution environment also matter.&lt;sup id="fnref1:1"&gt;&lt;a href="#fn:1" class="footnote-ref" role="doc-noteref"&gt;1&lt;/a&gt;&lt;/sup&gt;&lt;/p&gt;
&lt;p&gt;Small improvements in step reliability can produce larger changes in the feasibility of delegation. In an illustrative workflow requiring all 50 independent steps to succeed, 98% reliability per step yields 0.98 to the power of 50, or 36.4%. At 99.5%, the corresponding figure is 77.8%. Real failures are correlated, and checkpoints, retries and recovery change the result. These numbers illustrate compounding; they are not estimates of Astra&amp;rsquo;s real-world success rate.&lt;/p&gt;
&lt;p&gt;Nor should a company implement every task through screen clicks. A stable application programming interface, or API, should generally be preferred when available, with computer use bridging unsupported gaps. The opportunity is to connect a workflow without replacing every legacy system.&lt;/p&gt;
&lt;h2 id="capability-is-being-packaged-with-data-access-and-operating-infrastructure"&gt;Capability is being packaged with data access and operating infrastructure
&lt;/h2&gt;&lt;p&gt;ChatGPT Financial Services, announced September 10, attempts to connect financial data and document work with business templates and controls. The data-analysis offering connects approved data with business context. Both demonstrate why a capable model alone is insufficient: access and validation must accompany it.&lt;sup id="fnref1:2"&gt;&lt;a href="#fn:2" class="footnote-ref" role="doc-noteref"&gt;2&lt;/a&gt;&lt;/sup&gt;&lt;sup id="fnref1:3"&gt;&lt;a href="#fn:3" class="footnote-ref" role="doc-noteref"&gt;3&lt;/a&gt;&lt;/sup&gt;&lt;/p&gt;
&lt;p&gt;The Agents API public beta offers infrastructure for context management, tool execution and coordination between agents. Reducing the need for every enterprise to build durable execution infrastructure could lower the cost of turning model improvements into paid workflows. The launch and described features are observable; aggregate customer spending and deployment returns remain separate questions.&lt;sup id="fnref:5"&gt;&lt;a href="#fn:5" class="footnote-ref" role="doc-noteref"&gt;5&lt;/a&gt;&lt;/sup&gt;&lt;/p&gt;
&lt;p&gt;The likely initial opportunity is not the unsupervised replacement of every office role. Bounded, inspectable activities such as document comparisons, sales preparation, recurring reports and expense reconciliation are more plausible candidates for habitual delegation. Payments, binding contracts and employment decisions should retain appropriate human authorization because the cost of an error is much higher.&lt;/p&gt;
&lt;h2 id="a-short-instruction-can-initiate-a-long-chain-of-computation"&gt;A short instruction can initiate a long chain of computation
&lt;/h2&gt;&lt;p&gt;“Compare three suppliers and prepare the next meeting pack” is a short request. Carrying it out may require finding emails and quotations, normalizing commercial terms, building a table, verifying figures and sources, and formatting the result. Missing information can trigger another pass through part of the process.&lt;/p&gt;
&lt;p&gt;The final answer need not be long for the model to be called many times. Retrieved material, tool responses, intermediate checks and revisions consume tokens. Tokens are units of model input and output, not counts of user questions or sentences.&lt;/p&gt;
&lt;p&gt;Anthropic reported in June 2025 that, in its own research-system data, agents typically used around four times the tokens of chat interactions and multi-agent systems around fifteen times. This was a comparison within a particular company&amp;rsquo;s systems, not a universal multiplier for business work. It nevertheless provides evidence for the mechanism through which conversation can become substantially more intensive execution.&lt;sup id="fnref:6"&gt;&lt;a href="#fn:6" class="footnote-ref" role="doc-noteref"&gt;6&lt;/a&gt;&lt;/sup&gt;&lt;/p&gt;
&lt;p&gt;The durable opportunity is not inefficiently spending longer on the same task. It is undertaking work previously omitted because preparation was too costly: more customer-specific proposals, more frequent inventory reconciliations or broader competitor research. Retry loops that merely inflate consumption do not establish sustainable willingness to pay.&lt;/p&gt;
&lt;h2 id="delegation-frequency-matters-more-than-a-headline-workforce-number"&gt;Delegation frequency matters more than a headline workforce number
&lt;/h2&gt;&lt;p&gt;A useful decomposition is:&lt;/p&gt;

 &lt;blockquote&gt;
 &lt;p&gt;Monthly tokens = eligible workers × active-use share × delegated tasks per active user per day × working days × cumulative tokens per task.&lt;/p&gt;

 &lt;/blockquote&gt;
&lt;p&gt;Cumulative tokens per task include all input and output across model calls, retries and verification. Calls made by additional agents belong in this total. Multiplying again by the number of agents or a retry factor would double-count the workload.&lt;/p&gt;
&lt;p&gt;The following sensitivity analysis keeps the organization at 1,000 employees and the month at 22 working days. These are assumptions, not observations from a customer or forecasts for 2027.&lt;/p&gt;
&lt;table&gt;
 &lt;thead&gt;
 &lt;tr&gt;
 &lt;th&gt;Alternative operating state&lt;/th&gt;
 &lt;th style="text-align: right"&gt;Active-use share&lt;/th&gt;
 &lt;th style="text-align: right"&gt;Daily activity&lt;/th&gt;
 &lt;th style="text-align: right"&gt;Cumulative tokens per interaction/task&lt;/th&gt;
 &lt;th style="text-align: right"&gt;Monthly total&lt;/th&gt;
 &lt;th style="text-align: right"&gt;Relative to chat baseline&lt;/th&gt;
 &lt;/tr&gt;
 &lt;/thead&gt;
 &lt;tbody&gt;
 &lt;tr&gt;
 &lt;td&gt;Chat baseline&lt;/td&gt;
 &lt;td style="text-align: right"&gt;20%&lt;/td&gt;
 &lt;td style="text-align: right"&gt;10 questions&lt;/td&gt;
 &lt;td style="text-align: right"&gt;2,000&lt;/td&gt;
 &lt;td style="text-align: right"&gt;88 million&lt;/td&gt;
 &lt;td style="text-align: right"&gt;1.0x&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Limited delegation&lt;/td&gt;
 &lt;td style="text-align: right"&gt;30%&lt;/td&gt;
 &lt;td style="text-align: right"&gt;1 task&lt;/td&gt;
 &lt;td style="text-align: right"&gt;20,000&lt;/td&gt;
 &lt;td style="text-align: right"&gt;132 million&lt;/td&gt;
 &lt;td style="text-align: right"&gt;1.5x&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Routine delegation&lt;/td&gt;
 &lt;td style="text-align: right"&gt;50%&lt;/td&gt;
 &lt;td style="text-align: right"&gt;3 tasks&lt;/td&gt;
 &lt;td style="text-align: right"&gt;40,000&lt;/td&gt;
 &lt;td style="text-align: right"&gt;1.32 billion&lt;/td&gt;
 &lt;td style="text-align: right"&gt;15.0x&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Broad delegation&lt;/td&gt;
 &lt;td style="text-align: right"&gt;70%&lt;/td&gt;
 &lt;td style="text-align: right"&gt;5 tasks&lt;/td&gt;
 &lt;td style="text-align: right"&gt;80,000&lt;/td&gt;
 &lt;td style="text-align: right"&gt;6.16 billion&lt;/td&gt;
 &lt;td style="text-align: right"&gt;70.0x&lt;/td&gt;
 &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;The baseline calculation is 1,000 × 20% × 10 × 22 × 2,000 = 88,000,000 tokens. Routine delegation is 1,000 × 50% × 3 × 22 × 40,000 = 1,320,000,000. The rows are alternative operating states, not increments to add to the baseline. A chat interaction and a completed assignment also do not represent identical output.&lt;/p&gt;
&lt;p&gt;The 15x result does not require fifteen times as many users. It combines 2.5 times as many active users with six times as many daily tokens per active user, rising from 20,000 to 120,000. Conversely, adoption that stops at one daily summary would generate a much smaller increase.&lt;/p&gt;
&lt;p&gt;A larger non-developer population alone cannot establish when office demand overtakes coding demand. Adoption, delegation frequency and tokens per task must be sufficient. Classification also needs care: code created by a business user and documentation produced by a developer should not be counted twice.&lt;/p&gt;
&lt;h2 id="a-056-model-cost-example-illustrates-the-economic-threshold"&gt;A $0.56 model-cost example illustrates the economic threshold
&lt;/h2&gt;&lt;p&gt;Astra Standard API pricing is $10 per million input tokens and $50 per million output tokens. Consider an uncached task using 36,000 input tokens and 4,000 output tokens.&lt;sup id="fnref2:1"&gt;&lt;a href="#fn:1" class="footnote-ref" role="doc-noteref"&gt;1&lt;/a&gt;&lt;/sup&gt;&lt;/p&gt;

 &lt;blockquote&gt;
 &lt;p&gt;Model cost = 36,000 / 1,000,000 × $10 + 4,000 / 1,000,000 × $50 = $0.56.&lt;/p&gt;

 &lt;/blockquote&gt;
&lt;p&gt;The 4,000 output tokens are the billable total across all calls, not the length of the final report. Where reasoning tokens are billed as output, they must be included. Caching, cheaper-model routing and separate tool fees change the actual bill.&lt;/p&gt;
&lt;p&gt;The routine-delegation scenario produces 500 active users × 3 tasks × 22 days = 33,000 monthly tasks. At this assumed cost, the model bill would be $18,480 a month, or $36.96 per active user. Browser execution, data licensing, integration, security, training and incident response are excluded.&lt;/p&gt;
&lt;p&gt;The more important equation is:&lt;/p&gt;

 &lt;blockquote&gt;
 &lt;p&gt;Net value of delegation = realized labor savings or additional commercial value − model and tool costs − human review and rework − error losses − allocated operating and integration costs.&lt;/p&gt;

 &lt;/blockquote&gt;
&lt;p&gt;Suppose a low-risk task with equivalent output takes a person 30 minutes and their time is worth $50 an hour. The original cost is $25. If AI requires five minutes of review, and the failed 20% of assignments must be redone from scratch, expected additional rework is six minutes. Human cost is about $9.17; adding $0.56 in model cost produces $9.73, leaving approximately $15.27 per task for other costs and error losses. The 80% success assumption is illustrative and is not derived from a benchmark score.&lt;/p&gt;
&lt;p&gt;If review takes 25 minutes instead, expected cost rises to $26.39. Cheap inference does not rescue the workflow. Saved time is also not automatically a cash reduction in payroll: it must translate into redeployment, higher throughput or avoided additional hiring.&lt;/p&gt;
&lt;p&gt;The decisive long-horizon metric is therefore not how long an agent can keep running. It is how much human intervention remains per completed assignment.&lt;/p&gt;
&lt;h2 id="tokens-revenue-compute-and-memory-are-different-quantities"&gt;Tokens, revenue, compute and memory are different quantities
&lt;/h2&gt;&lt;p&gt;The easiest analytical mistake is to convert 15x token growth into 15x semiconductor demand. Several transformations intervene.&lt;/p&gt;
&lt;table&gt;
 &lt;thead&gt;
 &lt;tr&gt;
 &lt;th&gt;Measure&lt;/th&gt;
 &lt;th&gt;What it captures&lt;/th&gt;
 &lt;th&gt;Why it diverges from token growth&lt;/th&gt;
 &lt;/tr&gt;
 &lt;/thead&gt;
 &lt;tbody&gt;
 &lt;tr&gt;
 &lt;td&gt;Logical tokens&lt;/td&gt;
 &lt;td&gt;Input and output across all calls&lt;/td&gt;
 &lt;td&gt;Reused input can still appear in the total&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Actual bill&lt;/td&gt;
 &lt;td&gt;Spending at input, cache and output rates&lt;/td&gt;
 &lt;td&gt;Discounts, subscriptions and cheaper models change realized pricing&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Physical computation&lt;/td&gt;
 &lt;td&gt;The calculations actually performed&lt;/td&gt;
 &lt;td&gt;Model architecture, size, reuse and efficiency differ&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Memory and storage&lt;/td&gt;
 &lt;td&gt;Capacity and bandwidth for models, working state and data&lt;/td&gt;
 &lt;td&gt;Concurrency, context length, retention and storage tier matter&lt;/td&gt;
 &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;Holding the workload mix fixed for illustration, 15x more logical tokens with an 80% reduction in computation per token implies 3x compute demand. With a 95% reduction, it implies 0.75x. The calculations are 15 × 0.20 and 15 × 0.05. These are sensitivities, not forecasts of engineering progress.&lt;/p&gt;
&lt;p&gt;Lower cost can induce more usage, but it does not automatically increase total spending. Other things equal, quantity must rise sufficiently to offset the lower unit price. New uses unlocked by affordability must be assessed alongside reductions in computation needed for existing work.&lt;/p&gt;
&lt;p&gt;Concurrency deserves separate treatment. The same daily volume can require more capacity if requests cluster at the start of the working day. Conversely, batching delay-tolerant overnight jobs can improve utilization and postpone new construction. Multiplying agent count by 24 hours does not create demand by itself.&lt;/p&gt;
&lt;h2 id="the-memory-thesis-spans-tiers-not-hbm-alone"&gt;The memory thesis spans tiers, not HBM alone
&lt;/h2&gt;&lt;p&gt;AI systems retain both model weights and intermediate results from earlier context. Reusable intermediate attention state is commonly called the KV cache. Long-running work and concurrent users can increase the importance of where this state is stored and how quickly it can be retrieved.&lt;/p&gt;
&lt;p&gt;Not all of it needs to remain continuously in expensive high-bandwidth memory, or HBM. NVIDIA&amp;rsquo;s March 2026 STX and CMX announcements describe an additional context-storage layer. The objective includes expanding accessible context while improving utilization of existing GPUs. Vendor performance figures are configuration-dependent claims, not measured efficiency gains across the entire industry.&lt;sup id="fnref:7"&gt;&lt;a href="#fn:7" class="footnote-ref" role="doc-noteref"&gt;7&lt;/a&gt;&lt;/sup&gt;&lt;/p&gt;
&lt;p&gt;HyMCache, first submitted in July and revised in August under the title A CXL Memory Rack for Multi-Turn LLM Serving, explores combining DRAM and SSD capacity for reusable state. One comparison traded some performance for substantially less DRAM. This is research rather than evidence of large-scale commercial adoption, but it is a counterexample to one-for-one extrapolation from tokens to a particular memory product.&lt;sup id="fnref:8"&gt;&lt;a href="#fn:8" class="footnote-ref" role="doc-noteref"&gt;8&lt;/a&gt;&lt;/sup&gt;&lt;/p&gt;
&lt;p&gt;The relevant Korean investment hypothesis is not that long-running agents either eliminate HBM or guarantee an explosion in it. It is that fast HBM, server DRAM and high-capacity enterprise SSDs divide the workload, and suppliers may capture value across those tiers. Tokens and stored documents are not interchangeable either: rereading the same file increases logical token volume without necessarily increasing the amount of source data stored.&lt;/p&gt;
&lt;h2 id="four-distinct-earnings-mechanisms-in-korean-listed-equities"&gt;Four distinct earnings mechanisms in Korean listed equities
&lt;/h2&gt;&lt;p&gt;This is an exposure map and research sequence, not a valuation-adjusted buy ranking.&lt;/p&gt;
&lt;table&gt;
 &lt;thead&gt;
 &lt;tr&gt;
 &lt;th&gt;Company&lt;/th&gt;
 &lt;th&gt;Exposure to office-agent demand&lt;/th&gt;
 &lt;th&gt;Evidence of conversion to profit&lt;/th&gt;
 &lt;th&gt;What weakens the thesis&lt;/th&gt;
 &lt;/tr&gt;
 &lt;/thead&gt;
 &lt;tbody&gt;
 &lt;tr&gt;
 &lt;td&gt;SK hynix (000660)&lt;/td&gt;
 &lt;td&gt;HBM, server DRAM and enterprise SSDs&lt;/td&gt;
 &lt;td&gt;Product shipments, average selling prices, premium mix and cash flow after investment&lt;/td&gt;
 &lt;td&gt;Customer efficiency offsets volume growth, or supply expansion lowers prices&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Samsung Electronics (005930)&lt;/td&gt;
 &lt;td&gt;Broad HBM, server DRAM and enterprise SSD exposure&lt;/td&gt;
 &lt;td&gt;Customer production ramps, yields, mix and memory earnings versus other divisions&lt;/td&gt;
 &lt;td&gt;Technical progress fails to become profitable volume, or losses elsewhere deepen&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;LS ELECTRIC (010120)&lt;/td&gt;
 &lt;td&gt;Data-center power distribution and electrical equipment&lt;/td&gt;
 &lt;td&gt;Signed orders, delivery schedules, backlog conversion, project margins and cash collection&lt;/td&gt;
 &lt;td&gt;Higher utilization replaces new construction, or projects are delayed or cancelled&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Samsung SDS (018260)&lt;/td&gt;
 &lt;td&gt;Enterprise data integration, agent operation and controls, workflow platforms&lt;/td&gt;
 &lt;td&gt;Paid repeat usage, renewals, earnings after model costs and external customers&lt;/td&gt;
 &lt;td&gt;Usage grows but resale costs and integration labor absorb the revenue&lt;/td&gt;
 &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;h3 id="sk-hynix-look-beyond-hbm-without-ignoring-the-cycle"&gt;SK hynix: look beyond HBM without ignoring the cycle
&lt;/h3&gt;&lt;p&gt;SK hynix&amp;rsquo;s August 26 product presentation included HBM, server DRAM and enterprise SSDs. The portfolio provides multiple ways to address a more differentiated memory hierarchy. A displayed product is not evidence that every offering contributes equal revenue on the same schedule.&lt;sup id="fnref:9"&gt;&lt;a href="#fn:9" class="footnote-ref" role="doc-noteref"&gt;9&lt;/a&gt;&lt;/sup&gt;&lt;/p&gt;
&lt;p&gt;Investors should track product-level volumes and prices rather than aggregate revenue labelled AI. Higher HBM supply does not fully describe the outcome if server DRAM or NAND pricing weakens. Operating cash flow should also be assessed after capital expenditure and working-capital requirements.&lt;/p&gt;
&lt;p&gt;SK hynix is a direct research candidate for this thesis, but direct exposure is not the same as undervaluation. Attractive industry conditions can produce poor returns when expectations already reflect the supply advantage.&lt;/p&gt;
&lt;h3 id="samsung-electronics-broad-exposure-with-offsetting-business-effects"&gt;Samsung Electronics: broad exposure, with offsetting business effects
&lt;/h3&gt;&lt;p&gt;In its July 30 second-quarter results, Samsung connected its second-half outlook for agentic AI with server DRAM, enterprise SSD and HBM demand. This establishes that a supplier is advancing a similar demand hypothesis. It remains management&amp;rsquo;s outlook rather than independent proof of future supply-demand conditions.&lt;sup id="fnref:10"&gt;&lt;a href="#fn:10" class="footnote-ref" role="doc-noteref"&gt;10&lt;/a&gt;&lt;/sup&gt;&lt;/p&gt;
&lt;p&gt;The breadth of Samsung&amp;rsquo;s memory business offers several channels of exposure, but shareholder earnings depend on profitable production after customer qualification, yields and product mix. Other semiconductor and consumer businesses also influence consolidated results.&lt;/p&gt;
&lt;p&gt;Astra&amp;rsquo;s launch should therefore not simply be added to an earnings model. Investors must identify demand beyond the AI spending already assumed, and establish whether it changes volume, pricing or profitability. A new narrative describing an existing order is not incremental earnings.&lt;/p&gt;
&lt;h3 id="ls-electric-invest-in-construction-and-orders-not-a-token-royalty"&gt;LS ELECTRIC: invest in construction and orders, not a token royalty
&lt;/h3&gt;&lt;p&gt;On August 24, LS ELECTRIC announced an expanded North American AI data-center power-equipment contract with a total value of approximately KRW 230.9 billion. That is the amended total, not an entirely additional order to add to the earlier contract. It supports the business connection to data centers, but predates Astra&amp;rsquo;s announcement and cannot be attributed to Astra.&lt;sup id="fnref:11"&gt;&lt;a href="#fn:11" class="footnote-ref" role="doc-noteref"&gt;11&lt;/a&gt;&lt;/sup&gt;&lt;/p&gt;
&lt;p&gt;If office inference merely raises utilization of existing infrastructure, new electrical-equipment orders may not increase immediately. Persistent demand must lead to financed construction with access to electricity. High financing costs and construction delays can coexist with a favorable long-term demand outlook.&lt;/p&gt;
&lt;p&gt;The checks are new orders, conversion of backlog into revenue, project profit and cash collection. Treating the company as earning a royalty every time an AI token is processed ignores both the time lag and execution risk.&lt;/p&gt;
&lt;h3 id="samsung-sds-operating-responsibility-matters-more-than-model-resale"&gt;Samsung SDS: operating responsibility matters more than model resale
&lt;/h3&gt;&lt;p&gt;Samsung SDS describes FabriX as a platform connecting multiple models with enterprise systems and supporting agent creation, operation and governance. Brity Copilot connects AI to collaboration work such as email and meetings. Data access, permissions and operational responsibility provide plausible exposure to office adoption.&lt;sup id="fnref:12"&gt;&lt;a href="#fn:12" class="footnote-ref" role="doc-noteref"&gt;12&lt;/a&gt;&lt;/sup&gt;&lt;sup id="fnref:13"&gt;&lt;a href="#fn:13" class="footnote-ref" role="doc-noteref"&gt;13&lt;/a&gt;&lt;/sup&gt;&lt;/p&gt;
&lt;p&gt;Frontier-model vendors, however, are also expanding their own enterprise products and data connections. A chat interface or resold model access may face pricing pressure. Defensibility is more likely where a provider manages legacy exceptions, permissions, audit records and incident response while earning recurring renewals.&lt;sup id="fnref2:2"&gt;&lt;a href="#fn:2" class="footnote-ref" role="doc-noteref"&gt;2&lt;/a&gt;&lt;/sup&gt;&lt;sup id="fnref1:5"&gt;&lt;a href="#fn:5" class="footnote-ref" role="doc-noteref"&gt;5&lt;/a&gt;&lt;/sup&gt;&lt;/p&gt;
&lt;p&gt;Higher paid usage does not ensure higher shareholder profit when model costs and integration labor rise faster. Investors should ask for repeat external-customer contracts, retention after implementation and margins after inference costs. The product material examined here does not establish a standalone AI-business profit margin.&lt;/p&gt;
&lt;h2 id="a-correct-industry-thesis-can-still-lose-money-in-equities"&gt;A correct industry thesis can still lose money in equities
&lt;/h2&gt;&lt;p&gt;A simple decomposition is:&lt;/p&gt;

 &lt;blockquote&gt;
 &lt;p&gt;Share-price factor = earnings-per-share factor × valuation-multiple factor.&lt;/p&gt;

 &lt;/blockquote&gt;
&lt;p&gt;If earnings per share rise 30% but the price/earnings ratio falls from 20x to 15x, the share price becomes 1.30 × 15 / 20 = 0.975 of its previous value, a 2.5% decline. These are illustrative numbers, not current valuations of any named company. The comparison assumes a consistent definition and horizon for earnings.&lt;/p&gt;
&lt;p&gt;Peak quarterly memory earnings should not be multiplied by four and treated as permanent. Evaluation requires normalized pricing, supply additions, depreciation, investment and cash flow. Equipment contract value is not immediate profit, and software revenue must be considered after resale costs.&lt;/p&gt;
&lt;p&gt;This article is not a stock-by-stock valuation report with comprehensively verified same-day prices and consensus forecasts. It does not invent target prices or entry levels. A purchase decision requires reverse-engineering the growth already embedded in enterprise value and checking whether new orders and repeat usage exceed it.&lt;/p&gt;
&lt;p&gt;Memory suppliers merit early research because investors need not identify a single winning office application to gain exposure to adoption across models. But diversification across model vendors is not diversification across the memory cycle. Owning both major Korean suppliers still leaves common exposure to AI capital expenditure and memory prices.&lt;/p&gt;
&lt;h2 id="from-q4-2026-watch-repeat-usage-and-intervention-rather-than-launches"&gt;From Q4 2026, watch repeat usage and intervention rather than launches
&lt;/h2&gt;&lt;p&gt;Q4 2026 and Q1 2027 provide a useful observation window. This is a research timetable, not a promised adoption schedule.&lt;/p&gt;
&lt;table&gt;
 &lt;thead&gt;
 &lt;tr&gt;
 &lt;th&gt;What to observe&lt;/th&gt;
 &lt;th&gt;Evidence strengthening the thesis&lt;/th&gt;
 &lt;th&gt;Evidence requiring a downgrade&lt;/th&gt;
 &lt;/tr&gt;
 &lt;/thead&gt;
 &lt;tbody&gt;
 &lt;tr&gt;
 &lt;td&gt;Paid non-developer usage&lt;/td&gt;
 &lt;td&gt;Persistent growth in repeat assignments and paid usage within the same adoption cohort&lt;/td&gt;
 &lt;td&gt;More accounts, but declining activity after trials&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Task economics&lt;/td&gt;
 &lt;td&gt;Falling review and rework time per completed assignment&lt;/td&gt;
 &lt;td&gt;Verification takes roughly as long as doing the original work&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Physical infrastructure&lt;/td&gt;
 &lt;td&gt;Compute or context-storage demand rises after caching and model routing&lt;/td&gt;
 &lt;td&gt;Logical token totals rise while physical resource use stagnates&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Korean memory suppliers&lt;/td&gt;
 &lt;td&gt;Unexpected incremental orders change product volumes, prices and earnings&lt;/td&gt;
 &lt;td&gt;Old orders are relabelled; inventories rise and prices decline&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Electrical equipment and enterprise software&lt;/td&gt;
 &lt;td&gt;Orders turn into cash; renewals and margins improve&lt;/td&gt;
 &lt;td&gt;Construction delays, low-margin resale and one-off implementation dominate&lt;/td&gt;
 &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;If repeat non-developer use and unit economics remain unproven over two quarters, the timing of the next demand wave should be pushed back. If adoption grows but efficiency offsets hardware requirements, the software-adoption thesis can survive while the semiconductor thesis becomes less powerful. Security incidents that restrict permissions, or continuing dependence on human rescue and approval, are additional falsification signals.&lt;/p&gt;
&lt;p&gt;Astra could expand the market for completed assignments more than the market for longer answers. For Korean equities, the buying case emerges when that possibility creates recurring spending and cash flow beyond expectations already embedded in the share price.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;Information checked as of September 11, 2026. Product and corporate disclosures are issuer statements, distinguished from independent production validation. Token, cost and share-price sensitivities are the author&amp;rsquo;s assumed examples. This is not personalized investment advice.&lt;/p&gt;
&lt;h3 id="sources"&gt;Sources
&lt;/h3&gt;&lt;div class="footnotes" role="doc-endnotes"&gt;
&lt;hr&gt;
&lt;ol&gt;
&lt;li id="fn:1"&gt;
&lt;p&gt;OpenAI, &lt;a class="link" href="https://openai.com/index/gpt-6-astra/" target="_blank" rel="noopener"
 &gt;GPT-6 Astra: A new generation of intelligence&lt;/a&gt;, September 2026. Evaluation conditions and Standard API prices.&amp;#160;&lt;a href="#fnref:1" class="footnote-backref" role="doc-backlink"&gt;&amp;#x21a9;&amp;#xfe0e;&lt;/a&gt;&amp;#160;&lt;a href="#fnref1:1" class="footnote-backref" role="doc-backlink"&gt;&amp;#x21a9;&amp;#xfe0e;&lt;/a&gt;&amp;#160;&lt;a href="#fnref2:1" class="footnote-backref" role="doc-backlink"&gt;&amp;#x21a9;&amp;#xfe0e;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id="fn:2"&gt;
&lt;p&gt;OpenAI, &lt;a class="link" href="https://openai.com/index/introducing-chatgpt-financial-services/" target="_blank" rel="noopener"
 &gt;Introducing ChatGPT Financial Services&lt;/a&gt;, September 10, 2026.&amp;#160;&lt;a href="#fnref:2" class="footnote-backref" role="doc-backlink"&gt;&amp;#x21a9;&amp;#xfe0e;&lt;/a&gt;&amp;#160;&lt;a href="#fnref1:2" class="footnote-backref" role="doc-backlink"&gt;&amp;#x21a9;&amp;#xfe0e;&lt;/a&gt;&amp;#160;&lt;a href="#fnref2:2" class="footnote-backref" role="doc-backlink"&gt;&amp;#x21a9;&amp;#xfe0e;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id="fn:3"&gt;
&lt;p&gt;OpenAI, &lt;a class="link" href="https://openai.com/index/put-data-to-work/" target="_blank" rel="noopener"
 &gt;Put data to work&lt;/a&gt;, September 10, 2026.&amp;#160;&lt;a href="#fnref:3" class="footnote-backref" role="doc-backlink"&gt;&amp;#x21a9;&amp;#xfe0e;&lt;/a&gt;&amp;#160;&lt;a href="#fnref1:3" class="footnote-backref" role="doc-backlink"&gt;&amp;#x21a9;&amp;#xfe0e;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id="fn:4"&gt;
&lt;p&gt;Bai et al., &lt;a class="link" href="https://arxiv.org/abs/2604.22750v2" target="_blank" rel="noopener"
 &gt;How Do AI Agents Spend Your Money? Analyzing and Predicting Token Consumption in Agentic Coding Tasks&lt;/a&gt;, first submitted April 24 and revised April 29, 2026. Abstract-level evidence, not generalized to all office work.&amp;#160;&lt;a href="#fnref:4" class="footnote-backref" role="doc-backlink"&gt;&amp;#x21a9;&amp;#xfe0e;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id="fn:5"&gt;
&lt;p&gt;OpenAI, &lt;a class="link" href="https://openai.com/index/introducing-the-agents-api/" target="_blank" rel="noopener"
 &gt;Introducing the Agents API&lt;/a&gt;, September 10, 2026.&amp;#160;&lt;a href="#fnref:5" class="footnote-backref" role="doc-backlink"&gt;&amp;#x21a9;&amp;#xfe0e;&lt;/a&gt;&amp;#160;&lt;a href="#fnref1:5" class="footnote-backref" role="doc-backlink"&gt;&amp;#x21a9;&amp;#xfe0e;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id="fn:6"&gt;
&lt;p&gt;Anthropic, &lt;a class="link" href="https://www.anthropic.com/engineering/multi-agent-research-system" target="_blank" rel="noopener"
 &gt;How we built our multi-agent research system&lt;/a&gt;, June 13, 2025.&amp;#160;&lt;a href="#fnref:6" class="footnote-backref" role="doc-backlink"&gt;&amp;#x21a9;&amp;#xfe0e;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id="fn:7"&gt;
&lt;p&gt;NVIDIA, &lt;a class="link" href="https://nvidianews.nvidia.com/news/nvidia-launches-bluefield-4-stx-storage-architecture-with-broad-industry-adoption" target="_blank" rel="noopener"
 &gt;NVIDIA Launches BlueField-4 STX Storage Architecture With Broad Industry Adoption&lt;/a&gt;, March 16, 2026.&amp;#160;&lt;a href="#fnref:7" class="footnote-backref" role="doc-backlink"&gt;&amp;#x21a9;&amp;#xfe0e;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id="fn:8"&gt;
&lt;p&gt;Jang et al., &lt;a class="link" href="https://arxiv.org/abs/2607.18141v3" target="_blank" rel="noopener"
 &gt;A CXL Memory Rack for Multi-Turn LLM Serving&lt;/a&gt;, first submitted July 20; v3 August 5, 2026. Experimental results are not commercial adoption evidence.&amp;#160;&lt;a href="#fnref:8" class="footnote-backref" role="doc-backlink"&gt;&amp;#x21a9;&amp;#xfe0e;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id="fn:9"&gt;
&lt;p&gt;SK hynix, &lt;a class="link" href="https://news.skhynix.com/en/dtf-2026/" target="_blank" rel="noopener"
 &gt;Memory Solutions at DTF 2026&lt;/a&gt;, August 26, 2026.&amp;#160;&lt;a href="#fnref:9" class="footnote-backref" role="doc-backlink"&gt;&amp;#x21a9;&amp;#xfe0e;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id="fn:10"&gt;
&lt;p&gt;Samsung Electronics, &lt;a class="link" href="https://news.samsung.com/global/samsung-electronics-announces-second-quarter-2026-results" target="_blank" rel="noopener"
 &gt;Second Quarter 2026 Results&lt;/a&gt;, July 30, 2026.&amp;#160;&lt;a href="#fnref:10" class="footnote-backref" role="doc-backlink"&gt;&amp;#x21a9;&amp;#xfe0e;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id="fn:11"&gt;
&lt;p&gt;LS ELECTRIC, &lt;a class="link" href="https://www.ls-electric.com/ko/pr/news/view/401457?b_date=&amp;amp;e_date=&amp;amp;k_type=both&amp;amp;k_word=&amp;amp;page=1&amp;amp;rowsPerPage=10&amp;amp;visiblePage=10" target="_blank" rel="noopener"
 &gt;North American AI data-center power-equipment contract announcement&lt;/a&gt;, August 24, 2026. Amended contract total, not wholly incremental order value.&amp;#160;&lt;a href="#fnref:11" class="footnote-backref" role="doc-backlink"&gt;&amp;#x21a9;&amp;#xfe0e;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id="fn:12"&gt;
&lt;p&gt;Samsung SDS, &lt;a class="link" href="https://www.samsungsds.com/kr/ai-fabrix/fabrix.html" target="_blank" rel="noopener"
 &gt;FabriX&lt;/a&gt;, accessed September 11, 2026.&amp;#160;&lt;a href="#fnref:12" class="footnote-backref" role="doc-backlink"&gt;&amp;#x21a9;&amp;#xfe0e;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id="fn:13"&gt;
&lt;p&gt;Samsung SDS, &lt;a class="link" href="https://www.samsungsds.com/kr/ai-agent/ai-agent.html" target="_blank" rel="noopener"
 &gt;Samsung SDS AI Agent&lt;/a&gt;, accessed September 11, 2026.&amp;#160;&lt;a href="#fnref:13" class="footnote-backref" role="doc-backlink"&gt;&amp;#x21a9;&amp;#xfe0e;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;/div&gt;</description></item><item><title>After Astra: How Looped Transformers Change AI Infrastructure Economics</title><link>https://koreainvestinsights.com/post/astra-loop-transformers-inference-economics-2026-09-08/</link><pubDate>Tue, 08 Sep 2026 18:00:00 +0900</pubDate><guid>https://koreainvestinsights.com/post/astra-loop-transformers-inference-economics-2026-09-08/</guid><description>&lt;p&gt;Can a small model solve harder problems by passing through the same neural network repeatedly? Lablup CEO Jeongkyu Shin revisits this question in his &lt;a class="link" href="https://www.facebook.com/jeongkyu.shin/posts/pfbid02jm4iibHsgU11P7gY8e4SZKtKYNNdtHF6wdTQK2EdyDeTMEwxiDvHnHyhyAKphLml" target="_blank" rel="noopener"
 &gt;Facebook essay on looped transformers after Astra&lt;/a&gt;. Behind the architecture question lies an infrastructure decision: how many accelerators and how much memory to buy, and how to operate them.&lt;/p&gt;
&lt;p&gt;Public research demonstrates that a model can improve problem solving while keeping its stored weights fixed, by executing a shared computational block repeatedly. But repetition consumes time and energy. A smaller model does not automatically mean a cheaper service.&lt;/p&gt;
&lt;p&gt;This is an independent analysis prompted by Shin&amp;rsquo;s essay, supplemented with original papers and model cards. We separate the essay&amp;rsquo;s interpretation of Astra from publicly established facts, and treat industry implications as conditional analysis. Sources were checked on September 8, 2026.&lt;/p&gt;
&lt;h2 id="astras-results-do-not-disclose-its-architecture"&gt;Astra&amp;rsquo;s results do not disclose its architecture
&lt;/h2&gt;&lt;p&gt;OpenAI&amp;rsquo;s September 3 announcement confirms GPT-6 Astra&amp;rsquo;s launch and capability improvements. However, the &lt;a class="link" href="https://openai.com/index/gpt-6-astra/" target="_blank" rel="noopener"
 &gt;announcement&lt;/a&gt; and &lt;a class="link" href="https://deploymentsafety.openai.com/gpt-6-astra" target="_blank" rel="noopener"
 &gt;system card&lt;/a&gt; reviewed here do not disclose a looped-transformer architecture, recursion counts, or total and active parameter counts.&lt;/p&gt;
&lt;p&gt;We therefore do not treat a connection between Astra and a Huginn-style design as an established fact. The essay&amp;rsquo;s 10T/1T size claims and AGI quotation are also excluded from the premises of this analysis. Better performance alone cannot identify the internal architecture.&lt;/p&gt;
&lt;p&gt;There is still a strong reason to examine recurrence. Public models already show attempts to vary stored parameter capacity and inference computation separately. That development can be assessed without relying on a frontier model&amp;rsquo;s undisclosed design.&lt;/p&gt;
&lt;h2 id="storing-more-and-computing-longer-are-different-choices"&gt;Storing more and computing longer are different choices
&lt;/h2&gt;&lt;p&gt;Parameters are the numerical weights adjusted during training. Enlarging a model usually increases the amount stored. Mixture of Experts, or MoE, selects some expert modules for each input, aiming to execute less computation relative to total model capacity.&lt;/p&gt;
&lt;p&gt;MoE did not originate with Switch Transformer alone. The &lt;a class="link" href="https://arxiv.org/abs/1701.06538" target="_blank" rel="noopener"
 &gt;2017 sparsely gated MoE paper&lt;/a&gt; preceded &lt;a class="link" href="https://arxiv.org/abs/2101.03961" target="_blank" rel="noopener"
 &gt;Switch Transformer in 2021&lt;/a&gt;, which simplified routing and training at scale. A larger total parameter count also does not necessarily mean more layers.&lt;/p&gt;
&lt;p&gt;Chain-of-thought (CoT) generates intermediate tokens that extend the context for later computation. A looped model feeds its internal state through a block with shared weights again. Intermediate computation need not be converted into a word at every step. These approaches can also be combined.&lt;/p&gt;
&lt;p&gt;Comparing what each approach adds clarifies the trade-off.&lt;/p&gt;
&lt;table&gt;
 &lt;thead&gt;
 &lt;tr&gt;
 &lt;th&gt;Approach&lt;/th&gt;
 &lt;th&gt;What increases&lt;/th&gt;
 &lt;th&gt;Potential cost&lt;/th&gt;
 &lt;/tr&gt;
 &lt;/thead&gt;
 &lt;tbody&gt;
 &lt;tr&gt;
 &lt;td&gt;Larger model&lt;/td&gt;
 &lt;td&gt;Weights or expert capacity&lt;/td&gt;
 &lt;td&gt;Storage, active computation, communication&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Chain-of-thought&lt;/td&gt;
 &lt;td&gt;Intermediate reasoning tokens&lt;/td&gt;
 &lt;td&gt;Generation time, context and cache&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Recurrent depth&lt;/td&gt;
 &lt;td&gt;Passes through a shared block&lt;/td&gt;
 &lt;td&gt;Repeated computation, latency, state management&lt;/td&gt;
 &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;This is a conceptual comparison. Actual economics require measurements at matched accuracy, input length and hardware conditions.&lt;/p&gt;
&lt;h2 id="research-on-thinking-before-speaking-uses-distinct-mechanisms"&gt;Research on thinking before speaking uses distinct mechanisms
&lt;/h2&gt;&lt;p&gt;&lt;a class="link" href="https://arxiv.org/abs/2310.02226" target="_blank" rel="noopener"
 &gt;Pause tokens&lt;/a&gt; provide additional computation before an answer. &lt;a class="link" href="https://arxiv.org/abs/2403.09629" target="_blank" rel="noopener"
 &gt;Quiet-STaR&lt;/a&gt; learns to generate intermediate rationales that help predict subsequent tokens. Its name should not be read as proof that it uses nonverbal continuous-state reasoning.&lt;/p&gt;
&lt;p&gt;&lt;a class="link" href="https://arxiv.org/abs/2412.06769" target="_blank" rel="noopener"
 &gt;Coconut&lt;/a&gt; feeds the final hidden state back as an input without converting it into a word. It explores retaining possibilities in an internal representation before committing to language. This is not evidence of human consciousness or continuously running autonomous thought.&lt;/p&gt;
&lt;p&gt;Recurrent-depth research includes the &lt;a class="link" href="https://arxiv.org/abs/1807.03819" target="_blank" rel="noopener"
 &gt;2018 Universal Transformer&lt;/a&gt;, which repeats a transformation and can allocate computation differently across positions. The difficulty is training useful repetition: a shared block must handle states from different stages, and another pass must improve the result. Conflicting layer roles are a useful intuition, not a universal explanation for every failure.&lt;/p&gt;
&lt;h2 id="copying-layers-differs-from-sharing-the-same-weights"&gt;Copying layers differs from sharing the same weights
&lt;/h2&gt;&lt;p&gt;Upstage&amp;rsquo;s &lt;a class="link" href="https://arxiv.org/abs/2312.15166" target="_blank" rel="noopener"
 &gt;SOLAR 10.7B&lt;/a&gt; introduced depth up-scaling, or DUS: copy existing layers, remove some, connect them into a deeper model, and continue training. Copies that begin identically can develop different weights. The resulting model stores more parameters.&lt;/p&gt;
&lt;p&gt;A recurrent model continues to share the same weights. DUS reuses prior training to build a deeper model; looping increases execution depth without a corresponding expansion in stored weights. Treating both as the same memory-saving technique gives the wrong cost model.&lt;/p&gt;
&lt;h2 id="read-huginn-and-ouro-numbers-with-their-comparison-conditions"&gt;Read Huginn and Ouro numbers with their comparison conditions
&lt;/h2&gt;&lt;p&gt;Geiping and colleagues&amp;rsquo; &lt;a class="link" href="https://arxiv.org/abs/2502.05171" target="_blank" rel="noopener"
 &gt;Huginn research&lt;/a&gt; separates input processing, a recurrent core and output processing. The core refines the internal state through repeated execution. The authors trained a 3.5B-parameter model on 800B tokens and reported improved reasoning-task performance as recurrent computation increased.&lt;/p&gt;
&lt;p&gt;The abstract&amp;rsquo;s 50B figure needs care. It describes improvements up to a computational load equivalent to 50B parameters. It does not guarantee the quality of a 50B model on every task, or that such quality is obtained at the same cost. A small set of weights using more computation is a research result; service economics require separate measurement.&lt;/p&gt;
&lt;p&gt;&lt;a class="link" href="https://arxiv.org/abs/2510.25741" target="_blank" rel="noopener"
 &gt;Ouro&lt;/a&gt;, from ByteDance and collaborators, was released in October 2025. The paper covers a family of 1.4B and 2.6B models and reports comparisons with models up to 12B across benchmarks. The &lt;a class="link" href="https://huggingface.co/ByteDance/Ouro-1.4B" target="_blank" rel="noopener"
 &gt;official Ouro-1.4B model card&lt;/a&gt;, however, describes that particular model as matching conventional 3–4B models. Saying that 1.4B always replaces 12B would overstate the comparison.&lt;/p&gt;
&lt;p&gt;Parameter counts should be read alongside training data, recurrence counts and evaluation tasks. Apparent inference efficiency may also follow substantial pretraining investment.&lt;/p&gt;
&lt;h2 id="fewer-stored-weights-do-not-remove-memory-bottlenecks"&gt;Fewer stored weights do not remove memory bottlenecks
&lt;/h2&gt;&lt;p&gt;Consider an illustrative calculation. Storing 3.5B parameters at 2 bytes each requires approximately 7GB for weights. Repeated use of those weights does not multiply their storage requirement by the recurrence count. This is arithmetic, not a measurement of Huginn&amp;rsquo;s total GPU memory usage.&lt;/p&gt;
&lt;p&gt;Total inference memory also includes the KV cache used to reuse prior context, intermediate states and execution workspace. &lt;a class="link" href="https://developer.nvidia.com/blog/mastering-llm-techniques-inference-optimization/" target="_blank" rel="noopener"
 &gt;NVIDIA&amp;rsquo;s inference optimization guide&lt;/a&gt; distinguishes weights and KV cache as major memory components. Longer contexts and more concurrent requests increase cache pressure. Whether caches can be shared between recurrent steps depends on the design.&lt;/p&gt;
&lt;p&gt;Weight reuse is also different from reduced data movement. If weights cannot remain in fast on-chip memory, another pass may require reading them from HBM again. Repetition can increase bandwidth demand along with computation. Without examining the memory hierarchy and implementation, looping cannot be declared a reason HBM becomes unnecessary.&lt;/p&gt;
&lt;h2 id="software-must-realize-the-savings-from-early-exit"&gt;Software must realize the savings from early exit
&lt;/h2&gt;&lt;p&gt;&lt;a class="link" href="https://arxiv.org/abs/2507.10524" target="_blank" rel="noopener"
 &gt;Mixture-of-Recursions (MoR)&lt;/a&gt; varies recursive depth by token and manages computation and caching around tokens still active at a given depth. The aim is to direct computation toward harder tokens rather than spend it on easy ones.&lt;/p&gt;
&lt;p&gt;Serving makes this harder. Different recurrence requirements across requests can reduce batching efficiency. Schedulers need to let other work use the resources freed by early completion. This is an anticipated operational challenge, not a measured result for a particular commercial product.&lt;/p&gt;
&lt;p&gt;The Ouro model card provides a concrete example. The model supports early exit, but the card states that vLLM does not support this feature and instead executes the configured full recurrence count. An architectural capability is not automatically implemented in a serving engine.&lt;/p&gt;
&lt;p&gt;This gives specific questions for AI infrastructure software companies such as Lablup: can the platform batch jobs with different recurrence depths, reuse caches, and reduce completion time and energy cost at matched quality? These are questions for assessing an opportunity, not claims that Lablup already supports those features or has demonstrated revenue growth from them.&lt;/p&gt;
&lt;h2 id="korean-semiconductors-face-both-resource-savings-and-usage-expansion"&gt;Korean semiconductors face both resource savings and usage expansion
&lt;/h2&gt;&lt;p&gt;The following are conditional scenarios for broader looped-model adoption, not earnings forecasts.&lt;/p&gt;
&lt;table&gt;
 &lt;thead&gt;
 &lt;tr&gt;
 &lt;th&gt;Condition&lt;/th&gt;
 &lt;th&gt;Possible industry effect&lt;/th&gt;
 &lt;th&gt;Evidence needed&lt;/th&gt;
 &lt;/tr&gt;
 &lt;/thead&gt;
 &lt;tbody&gt;
 &lt;tr&gt;
 &lt;td&gt;Fewer weights and less cache at matched quality&lt;/td&gt;
 &lt;td&gt;Lower memory pressure per request&lt;/td&gt;
 &lt;td&gt;Measured memory at equal context and concurrency&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;More recurrence on difficult problems&lt;/td&gt;
 &lt;td&gt;More accelerator time and energy demand&lt;/td&gt;
 &lt;td&gt;GPU time and energy per successful task&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Lower cost expands usage&lt;/td&gt;
 &lt;td&gt;Stable or higher aggregate infrastructure demand&lt;/td&gt;
 &lt;td&gt;Actual customer usage and purchasing plans&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Recurrence and cache management reduce batching efficiency&lt;/td&gt;
 &lt;td&gt;Delayed commercialization&lt;/td&gt;
 &lt;td&gt;Throughput at the same latency target&lt;/td&gt;
 &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;For memory suppliers such as Samsung Electronics and SK hynix, aggregate demand depends on both resources per request and the number of requests. Efficiency may stimulate adoption, but that growth cannot be assumed to exceed the savings. This analysis alone is insufficient to revise HBM demand or company earnings forecasts.&lt;/p&gt;
&lt;p&gt;A more useful comparison is the cost of completing one successful task. High benchmark scores can still be expensive if repetition takes too long or retries are frequent. Conversely, extra computation can lower total cost if it improves first-attempt success enough.&lt;/p&gt;
&lt;h2 id="the-next-test-is-task-cost-not-parameter-count"&gt;The next test is task cost, not parameter count
&lt;/h2&gt;&lt;p&gt;Testing the industrial case requires comparing total memory, completion time, energy and concurrent throughput at matched accuracy. Tail latency matters alongside averages for easy questions. If more recurrence stops improving quality, or batching losses exceed resource savings, the commercialization case weakens.&lt;/p&gt;
&lt;p&gt;Further architectural disclosure could establish whether Astra belongs in this research lineage. Meanwhile, a verifiable change remains: infrastructure planning must consider how long to compute on each problem and when to stop, alongside the size of the stored model. Turning that flexibility into lower actual costs is a joint test of hardware and software.&lt;/p&gt;</description></item></channel></rss>