Samsung Electronics and SK Hynix each fell more than 3% in Seoul on Friday after DeepSeek said its newest AI model needs a fraction of the memory required by its predecessor. Both stocks had been trying to recover from July's selloff and remain more than 25% below their highs. Local retail traders, who helped drive the rally earlier this year, have sold around $10 billion of the pair this month alone.
The trigger came Thursday out of Hangzhou. DeepSeek's V4.1-Flash is fast, cheap and, by the company's own numbers, stronger than its flagship model. But the title of DeepSeek's paper had nothing to do with intelligence benchmarks. It called the release "Pushing the Limits of KV Cache Compression," and its abstract identifies memory consumption during long AI sessions as the main obstacle to making these models cheaper to run.
What DeepSeek Shipped
V4.1-Flash is a 552-billion-parameter model (parameters are the numerical values a model learns during training). DeepSeek does not activate all 552 billion for every word. It uses 8 billion parameters while processing input and 16 billion while generating output, a design that lowers the amount of computing needed for each step.
The model can handle a context window of one million tokens, meaning roughly a million small pieces of text or other input can remain available to it during a session. It also reads images natively, and its model weights are available under an MIT license, allowing anyone with sufficient hardware to run it. Nine providers were serving it through OpenRouter within a day of release.
The memory problem needs a little more explanation. As an AI model works through a long document, conversation or agent task, it keeps a running record of what it has already processed so it does not have to recalculate everything from scratch each time it generates another token. That record is called the key-value cache, or KV cache.
The cache is normally stored in high-bandwidth memory, or HBM. HBM consists of stacks of DRAM placed next to the processor to move data extremely quickly. It is among the fastest and most expensive memory in production, and booming AI demand for it helped turn SK Hynix, Samsung and Micron into some of the biggest semiconductor trades of 2026. As an AI session gets longer, however, the cache keeps growing. In long-running agent workloads, the memory needed for the cache can eventually exceed the memory occupied by the model itself.
DeepSeek has attacked that problem directly. Its new architecture reuses cached information across layers of the network, stores the cache at 4-bit precision and reconstructs part of it when needed rather than storing the entire record. The result, according to DeepSeek, is 890 bytes of cache per token, one-quarter of what the previous Flash model kept in HBM and one-eighth of what it wrote to SSD storage.
DeepSeek's post on X gave the commercial reason plainly. Cached input can account for a large share of the cost of running an AI agent. Shrink the cache, and the bill falls with it.
Memory Per Token, Down 437-Fold
The figure in DeepSeek's paper that matters most for Seoul is the second one. It tracks the amount of cache memory required for every token across successive generations of DeepSeek models. From the company's first release in January 2024 to V4.1-Flash, that figure has fallen 437-fold.

The decline did not begin this week. April's V4 had already cut KV-cache requirements to one-tenth of its predecessor at the full context window. September's model reduced them by another three-quarters. DeepSeek describes V4.1-Flash as the smallest member of a new architecture family designed to scale to larger models.
On the other side of that chart are the growth assumptions embedded in the memory trade. Micron says its entire calendar 2026 HBM supply is already contracted on both price and volume. In December, the company told investors the HBM market could grow from about $35 billion in 2025 to roughly $100 billion in 2028. By June it had moved the $100 billion estimate forward to 2027.
Micron reports earnings on September 30 and has guided the quarter to roughly $50 billion in revenue, about 350% above the year-earlier period, with an 86% gross margin.
How much premium memory each unit of AI work requires is the whole point.
Out With The 'Old'
This seems a little risky, but beginning September 14, DeepSeek will route every request for V4-Pro, its 1.6-trillion-parameter flagship, to V4.1-Flash and charge the lower Flash rate until a V4.1-Pro arrives. No date has been given for the larger model.
DeepSeek says the smaller model now beats V4-Pro on performance, cost, speed and the time required to complete a task. Model size alone is becoming a worse guide to how much computing and memory a useful AI system will consume.
Against the leading American models, V4.1-Flash fits the pattern that has held for much of this year. On DeepSeek's own benchmark tables, it narrowly beats the best reported score from Claude Opus 5 or GPT-5.6 Sol on Terminal-Bench 2.1, DeepSWE, CyberGym and Humanity's Last Exam when tools are allowed.
On newer and harder tests, however, the American frontier remains well ahead. V4.1-Flash scores 30.0 against Opus's 43.3 on Terminal-Bench 3.0, 31.2 against 51.8 on version 4.0, 20.3 against 37.0 on ProgramBench, and 15.3 against GPT-5.6 Sol's 33.7 on ExploitGym. The results have not been independently verified.
DeepSeek's open model is clearing benchmarks that defined the frontier last year at a fraction of the price, while the American frontier keeps moving to harder tests. It is the same pattern seen with V4-Flash earlier this year.
86x Cheaper!?
According to X user NIK (@ns123abc), V4.1-Flash is 86 times cheaper than the American flagships. This is because AI providers charge separately for input tokens, the material sent into the model; output tokens, the material the model generates; and cached input, previously processed material that can be reused without running the full computation again.
At peak hours, V4.1-Flash costs $0.30 per million input tokens and $1.20 per million output tokens, with prices cut in half during off-peak hours. Anthropic charges $5 and $25 for Claude Opus 5. On output, that makes DeepSeek roughly 20 times cheaper at peak and about 40 times cheaper off-peak.
Deepseek just dropped v4.1 flash, fully open weights
— NIK (@ns123abc) September 10, 2026
it beats gpt 5.6 sol, opus 5 and every chinese model on coding and cybersecurity at ~86x cheaper cost per million tokens running at 420-507 tok/s = faster than gemini 3.8 flash
"smallest model in our new architecture family"… pic.twitter.com/02Fft2OKTB
The 86-fold figure comes from cached input. Anthropic charges 50 cents per million cached tokens, while DeepSeek charges six-tenths of a cent. Cached input is precisely the cost DeepSeek's new memory architecture was built to reduce.

Meanwhile, DeepSeek raised its prices only a month ago. On August 16 it moved V4-Flash from flat rates of $0.14 for input and $0.28 for output to peak rates of $0.44 and $1.32. Thursday's release brought those prices back down, although not to July's levels, while giving V4-Pro customers a reduction of roughly 70%.

The timing also comes as DeepSeek moves toward the public markets. On Wednesday, Reuters reported that the company had hired CITIC Securities to prepare for a Shanghai listing. That followed a June financing round of about $7.4 billion involving investors including Tencent and CATL, along with reports of another potential round at a valuation near 500 billion yuan.
July Was Supply, September Is Demand
The July selloff in Korean memory stocks centered on supply. SK Hynix signaled a major increase in spending, while Chinese memory manufacturers continued ramping cheaper output.
Friday's scare came from the other side of the market: each unit of AI activity needing fewer memory chips.
Micron and SanDisk, which held up overnight, were rising in Friday's premarket after Oracle's cloud results. The Korean names, where leverage and local retail participation are greater, took the immediate hit.
Moves of 5% or more in a single day remain common for both Samsung and SK Hynix. Bloomberg notes that volatility remains near levels last seen during the 2008 financial crisis and the Covid shock. The shares also look inexpensive on conventional measures. Samsung trades at about 2.7 times book value and Hynix at 5 times, compared with roughly 11 times for the Philadelphia Semiconductor Index. Both Korean companies trade near 4 times forward earnings, versus 19 times for the index.
Within hours, Fibonacci Asset Management's Jung In Yun said cheaper AI could drive more usage and offset the efficiency gains. Eugene Asset Management's Ha SeokKeun called the issue a near-term concern.
When Seoul reopened after the Lunar New Year on January 31 last year, eleven days after DeepSeek's R1 release, SK Hynix fell as much as 12% in a day and Samsung dropped 4%. Investors initially feared that more efficient models would weaken demand for AI hardware. Instead, AI spending kept climbing, and the argument that lower costs stimulate greater usage won the year that followed.
R1 challenged the amount of computing required to produce useful AI. V4.1-Flash is attacking the amount of memory required for every token, and DeepSeek has published a chart, running back to January 2024, showing that requirement falling by more than 400-fold. The company also says the same architecture is intended for larger models.
Export Controls And Huawei's Own HBM
Since December 2024, U.S. export controls have barred sales of advanced HBM to China, making access to fast memory one of the hardware constraints on Chinese AI developers.
Huawei has been working on a domestic alternative. Its Ascend 950DT uses Huawei-made HiZQ 2.0 memory, with 144 gigabytes of capacity and bandwidth of 4 terabytes per second. That remains well behind the HBM SK Hynix supplies for Nvidia's leading accelerators.
DeepSeek's V4 in April was the first frontier model validated on Huawei Ascend hardware alongside Nvidia chips. The company also has a reported order for 160,000 Ascend 950DT processors for a gigawatt-scale data center in Ulanqab.
So - a model that needs one-quarter as much memory per token can be deployed more broadly on hardware that has less memory to offer. The V4.1-Flash model card does not identify the hardware used to train the model.
DeepSeek has not announced a date for the larger member of the V4.1 family. Micron reports on September 30.

No comments:
Post a Comment
Note: Only a member of this blog may post a comment.