AI servers packed with Nvidia accelerators can no longer hold their temporary context data in high-bandwidth memory, according to Wccftech. The volume of key-value cache generated during inference now overruns the HBM budget, forcing operators to tier that data down to storage. That single engineering fact relocates this cycle's bottleneck from compute logic to the memory-and-storage stack that feeds it.

The KV cache is the running record of a model's attention state during a session. Each additional token of context multiplies that cache, and long-context inference multiplies it fast. When the cache exceeds HBM capacity, throughput collapses unless the data spills to a faster storage tier — which is precisely the gap Samsung's V10 NAND is being positioned to close.

The Spec

Samsung's V10 NAND targets next-generation CMX storage built to absorb overflow KV cache at speeds high enough to keep accelerators fed. Higher layer-count NAND raises bits per wafer, which lowers cost per terabyte and makes cache-tiering economically viable at rack scale. The engineering benefit is a wider memory hierarchy; the financial benefit is more usable accelerator hours per dollar of installed GPU.

This matters because idle silicon is the most expensive line in any datacenter. A GPU stalled waiting on cache is capital earning nothing. If V10-based CMX storage recovers even a fraction of those stalled cycles, the effective cost per inference on the same hardware falls.

The Bottleneck

The constraint has moved, but it has not disappeared. HBM remains supply-constrained and priced by a handful of manufacturers; NAND now inherits the burden of everything HBM cannot hold. The relief comes at the industry's expense, as Wccftech frames it, because wafer starts redirected toward high-layer NAND for AI storage compete with the same fab capacity serving consumer and enterprise SSDs.

Every terabyte of NAND allocated to KV-cache tiering is a terabyte not allocated to the broader storage market — and that allocation shift sets the price floor for the entire category.

Naver's plan illustrates how expensive the full stack has become. The Korea JoongAng Daily reports Naver's AI factory with Nvidia requires deep pockets, sending the company to North America to raise capital. When an operator building on Nvidia silicon must cross an ocean for funding, the bill of materials — accelerators, HBM, NAND, networking, and power — has outrun domestic balance sheets.

The Unit Economics

Compare two paths to the same inference workload. Buying more GPUs to expand HBM headroom adds the most expensive component in the system to solve a data-volume problem. Tiering the overflow to high-layer NAND solves the same problem with a component that costs a fraction per terabyte.

The cost curve favors the tiered architecture wherever context lengths keep growing. A GPU-only approach scales cost linearly with cache size, since HBM is bonded to the accelerator and cannot be expanded independently. A storage-tiered approach decouples cache capacity from accelerator count, which is what bends the cost per inference downward. The winner of the inference cost curve is whoever controls the cheapest tier that keeps the accelerator busy.

That points less to the obvious beneficiary than the market assumes. Nvidia captures the headline because the bottleneck appears on its servers; that position is already priced into every AI infrastructure discussion. The less-recognized beneficiary is the NAND supplier whose high-layer product becomes mandatory infrastructure rather than optional storage — Samsung, if V10 reaches volume before competitors match the layer count.

The Inflection

The inflection arrives when V10-based CMX storage ships in qualified volume for AI servers. Qualification is the proof point, not the announcement, because hyperscalers do not deploy unqualified storage into inference clusters. Watch for design wins in Nvidia-based systems as the signal that cache-tiering has moved from workaround to standard architecture.

The claim would reverse if HBM capacity expands fast enough to absorb KV cache on-package, which would return the bottleneck to compute and strand the storage-tiering thesis. It would also weaken if a competing NAND supplier reaches comparable layer counts first, since fab capacity and yield decide who supplies the tier. At scale, whoever ships the lowest cost per usable terabyte of cache storage owns this node.