FMS: The Future of Memory and Storage 2026 opens Tuesday at the Santa Clara Convention Center and runs through August 6, with a debate running through all three days that will shape how AI inference infrastructure gets built for the next decade. The question is which technology fills the gap between the GPU’s small, expensive high-bandwidth memory and the vast cold storage behind it, and who profits from that gap. The three competing answers are high-bandwidth memory for the fastest tier near the GPU, CXL-attached DRAM for a warm expansion tier, and PCIe 6.0 NVMe flash for KV cache and bulk storage, and the conference program suggests the industry is no longer confident in any single one of those answers.
The Opening Act: Kioxia’s CM10
The conference has a concrete opening product. On July 30, Kioxia announced the CM10 Series, the first enterprise SSDs built on its 332-layer BiCS FLASH generation 10 TLC memory, and the drives arrive at FMS as the first enterprise SSDs to offer direct cold-plate liquid cooling alongside conventional air cooling. The liquid cooling capability is not a novelty feature. As flash storage is increasingly co-located with accelerator racks running hundreds of watts per chip, thermal management has become as structurally important as bandwidth. The CM10’s E3.S and 9.5mm E1.S form factors are the first enterprise SSDs designed for dense GPU pod deployments where cold plates, not fans, handle heat removal.
The CM10 uses a PCIe 6.0 interface, the generation that doubles per-lane bandwidth to 64 GT/s by switching from binary NRZ signaling to PAM-4 four-level encoding while adding forward error correction for the first time in PCIe’s history. Compared to Kioxia’s previous-generation CM9, the CM10 delivers up to approximately 92% higher sequential read performance and up to approximately 85% higher random read performance. Peak sequential read bandwidth reaches 28.4 GB/s; random read IOPS tops out at 6.29 million. Price and availability have not yet been disclosed. The drives are currently sampling to select customers and will be publicly demonstrated at the FMS exhibition floor for the first time beginning Tuesday.
What Is Driving PCIe 6.0 Adoption
The shift to PCIe 6.0 is motivated by a specific arithmetic problem in AI inference. When a large language model generates responses to concurrent users, it builds an attention memory called the key-value cache during each inference pass, storing the intermediate computation results it needs to avoid reprocessing earlier tokens. A 70-billion-parameter model running a one-million-token context window generates approximately 320 gigabytes of KV cache for a single user, four times the entire high-bandwidth memory capacity of an NVIDIA H100 GPU. That data has to go somewhere, and the nearest candidate is an NVMe SSD.
At PCIe Gen 5 bandwidth, about 14 GB/s for a typical enterprise drive, feeding that storage tier during active inference is a bottleneck. PCIe 6.0’s 28 GB/s ceiling for a standard x4 drive changes the equation, not by eliminating the bottleneck entirely, but by making flash storage a viable warm-tier participant in the AI memory hierarchy rather than a cold-tier afterthought. The engineering tradeoff is the forward error correction latency overhead: FEC adds approximately four to eight nanoseconds of latency per direction. For bulk throughput workloads this is negligible. For latency-sensitive inference pipelines where time-to-first-token matters, that overhead deserves explicit evaluation.
The CXL Counter-Argument
While PCIe 6.0 flash dominates the product announcement narrative, the CXL Consortium is sponsoring a full track at FMS 2026, and the data being presented challenges the assumption that NVMe SSD is the right answer for all KV cache storage. Micron’s Luis Ancajas is scheduled to present preliminary results from a CXL-based disaggregated memory architecture that connects a CXL JBOM module to NVIDIA’s Dynamo inference stack via Micron’s FAMFS, a fabric-attached memory file system that allows CXL-attached DRAM to operate as a warm memory pool accessible via standard file system semantics. Preliminary results from this architecture show a five- to ten-fold speedup over traditional storage-backed KV cache implementations in HPC and data center environments.
If CXL-attached DRAM can deliver five to ten times the effective KV cache performance of NVMe SSD at meaningful scale, that shifts the engineering question from which SSD is fast enough for KV cache to at what context length and batch size SSD becomes the right tier, and at what point does CXL DRAM win. Samsung researchers are also presenting an XGBoost-based hot-page placement framework that proactively routes memory to DRAM based on predicted access patterns, and a panel organized by the CXL Consortium features Meta’s hardware systems engineers presenting practical lessons from CXL memory in production hyperscale environments.
The CM10 is the first enterprise SSD with cold-plate liquid cooling support in the E3.S and 9.5mm E1.S form factors, designed for dense GPU pod deployments where air cooling is insufficient for co-located storage.
The Three-Day Program
Day 1 keynotes put Samsung and Kioxia on the same stage as NVIDIA. Kioxia SVP Neville Ichhaporia and GM of Memory Technical Marketing Katsuki Matsudera present the CM10 and BiCS10 portfolio alongside NVIDIA VP of Storage Technology Jason Hardy, who is expected to discuss how the shift from AI training to inference is driving new data center storage designs, including NVIDIA’s CMX architecture for shared pod-level KV cache. Samsung Electronics EVP Jin-Yub Lee then presents the company’s vision for memory-centric AI infrastructure covering next-generation HBM and flash alongside Samsung’s integrated device manufacturer advantages. SK hynix EVP Chunsung Kim outlines a tiered memory orchestration vision extending the traditional memory hierarchy with new intermediate tiers implemented via 3D-stacked DRAM on accelerators, high-bandwidth flash, and CXL-based memory pooling.
Day 2 turns to the economics of AI storage at scale. ScaleFlux CEO Hao Zhong and NVIDIA’s Jason Hardy co-present Wednesday’s headline keynote on how memory capacity has become one of AI’s most critical bottlenecks. A Marvell session makes the case that HDDs remain economically indispensable in the AI data center for cold storage, archive, and bulk training data. Day 3 opens with ScaleFlux Chief Scientist Prof. Tong Zhang and NVIDIA’s Vikram Sharma Mailthody revisiting the Five-Minute Rule from the 1980s, which defined when it is economically rational to keep data in memory versus store it to disk. The economics of that rule have shifted dramatically with every generation of storage technology, and the FMS session explicitly proposes to re-draw it. The cross-cutting theme tying all three days together is whether Samsung, SK hynix, and Micron’s aggressive HBM production scaling can keep pace with the AI inference workloads the rest of the program is describing, or whether the HBM-centric AI infrastructure narrative runs into supply or cost limits first.

