By Abed Mohammad Kamaluddin, Director, Custom Cloud Solutions, Marvell, and Vienna Alexander, Marketing Content Professional, Marvell

For most of the last decade, scaling AI meant scaling compute. If you built faster accelerators and wired enough of them together, the models would follow.
That is no longer the whole picture. Models now run to hundreds of billions of parameters with extensive context windows, and whether an expensive accelerator is working or just waiting comes down to memory: how much you have, how fast you can reach it, and how much time you lose moving data around.
Once memory leaves the server and rides a switched fabric, is it still memory, or has it become a network? This was the central question of MemNetAI, the first workshop on Memory-Semantic Networking for AI-Scale Systems, launched by Marvell with researchers from IIT Hyderabad and IIIT Delhi. Held at ACM SIGCOMM 2026 in Denver and guided by a program committee spanning academia and industry, it brought speakers from Cornell, alongside industry experts and researchers presenting their work, to debate these questions on the bleeding edge of AI infrastructure.
By Chander Chadha, Director of Product Marketing, Storage Products, Marvell
As AI models grow larger and inference workloads scale, larger models, longer context windows and growing KV caches are driving demands on memory resources. Traditional CPU-centric networked JBOF (Just Bunch of Flash) can’t keep pace with the throughput, latency, and efficiency requirements of these modern AI clusters.
DPU (data processing unit) -based storage is a compelling alternative to traditional storage, acting as the broker between the network and SSD for remote storage. By offloading storage, networking and security processing from the host CPU onto a dedicated DPU, AI infrastructure can move data closer to compute, reduce latency, and free up valuable CPU cycles for AI workloads.
To address this need, Marvell offers the OCTEON DPU family, purpose-built for hyperscale cloud workloads and data center applications, extending its use specifically into network storage acceleration for AI environments.
By Khurram Malik, Associate Vice President, Custom Cloud Solutions, Marvell, and Kangkyu Park, Vice President, System Architecture, SK hynix
The rapid growth of AI workloads is creating unprecedented demands on data center architectures. Modern AI applications, including large language models (LLMs), generative AI, recommendation systems, and high-performance computing, require significantly higher memory capacity, bandwidth, and efficiency.
Traditional compute-centric architectures are increasingly limited by the movement of data between processors and memory. As AI models continue to scale, excessive data movement creates performance bottlenecks, increases latency, and drives higher power consumption.
To address these challenges, the industry is moving toward memory-centric computing architectures enabled by Compute Express Link® (CXL®). CXL enables flexible memory expansion, memory pooling, and new system architectures that allow compute resources to operate more efficiently.
Marvell and SK hynix are collaborating to enable the next generation of memory-centric computing by combining Marvell® Structera™ A CXL-based near-memory acceleration technology with SK hynix advanced memory solutions. Together, Marvell and SK hynix are helping accelerate the adoption of CXL-enabled architectures for AI data centers by delivering a highly efficient and scalable approach to memory processing.
By Ravi Mahatme, Senior Director, Product Management, Photonic Fabric Business Unit, Marvell
AI infrastructure is entering a new architectural era. As AI inference scales, performance increasingly depends not simply on adding more compute, but on how efficiently compute, memory and connectivity operate together as one unified AI infrastructure system. Modern AI inference workloads are insatiable consumers of memory. Large language models, reasoning models and agentic AI applications require rapid access to model parameters, embeddings and rapidly growing key-value (KV) caches that preserve conversational context. These working data sets are growing into the hundreds of gigabytes, and increasingly terabytes, making memory capacity, bandwidth and latency just as important as accelerator performance. As AI infrastructure scales, overall system performance increasingly depends on how efficiently accelerators can access, move and utilize memory resources rather than simply adding more compute.
Why AI Needs a New Memory Tier
Today's AI memory hierarchy was never designed for inference at the scale modern workloads demand. High bandwidth memory (HBM) attached directly to GPUs delivers exceptional performance but remains expensive and capacity constrained. System DRAM provides larger memory pools but cannot economically scale alongside every accelerator. NVMe SSDs offer abundant capacity, yet their latency makes them unsuitable for serving active inference workloads.
This challenge is especially visible in large language models, where growing KV caches must remain readily accessible to avoid repeatedly recomputing previous tokens. Keeping these caches entirely in HBM is prohibitively expensive, while moving them to storage introduces latency that reduces token generation performance. The result is that GPUs increasingly spend valuable cycles waiting for data rather than performing inference.
By Arifur Rahman, Director of Product Marketing, Custom Cloud Solutions, Marvell
![]()
Modern AI workloads are insatiable consumers of memory. Deep learning recommendation models (DLRM), large language model (LLM) inference, in-memory databases and vector search engines all share a common bottleneck: there is never enough DRAM, and what exists is very expensive.
At today's spot prices—$27–$37 per GB for server-grade DDR5 RDIMMs1—a 12TB memory pool requires nearly half a million dollars in DRAM alone. Meanwhile, AI infrastructure buildouts are consuming server DRAM capacity faster than fabs can produce it, driving prices up 300–400% since mid-2025.1, 2
CXL memory expansion was supposed to solve this. And it does—but there's a subtler lever that most solutions ignore: the data sitting in that memory is compressible, and most CXL controllers don't touch it.