AI infrastructure is entering a new architectural era. As AI inference scales, performance increasingly depends not simply on adding more compute, but on how efficiently compute, memory and connectivity operate together as one unified AI infrastructure system. Modern AI inference workloads are insatiable consumers of memory. Large language models, reasoning models and agentic AI applications require rapid access to model parameters, embeddings and rapidly growing key-value (KV) caches that preserve conversational context. These working data sets are growing into the hundreds of gigabytes, and increasingly terabytes, making memory capacity, bandwidth and latency just as important as accelerator performance. As AI infrastructure scales, overall system performance increasingly depends on how efficiently accelerators can access, move and utilize memory resources rather than simply adding more compute.
Why AI Needs a New Memory Tier
Today's AI memory hierarchy was never designed for inference at the scale modern workloads demand. High bandwidth memory (HBM) attached directly to GPUs delivers exceptional performance but remains expensive and capacity constrained. System DRAM provides larger memory pools but cannot economically scale alongside every accelerator. NVMe SSDs offer abundant capacity, yet their latency makes them unsuitable for serving active inference workloads.
This challenge is especially visible in large language models, where growing KV caches must remain readily accessible to avoid repeatedly recomputing previous tokens. Keeping these caches entirely in HBM is prohibitively expensive, while moving them to storage introduces latency that reduces token generation performance. The result is that GPUs increasingly spend valuable cycles waiting for data rather than performing inference.
Addressing this challenge requires extending the AI memory hierarchy with a new intermediate tier that delivers substantially greater capacity, typically expected from storage, while remaining much closer to compute with very low latency, that is very close to direct attached memory.

The Marvell Photonic Fabric Memory Appliance extends the traditional AI memory hierarchy with new shared memory tiers that bridge the gap between DRAM and storage, enabling higher-capacity, lower-latency AI inference.
Why Optical Shared Memory Matters
As memory resources extend beyond a single server, moving data efficiently becomes just as important as the memory itself. Traditional electrical interconnects become increasingly constrained as bandwidth, reach and power requirements continue to grow. Every additional electrical hop consumes energy, increases latency and limits the ability to make large memory resources appear as a unified pool.
Optical connectivity fundamentally changes this equation. By replacing long electrical paths with high-bandwidth optical links, shared memory can extend across multiple racks while maintaining the bandwidth and latency characteristics required for AI inference. Instead of treating memory as a resource attached to individual servers, optical fabrics enable memory to become shared infrastructure that can be dynamically allocated wherever workloads require it.
Keeping larger KV caches closer to compute reduces unnecessary data movement and memory stalls, allowing accelerators to spend more time generating tokens and less time waiting for memory accesses. This improves GPU utilization while enabling memory capacity to scale independently from compute, creating a more efficient foundation for next-generation AI infrastructure.
Introducing Marvell® Photonic Fabric™ Technology
Marvell® Photonic Fabric™ technology introduces a new optical connectivity architecture for AI inference that enables a shared-memory tier across AI infrastructure. Built around the Marvell® Photonic Fabric™ Memory Module, Photonic Fabric NIC and Marvell® Photonic Fabric™ chiplet, the platform enables memory sharing across multiple AI racks up to 30 meters apart, allowing hyperscalers to dramatically increase available memory without scaling compute resources proportionally.

The Marvell Photonic Fabric Memory Module combines shared DDR5 memory, HBM cache and high-bandwidth optical connectivity to enable memory expansion, pooling and disaggregation across AI infrastructure.
At the heart of the platform is the industry's first system-on-chip with optical I/O integrated into the middle of the silicon die. Designed specifically for AI memory disaggregation, each PFMM supports up to 8x DDR5 DIMMs, backed by 72GB HBM cache and 7.2 Tbps of optical fabric bandwidth. The architecture delivers approximately < 350 nanoseconds of near-NUMA memory access across multiple racks, allowing AI accelerators to access dramatically larger memory resources without the latency traditionally associated with external storage systems.
Complementing the PFMM, the PF-NIC provides a CXL® 3.1 and PCIe® Gen 6 optical network interface that connects AI servers to the shared memory fabric. Together, the PFMM and PF-NIC enable third-party shared-memory appliances capable of supporting growing KV caches, larger model contexts and increasingly demanding inference workloads while preserving compatibility with existing AI software frameworks.

Building the Next Generation of AI Infrastructure
As AI inference becomes the dominant workload, memory is emerging as one of the primary determinants of overall infrastructure performance. Increasingly, however, success will depend not simply on adding more compute, but on how efficiently compute, memory and connectivity operate together as one unified AI infrastructure system.
Photonic Fabric technology introduces a new optical connectivity architecture that enables a shared memory tier for AI infrastructure. By combining optical connectivity with memory disaggregation, it allows memory capacity to scale independently of compute while maintaining the bandwidth and latency required for large-scale AI inference. The result is larger warm KV caches, improved GPU utilization, higher token efficiency and lower energy associated with data movement.
AI infrastructure is evolving toward architectures where compute, memory and connectivity operate as one unified system. Photonic Fabric technology represents the first architectural proof point for Marvell, demonstrating how optical connectivity enables a new shared memory tier while laying the foundation for more scalable, efficient and cost-effective AI infrastructure.
Learn more: Explore the Marvell® Photonic Fabric™ technology platform.
# # #
This blog contains forward-looking statements within the meaning of the federal securities laws that involve risks and uncertainties. Forward-looking statements include, without limitation, any statement that may predict, forecast, indicate or imply future events or achievements. Actual events or results may differ materially from those contemplated in this blog. Forward-looking statements are only predictions and are subject to risks, uncertainties and assumptions that are difficult to predict, including those described in the “Risk Factors” section of our Annual Reports on Form 10-K, Quarterly Reports on Form 10-Q and other documents filed by us from time to time with the SEC. Forward-looking statements speak only as of the date they are made. Readers are cautioned not to put undue reliance on forward-looking statements, and no person assumes any obligation to update or revise any such forward-looking statements, whether as a result of new information, future events or otherwise.
Tags: AI, Data Center, Optical Module, server connectivity, Optical Connectivity, AI infrastructure