Primemas is attempting to dismantle the traditional server-side memory bottleneck by shifting from static, local memory constraints to a dynamic, rack-scale infrastructure model. At the Future of Memory and Storage (FMS) 2026 event, the company unveiled a CXL 3.0 portfolio designed to support massive AI and High-Performance Computing (HPC) workloads. Central to this strategy is the Abaco Project, a collaborative effort with Micron Technology to develop a disaggregated memory system. This system, intended for deployment at the Department of Energy’s Pacific Northwest National Laboratory (PNNL), aims to provide over 100TB of CXL-attached shared and pooled memory to satisfy the extreme data requirements of next-generation scientific AI.
Scaling Capacity via the Abaco Project and CXL 3.0
The Abaco Project addresses a fundamental infrastructure challenge: the inability of traditional hardware to scale memory capacity alongside expanding AI datasets without prohibitively expensive server over-provisioning. By utilizing CXL (Compute Express Link) technology, Primemas and Micron are positioning memory as a flexible, rack-scale layer rather than a fixed component of individual servers. This architecture allows CPU and GPU clusters to access hundreds of terabytes of pooled Micron DDR5 memory within a single rack. The foundation of this capability rests on Primemas’ proprietary Hublet® architecture, a high-density silicon platform designed to aggregate massive amounts of DRAM.
To facilitate this, Primemas is introducing the PMA14 and PMA16 CXL add-in cards. These PCIe-based expansion cards feature 14 or 16 RDIMM slots, supporting up to 3.5TB or 4TB of DRAM capacity when using 256GB RDIMMs. Primemas claims these products deliver up to four times the memory capacity per CXL port compared to conventional monolithic alternatives. The company expects to ship these cards to Micron in September for testing, qualification, and benchmarking within the Abaco system racks. By deploying three to four Abaco chassis per server rack, operators could potentially build shared memory pools exceeding 100TB.
SLiM Architecture and Near-Memory Computing Roadmap
Beyond large-scale rack deployments, Primemas is targeting the specific efficiency needs of AI inference through its new SLiM (Switchless Pooled Memory) architecture. SLiM is designed as a 1U system to manage KV-cache-intensive workloads, which require both high capacity and ultra-high bandwidth. The company is positioning SLiM as a space-efficient solution that enables multiple servers to access tens of terabytes of shared DRAM with ultra-low latency. Crucially, this architecture aims to bypass the cost and complexity typically associated with external CXL switches by utilizing a switchless design.
Looking toward future iterations of its hardware, Primemas is developing a roadmap for intelligent memory systems. The company intends to integrate CPU and FPGA acceleration capabilities directly into its Hublet® architecture. This move suggests a transition toward near-memory computing platforms specifically optimized for AI and database acceleration. By embedding compute functionality closer to the data, Primemas is signaling a strategic shift toward reducing the latency and energy overhead inherent in moving massive datasets between memory and processors. This evolution aims to further optimize how enterprise data centers handle the increasing computational density required by modern AI models.
Key Takeaways
- Primemas and Micron are collaborating on the Abaco Project to create a rack-scale system providing over 100TB of pooled CXL-attached memory for PNNL.
- The PMA14 and PMA16 CXL add-in cards support up to 4TB of DRAM capacity per card, aiming for four times the capacity of traditional monolithic options.
- The new SLiM (Switchless Pooled Memory) architecture is designed as a 1U system to provide tens of terabytes of shared DRAM for AI inference without external CXL switches.
TechInsyte's Take
In our view, the Primemas-Micron collaboration represents a critical pivot in data center architecture, moving away from the "server-as-a-silo" model toward true resource disaggregation. The Abaco Project is not merely a capacity upgrade; it is a direct response to the "memory wall" that currently limits the scaling of large-scale scientific AI and HPC workloads. By leveraging CXL 3.0 and proprietary Hublet® silicon, Primemas is attempting to prove that memory can be treated as a utility—pooled and distributed across a rack rather than trapped within individual nodes. If the September testing and subsequent PNNL deployment succeed, this could set a new standard for how enterprise IT leaders manage the massive memory footprints required by generative AI. The move toward near-memory computing via FPGA integration further suggests that the industry is preparing for a future where the distinction between storage, memory, and compute becomes increasingly blurred.
Questions & Answers
How does the Abaco Project solve the problem of server over-provisioning in AI workloads?
The Abaco Project utilizes CXL-attached shared and pooled memory to transform memory into a dynamic, rack-scale infrastructure layer. This allows compute clusters to access hundreds of terabytes of pooled Micron DDR5 memory within a single rack, preventing the need to over-provision every individual server node with expensive, static local memory.
What are the technical specifications of the PMA14 and PMA16 CXL add-in cards?
The PMA14 and PMA16 are PCIe-based CXL memory expansion cards. The PMA14 features 14 RDIMM slots, while the PMA16 features 16 RDIMM slots. When populated with 256GB RDIMMs, they support up to 3.5TB or 4TB of DRAM capacity, respectively, which Primemas states is up to four times the capacity of conventional monolithic alternatives per CXL port.
What distinguishes the SLiM architecture from traditional pooled memory solutions?
The SLiM (Switchless Pooled Memory) architecture is designed as a space-efficient 1U system specifically for KV-cache-intensive AI inference. It distinguishes itself by enabling multiple servers to access tens of terabytes of shared DRAM with ultra-low latency without requiring the cost and complexity of external CXL switches.
What is the strategic role of Primemas' Hublet® architecture in these new systems?
The Hublet® architecture serves as the proprietary high-density silicon foundation that enables the aggregation of terabytes of DRAM. According to the company, this silicon is essential for achieving the high memory density required by advanced scientific AI workloads that a single server cannot satisfy.
Source: Businesswire