IN Brief:
- Abaco is being developed for PNNL with more than 100TB of CXL-attached shared and pooled memory per rack.
- Primemas PMA14 and PMA16 cards provide up to 3.5TB and 4TB of DRAM respectively.
- The cards are due at Micron in September for installation, testing, qualification, and benchmarking.
Primemas has detailed a rack-scale CXL memory architecture being developed for the US Department of Energy’s Pacific Northwest National Laboratory, where scientific AI and high-performance computing workloads require memory capacities beyond those available inside a conventional server.
The Abaco project is designed to provide more than 100TB of CXL-attached shared and pooled memory within a single rack. Primemas is working with Micron and other ecosystem partners to combine high-density controller silicon, DDR5 memory, and fabric management into an architecture that separates part of the memory resource from individual CPU or GPU nodes.
At the hardware level, Primemas has introduced two PCIe-based CXL memory cards. The PMA14 carries 14 RDIMM slots and supports up to 3.5TB of DRAM when populated with 256GB modules, while the PMA16 uses 16 slots to reach 4TB per card.
Three or four Abaco chassis can then be deployed in one server rack to create a shared pool exceeding 100TB. The cards are scheduled to ship to Micron in September, where they will be installed in Abaco racks for testing, qualification, and benchmarking.
The controller technology is based on Primemas’s Hublet architecture, which is intended to aggregate large quantities of DRAM behind CXL interfaces. The company describes the approach as part of a wider move from fixed server memory towards infrastructure in which capacity can be allocated more independently from the processor sockets using it.
That separation addresses an increasingly awkward server-design constraint. Local DRAM capacity is normally determined when a system is configured, forcing operators either to provision every node for its largest likely workload or accept that applications requiring more memory must be divided across multiple machines.
Pooling allows several compute nodes to reach a larger external memory resource instead. It does not make the attached capacity identical to processor-local DRAM — latency, bandwidth, topology, and software allocation remain important — but it gives system architects another tier between local memory and slower storage.
The distinction becomes useful in scientific computing, where datasets and simulations can outgrow the practical memory capacity of a single server before arithmetic throughput itself is exhausted. PNNL’s target workloads include scientific AI, HPC, and other data-intensive applications with memory footprints that can stretch into tens or hundreds of terabytes.
CXL provides the interconnect layer intended to make such architectures practical. Built on PCI Express physical infrastructure, the standard supports coherent communication between processors, accelerators, and memory expansion devices, allowing additional capacity to be exposed to a host without treating it simply as block storage.
Large-scale pooling introduces its own engineering problems. Access to external memory has to remain predictable enough for the target workload, while controllers and fabric-management software must coordinate allocation, error handling, and access across several hosts.
The value of a 100TB pool therefore depends on more than its headline capacity. Workload behaviour will determine whether an application benefits from keeping a large dataset in shared DRAM or whether the latency penalty of accessing more distant memory erodes part of the advantage.
Primemas is also developing a separate architecture called SLiM, or Switchless Pooled Memory, for inference systems with large key-value caches. The 1U design is intended to give several servers access to tens of terabytes of shared DRAM without an external CXL switch.
The company’s roadmap extends further towards near-memory computing, where selected processing functions are moved closer to the data rather than repeatedly transferring information back to a host CPU or accelerator. That approach can be useful when data movement consumes a disproportionate share of the time or energy required by a workload.
None of these approaches replaces high-bandwidth memory beside an AI accelerator or low-latency DRAM directly attached to a CPU. They create additional choices within the memory hierarchy, particularly where capacity rather than absolute access speed is the dominant constraint.
The September shipment to Micron will put the Abaco hardware into a more revealing stage. Qualification and rack-level benchmarking should show how the high-density CXL architecture behaves once controllers, pooled DDR5, fabric management, and real scientific workloads are tested together rather than considered as separate specifications.



