XCENA has introduced the MX1 production memory lineup at FMS 2026 in Santa Clara, California, targeting the memory constraints currently limiting AI inference at scale. By utilizing Compute Express Link (CXL) technology, the company aims to provide hyperscalers and cloud service providers a method to scale memory efficiency alongside compute. This launch transitions XCENA from its MX1P prototype phase into production-level evaluations, specifically addressing the growing KV cache footprints and data-movement bottlenecks associated with larger generative AI models and longer context windows.
MX1 Production Lineup and CXL Architecture
The MX1 lineup consists of two distinct products designed to optimize AI infrastructure. MX1 Compute integrates CXL-based memory expansion with near-data processing powered by 2,048 RISC-V cores. This architecture places compute adjacent to memory to reduce data movement between CPUs, which XCENA claims improves inference performance and lowers power consumption. Complementing this, MX1 Expand features eight DRAM slots, supporting server scale-up and DRAM reuse to provide a cost-efficient path for expanding memory resources.
To demonstrate these capabilities, XCENA is showcasing a CXL memory pool of up to 20 TB and a KV cache sharing demonstration at FMS 2026. Additionally, the company is partnering with Intel to demonstrate a CXL-based memory architecture on the Intel Xeon 6 platform. This specific integration focuses on offloading KV cache to CXL-attached memory, which Debendra Das Sharma of Intel notes is increasingly important for supporting memory-intensive AI workloads in hyperscale environments.
Enterprise Validation and Implementation Path
XCENA is moving from proof-of-concept engagements with the MX1P prototype toward commercialization discussions with enterprise AI infrastructure customers and hyperscalers. To support this transition, the company is presenting technical sessions at FMS 2026 focused on implementation and validation. Chief Product Officer Harry Kim is detailing CXL-based KV cache sharing via DRAM-NAND converged memory, while Senior Marketing Manager Jay Yeon is introducing a formalized out-of-box evaluation methodology to standardize the performance and reliability testing of CXL memory devices.
The company's strategy focuses on "bringing compute to data" to mitigate the high cost and capacity limits of high-bandwidth memory. By addressing stranded DRAM and rising total cost of ownership, XCENA positions the MX1 as a solution for operators facing peaking pressures in inference complexity. The company did not disclose further details regarding specific pricing or general availability dates in the announcement.
Key Takeaways
- The MX1 Compute product utilizes 2,048 RISC-V cores to enable near-data processing and reduce CPU-to-memory data movement.
- XCENA demonstrated a CXL memory pool capable of reaching up to 20 TB during the FMS 2026 event.
- The MX1 architecture is compatible with Intel Xeon 6 platforms to support KV cache offloading for AI inference.
TechInsyte's Take
In our view, XCENA’s shift from the MX1P prototype to a production lineup signals that the industry is moving beyond the theoretical benefits of CXL toward practical, hardware-level deployment. By integrating RISC-V cores directly into the memory expansion layer, XCENA is not just adding capacity but attempting to redefine the memory-compute relationship. This suggests that for CIOs and infrastructure leaders, the bottleneck for AI scaling is shifting from raw GPU compute to memory orchestration. If the MX1 can successfully reduce the "stranded DRAM" problem, it could fundamentally alter how hyperscalers architect their inference clusters to handle massive context windows.
Source: BUSINESSWIRE