AI-CPU

Purpose-built for Inference AI Head-Nodes.

The NR1® Chip is the first AI-CPU – Creating perfect system harmony with any GPU to eliminate bottlenecks and achieve maximum GPU output.

Engineered to transform your Inference-as-a-Service unit economics by solving the fundamental architectural bottleneck caused by dated general-purpose CPU and NIC technologies, NR1 AI-CPU is a Heterogeneous Compute server on a chip featuring enhanced video, audio, DSP, and CPU processors, with a hardware-based AI-Hypervisor that orchestrate the harmony between those processors and any attached GPU, all fused with an embedded AI-over-Fabric network engine. It delivers the best system performance-per-dollar

Background

Benefits of NR1 Disruptive Technology

Drives GPU Utilization to it’s maximum

Up 10x more energy efficient than legacy CPU+ NIC

Up to 6x million tokens per dollar

Lowest total cost of ownership

Greater throughput, improved cost and energy efficiency

CPU Reliant System

Saturates with Networking, Hardware control, and data processing tasks causing system bottlenecks and reduced output.

AI-CPU Head Node

Offloads Networking, Hardware control, and Data processing to dedicated hardware achieving Fully Utilized GPUs

Background

NR1 AI-CPU unleashes the true potential of GPUs

The NR1 Chip drives GPU utilization to its theoretical maximum, while increased GPU density in legacy x86 CPU-based AI servers results in significant utilization loss

What’s Inside

The NR1 Chip combines a server-class Arm CPU, integrated media and data processors, a hardware-driven NR1® AI-Hypervisor®,and an embedded AI-NIC – all on a single chip.

First true AI-CPU engineered
for inference at scale

Pairs with any AI Accelerator –
GPU, FPGA, ASIC

Works with any model – single
and multi-modality

Experience the Difference Between NR1 and Legacy CPU Architecture

Inference Appliance

The world’s first enterprise-ready Server, purpose-built for AI inference – Compact, Plug and Play, and Deployable in less than an hour

Get cutting-edge, high-performance AI for your on-premises Datacenter or in the Cloud. Pre-loaded with generative and agentic AI models, NR1 Inference Appliance, seamlessly integrates into existing AI infrastructure, enabling rapid scale-up or scale-out for data centers and cloud environments.By eliminating legacy CPU/NIC bottlenecks, it ensures near 100% GPU utilization, significantly reducing complexity and cost, enabling high-performance, efficient AI at scale.

Three Treasures in a Box

Performance/$ Redefined

Reduced Total Cost of Ownership through Higher Scalability, Lower System Cost, and Lower System Energy Consumption, delivering Best-In-Class Inference Requests per Dollar

Reduced Total Cost of Ownership through Higher Scalability, Lower System Cost, and Lower System Energy Consumption, delivering Best-In-Class Inference Requests per Dollar

Ready to Go

Vertically Integrated Inference Appliance experience with User-friendly Software Development Kit and Runtime, with API’s for Inference Development, Deployment, and Serving

Vertically Integrated Inference Appliance experience with User-friendly Software Development Kit and Runtime, with API’s for Inference Development, Deployment, and Serving

Generative & Agentic AI Ready

Built-in Optimized Open-Source Models; Fine tuning and Bring your own Model support; OpenAI API Ready; Backend support to all Agentic AI frameworks

Built-in Optimized Open-Source Models; Fine tuning and Bring your own Model support; OpenAI API Ready; Backend support to all Agentic AI frameworks

Background

Competitive Advantage

Our first cloud computing and financial services customers are currently running NR1 Inference Appliances on site – demonstrating accelerated systems performance, energy efficiency and real estate/space savings versus traditional x86 CPU-reliant inference systems

2.5X

Energy efficiency

2X

SERVER DENSITY

6X

Cost efficiency

Background

Specifications

Mechanical Form Factor

4U, 19” Rack Mount

PCI Express Capacity

20 slots of dual slot FHFL x 16 PCIe Gen5 card, accommodating 4-10 NR1 Inference Modules (AI-CPU) and 10-16 GPUs in 1 chassis, 5 switches

Performance

  • Neural Network Processing of up to 14POPs, DDR of 2.05TB / 8.8TBps and SRAM of 9.22GB
  • General Purpose Processing of 80 x Arm Neoverse N1 cores
  • 160 General Purpose DSP cores
  • 160 Audio DSP cores
  • 40 Video engines for 40K FPS@HD, 20K FPS@FHD,
  • 5K FPS@4K, 1200 FPS@8K

Networking

Up to 1 Tbps + redundancy

Host Memory

Up to DDR capacity of 1.6TB; DDR Bandwidth 2.56TBps

Storage

Up to 10 x 3.84TB E1.SSDs

Power

2+2 Redundance mode, Typical: 2.85KW

Cooling

6 modules, each 2x60x60 dual rotor fans

Management & Monitoring

BMC with 2 x 1Gbps on rear panel (RJ-45) management ports

Software

Server configuration, monitoring, and network security

Want to Learn More?