icon

AI Inference Server
With NPU.

NEO is The New V-Raptor.

Securing Unrivaled Bandwidth of
1.5TB/s With HBM3

Up to 4X
RNGD NPU Appliable

Bottleneck-Free Computation
With Many Core Structure

icon

W433, H176.5, D790, mm (4U)

Processor

Up to 128-core Ampere¢ç Altra¢ç / Altra¢ç Max, 3.0 GHz, Arm¢ç v8.2+ ISA

Memory

DDR4-3200 ECC RDIMM 8 Slots (1 DPC)
Up to 2TB

Storage

2.5¡± / 3.5¡± SATA¢ç / SAS
Up to 4 Slots in Front Bay

Expansion

PCIe¢ç 4.0 ¡¿16
4 Slots

Power

2600W 3+1 Redundant PSU
80¢ç PLUS Platinum

Operating System

Ubuntu Server 20.04, 22.04, 24.04
Rocky Linux 8.x, 9.x

Furiosa RNGD NPU is
More Efficient than GPUs
in Power, Cost & Concurrency.

Furiosa RNGD NPU inference servers, lower initial introduction cost than GPUs,
enable stable service operation with an AI-optimized architecture design.

512TFLOPSFP8 Performance
at TV-level Power.

Furiosa RNGD delivers world-class performance and unrivaled deployment flexibility for AI inference. Its groundbreaking power efficiency provides a scalable, secure, and cost-effective solution for all use cases, from on-premises to enterprise data centers and the Sovereign AI Cloud. It is a highly energy-efficient AI inference accelerator solution.


48GB HBM3 Dedicated Memory
1.5TB/s Bandwidth


Up to 512TFLOPSFP8 Performance
BF16, FP8, INT8, INT4 Supported


Faster Perf/W than H100
From Llama 3 Token Test


180W TDP
TV-Level Consumption

Reduces Power Consumption with Unrivaled Inference Perf/W. FuriosaAI Internal Test Std.

Secure Concurrency for LLMs by Equipping Up to 8X RNGDs on a Single Server.

Comprehensive SW Toolkit & User-Friendly API for Optimizing LLMs in RNGD facilitate Seamless, State-of-the-art LLM Deployment.

LLM Inference,
Simply At Furiosa SDK.

The SDK provides modules that enable direct execution on the NPU,
including optimized DNN model architectures and pre-trained model images.

Enabling Easy Development
With C/Python-Based Environment

The FuriosaAI SDK is a software development kit for writing C/Python applications that utilize the NPU. Using this SDK, you can leverage various tools, libraries, and frameworks from the C/Python ecosystem, which is the most widely used in the AI/ML field, for developing NPU applications. The SDK consists of various modules and provides inference APIs, quantization APIs, command-line tools, and server programs for serving.

And, With Ampere Altra Arm Chip
Works On Everything.

Powered by up to 128 Arm cores running at 3GHz,
the Ampere Altra series is a SOTA chip
designed to handle everything, from server applications
like NGINX and MySQL to cutting-edge multi-modal AI workloads.

Up to 128x Arm v8.2+ Cores
TDP 183W of Low Power Drives
Ultra-High Performance Intelligence

200% Increased Bandwidth
Record-High Connection Speeds
With PCIe 4 Support

Up to 2TB Memory
DDR4 3200Mhz
8ch 8 Slots1DPC Support

High-Grade Performance
Now in Your Server

V-Raptor Q100 offers flexible core configurations to meet your performance needs. How? With 12 well-defined processor categories — from 32 cores at 1.7GHz to 128 cores at 3.0GHz, and TDP options ranging from 40W to 183W. The choice is yours.

BMC function enables remote access and management of servers using dedicated advanced processors from ASPEED. In other words, you can operate it even if you are far away from the server. It's like carrying the server in your pocket wherever you go.

Inconvenient Truth About GPU Servers?

AI Inference Isn't
Something GPU Do.

While model inference has a lighter computational load than training, using high-performance GPUs causes power consumption and operating costs to swell, leading to a sharp drop in resource efficiency. Consequently, it often becomes an uneconomical choice that results in a high TCO (Total Cost of Ownership) relative to low resource utilization. AI inference is a job for the NPU.

Using GPU Servers for AI Inference is both Costly and Excessive.

Relying on GPUs for the Inference Stage results in a Significant Waste of Power and Budget.

In Real-World AI Deployments, the Volume of Inference Tasks far outweighs That of Training.

See Configurations.

Available LLMs

Specs

Generative AI

Arm¢ç Based Ampere Altra¢ç Processor
Up to 128C 3.0GHz

FuriosaAI
RNGD NPU

DDR4 3200
8Ch 1DPC 8 Slots
Up to 2TB Memory

M.2 NVMe¢ç
8 Slots

Up to 20X Faster Transfer
USB 3 Supported

10 Gigabit Ethernet

Ubuntu Server 20.04 / 22.04 / 24.04
Rocky Linux 8 / 9 / 10

BMC Specialised
ASPEED¢ç 2500
RNGD results are based on internal measurements from FuriosaAI using SDK 2025.3.0.
GPU results were obtained using RunPod's vLLM 0.9.1 under similar test conditions.

Processor

Ampere Altra¢ç Series

Q32-17
32C 1.7GHz
TDP 40W

Q64-22
64C 2.2GHz
TDP 124W

Q80-30
80C 3.0GHz
TDP 161W

M128-30
128C 3.0GHz
TDP 183W

Arm¢ç v8.2+ & SBSA 4 Instructrion Sets

TSMC 7nm FinFET

1MB L2 Cache/Core

Operate from 0¡ÆC to 90¡ÆC

4,962-Pin FCLGA

NPU

FuriosaAI

RNGD
256 TFLOPS (BF16)
512 TFLOPS (FP8)
512 TOPS (INT8)
1,024 TOPS (INT4)
TDP 180W
HBM3 48GB

PCIe 5 x16 Connection

TSMC 5nm

12 VHPWR Power

PCIe Dual-Slot Total Height 3/4 Length