NEO is The New V-Raptor.
Up to 128-core Ampere¢ç Altra¢ç / Altra¢ç Max, 3.0 GHz, Arm¢ç v8.2+ ISA
DDR4-3200 ECC RDIMM 8 Slots (1 DPC)
Up to 2TB
2.5¡± / 3.5¡± SATA¢ç / SAS
Up to 4 Slots in Front Bay
PCIe¢ç 4.0 ¡¿16
4 Slots
2600W 3+1 Redundant PSU
80¢ç PLUS Platinum
Ubuntu Server 20.04, 22.04, 24.04
Rocky Linux 8.x, 9.x
Furiosa RNGD NPU is
More Efficient than GPUs
in Power, Cost & Concurrency.
Furiosa RNGD NPU inference servers, lower initial introduction cost than GPUs,
enable stable service operation with an AI-optimized architecture design.
512TFLOPSFP8 Performance
at TV-level Power.
Furiosa RNGD delivers world-class performance and unrivaled deployment flexibility for AI inference. Its groundbreaking power efficiency provides a scalable, secure, and cost-effective solution for all use cases, from on-premises to enterprise data centers and the Sovereign AI Cloud. It is a highly energy-efficient AI inference accelerator solution.
48GB HBM3 Dedicated Memory
1.5TB/s Bandwidth
Up to 512TFLOPSFP8 Performance
BF16, FP8, INT8, INT4 Supported
Faster Perf/W than H100
From Llama 3 Token Test
180W TDP
TV-Level Consumption
LLM Inference,
Simply At Furiosa SDK.
The SDK provides modules that enable direct execution on the NPU,
including optimized DNN model architectures and pre-trained model images.
Enabling Easy Development
With C/Python-Based Environment
The FuriosaAI SDK is a software development kit for writing C/Python applications that utilize the NPU. Using this SDK, you can leverage various tools, libraries, and frameworks from the C/Python ecosystem, which is the most widely used in the AI/ML field, for developing NPU applications. The SDK consists of various modules and provides inference APIs, quantization APIs, command-line tools, and server programs for serving.
And, With Ampere Altra Arm Chip
Works On Everything.
Powered by up to 128 Arm cores running at 3GHz,
the Ampere Altra series is a SOTA chip
designed to handle everything, from server applications
like NGINX and MySQL to cutting-edge multi-modal AI workloads.
Up to 128x Arm v8.2+ Cores
TDP 183W of Low Power Drives
Ultra-High Performance Intelligence
200% Increased Bandwidth
Record-High Connection Speeds
With PCIe 4 Support
Up to 2TB Memory
DDR4 3200Mhz
8ch 8 Slots1DPC Support
High-Grade Performance
Now in Your Server
V-Raptor Q100 offers flexible core configurations to meet your performance needs. How? With 12 well-defined processor categories — from 32 cores at 1.7GHz to 128 cores at 3.0GHz, and TDP options ranging from 40W to 183W. The choice is yours.
BMC function enables remote access and management of servers using dedicated advanced processors from ASPEED. In other words, you can operate it even if you are far away from the server. It's like carrying the server in your pocket wherever you go.
Inconvenient Truth About GPU Servers?
AI Inference Isn't
Something GPU Do.
While model inference has a lighter computational load than training, using high-performance GPUs causes power consumption and operating costs to swell, leading to a sharp drop in resource efficiency. Consequently, it often becomes an uneconomical choice that results in a high TCO (Total Cost of Ownership) relative to low resource utilization. AI inference is a job for the NPU.
See Configurations.
Available LLMs
Specs
|
Generative AI
|
|
Arm¢ç Based Ampere Altra¢ç Processor
Up to 128C 3.0GHz |
|
FuriosaAI
RNGD NPU |
|
DDR4 3200
8Ch 1DPC 8 Slots Up to 2TB Memory |
|
M.2 NVMe¢ç
8 Slots |
|
Up to 20X Faster Transfer
USB 3 Supported |
|
10 Gigabit Ethernet
|
|
Ubuntu Server 20.04 / 22.04 / 24.04
Rocky Linux 8 / 9 / 10 |
|
BMC Specialised
ASPEED¢ç 2500 |
|
RNGD results are based on internal measurements from FuriosaAI using SDK 2025.3.0.
GPU results were obtained using RunPod's vLLM 0.9.1 under similar test conditions. |
Processor
Ampere Altra¢ç Series
Q32-17
32C 1.7GHz
TDP 40W
Q64-22
64C 2.2GHz
TDP 124W
Q80-30
80C 3.0GHz
TDP 161W
M128-30
128C 3.0GHz
TDP 183W
Arm¢ç v8.2+ & SBSA 4 Instructrion Sets
TSMC 7nm FinFET
1MB L2 Cache/Core
Operate from 0¡ÆC to 90¡ÆC
4,962-Pin FCLGA
NPU
FuriosaAI
RNGD
256 TFLOPS (BF16)
512 TFLOPS (FP8)
512 TOPS (INT8)
1,024 TOPS (INT4)
TDP 180W
HBM3 48GB
PCIe 5 x16 Connection
TSMC 5nm
12 VHPWR Power
PCIe Dual-Slot Total Height 3/4 Length