Get Free Shipping on Optical Transceivers Orders Over US$300
Currency: USD
USD - US Dollar
EUR - Euro
GBP - British Pound
CAD - Canadian Dollar
AUD - Australian Dollar
JPY - Japanese Yen
SEK - Swedish Krona
NOK - Norwegian Krone
INR - Indian Rupee
BRL - Brazilian Real
RUB - Russian Ruble
Need Help?
  1. English
  2. Русский
  3. Português
  4. Español
  5. Français
  6. Deutsch
  7. 한국어
  8. العربية
  9. にほんご
Select Currency
USD - US Dollar
EUR - Euro
GBP - British Pound
CAD - Canadian Dollar
AUD - Australian Dollar
JPY - Japanese Yen
SEK - Swedish Krona
NOK - Norwegian Krone
INR - Indian Rupee
BRL - Brazilian Real
RUB - Russian Ruble
Help
Filter

Categories

Clear All

NVIDIA DGX Spark

Sort by :Newest First

Filter
Filter
Filter
1 Results
Product Overview
The NVIDIA DGX Spark is a compact, CUDA-native AI workstation that puts a petaFLOP of AI compute and 128 GB of unified memory on a single desk. As a standalone box, it runs large language models up to roughly 200 billion parameters locally. What most resellers don't tell you is that the Spark is only the first node: link two units over a 200G QSFP56 connection, and you reach 405-billion-parameter models, and cluster three or more over 400G switches to scale even further.

What Is the NVIDIA DGX Spark?

The NVIDIA DGX Spark is NVIDIA's personal AI supercomputer—a desktop-sized system built around the GB10 Grace Blackwell Superchip, which combines a Blackwell GPU with a 20-core Arm Grace CPU in a single package. It is designed for AI researchers, data scientists, and developers who need to prototype, fine-tune, and run large models locally without waiting in shared cloud queues or sending sensitive data off-site.

Its defining feature is memory capacity. With 128 GB of LPDDR5x unified memory, the DGX Spark holds models that simply do not fit on a 32 GB discrete GPU, doing so at a fraction of the size, power, and cost of a full DGX rack. A single unit runs roughly 200-billion-parameter models; a matched pair with 200G networking reaches the 405-billion-parameter class.

The DGX Spark runs NVIDIA DGX OS—an Ubuntu-based distribution preloaded with NVIDIA AI Enterprise, CUDA, TensorRT-LLM, and NIM inference microservices—ensuring developers are productive on day one.

NVIDIA DGX Spark Specifications

· Superchip: NVIDIA GB10 Grace Blackwell (Blackwell GPU + 20-core Arm Grace CPU)
· AI Performance: 1 PFLOPS (1,000 AI TOPS) FP4, 5th-gen Tensor Cores
· Memory: 128 GB LPDDR5x unified, 256-bit, ~273 GB/s bandwidth
· Storage: 4 TB NVMe M.2 SSD (self-encrypting)
· Networking: ConnectX-7 SmartNIC, 2× QSFP56 ports (200G each)
· Management: 10 GbE RJ-45
· Wireless: Wi-Fi 7, Bluetooth 5.3
· Ports: 4× USB-C, 1× HDMI 2.1a
· Power: 240 W power supply (GB10 TDP 140 W)
· Dimensions & Weight: 150 × 150 × 50.5 mm, ~1.2 kg
· OS/Software: NVIDIA DGX OS, NVIDIA AI Enterprise, NIM, CUDA

The Problem: A Powerful Node That Needs a Network to Scale

A single DGX Spark is impressive, but it is still just one node. The moment a team wants to run a 405-billion-parameter model, serve several developers at once, or pool memory across units, a single box is not enough—and that is where most deployments stall.

The Spark's two QSFP56 ports are the answer, but they introduce a networking decision most resellers never explain: what cable, what transceiver, and what switch do you actually need? Get the connector format or topology wrong, and you end up with an unrecognized port, a half-speed link, or a cluster that underperforms from day one.

Scale a Single Spark: Two-Node 200G QSFP56 Pairing

The highest-value, lowest-complexity upgrade is pairing two DGX Spark units. NVIDIA's validated configuration connects two systems directly over one 200 GbE link, typically using a short 200G QSFP56 DAC cable. This results in 256 GB of combined unified memory and inference capabilities up to roughly 405-billion-parameter models, without requiring a switch.
· 1 × DGX Spark: Provides 128 GB of combined memory to support ~200B parameter models (no interconnect needed).
· 2 × DGX Spark: Provides 256 GB of combined memory to support ~405B parameter models, utilizing a direct 200G QSFP56 DAC/AOC interconnect.
· 3+ × DGX Spark: Memory and model capacity scale out with the nodes, requiring a 400G switch and breakout DACs.
The second QSFP56 port adds flexibility: you can run two 100G links for a switchless ring topology, or reserve one port for a peer Spark and the other for high-speed NVMe-oF storage.

Cluster Multiple Sparks: 400G Switches and Breakout DACs

Beyond two units, NVIDIA's supported path is an Ethernet switch. Because the Spark exposes 200G QSFP56 ports, clustering three or more units requires a 200G or 400G switch, with 400G ports fanned out to 2×200G using breakout DACs.

A typical multi-node DGX Spark cluster requires a specific bill of materials. The compute nodes are the NVIDIA DGX Spark units (GB10, 128 GB). To pair two units directly, a 200G QSFP56 DAC cable is used. For switch uplinks, a 400G QSFP-DD transceiver (400GBASE-DR4 / FR4) is necessary. A 400G-to-2×200G breakout DAC is required to fan one 400G port to two Spark ports. A 400G Ethernet switch serves as the fabric for three or more nodes, and single-mode or multimode fiber provides the necessary reach between racks or rooms.

NVLink-C2C vs. Ethernet: How the DGX Spark Actually Scales

One distinction clears up most of the confusion around the Spark. NVLink-C2C is the on-chip interconnect inside the GB10 Superchip—the high-bandwidth, low-latency link between the Grace CPU and the Blackwell GPU. It does not extend between systems.

Between two DGX Spark units, traffic runs over 200G Ethernet with RDMA (RoCEv2) through the ConnectX-7 NIC. That is a far simpler and more standard fabric than NVLink, and it is exactly why the Spark clusters use ordinary 200G/400G Ethernet switches and DAC cables rather than proprietary cabling. Understanding this—NVLink stays on the chip, while Ethernet scales the cluster—is the difference between a clean deployment and an expensive mistake.

DGX Spark vs. Mac Studio and Other Local AI Workstations

Buyers comparing local AI hardware usually weigh the DGX Spark against a Mac Studio or a discrete-GPU build. The right choice depends on whether you need CUDA, how large a model you run, and whether inference speed or memory capacity is your bottleneck.
· NVIDIA DGX Spark: Features 128 GB of unified memory (~273 GB/s bandwidth) and full CUDA ecosystem support. Priced at approximately $4,699, it is best for CUDA development, fine-tuning, and 120B-class models.
· Mac Studio M5 Ultra: Offers up to 512 GB of unified memory (up to ~1.2 TB/s bandwidth) but lacks CUDA support (relying on MLX/Metal). Priced at $5,499, it is best for fast, large-model inference.
· AMD Strix Halo: Provides up to 128 GB of unified memory (~256 GB/s bandwidth) and lacks CUDA (relying on ROCm). Priced around $2,000–$2,300, it is the best option for budget unified memory.
· RTX 5090 Build: Features 32 GB of GDDR7 memory (~1.8 TB/s bandwidth) and full CUDA support. Costing approximately $3,500–$4,200, it is ideal for fast iteration on small models.
The DGX Spark's primary strength is not raw decode speed—its 273 GB/s memory bus is deliberate, keeping the unit small and power-efficient. Its true advantages are CUDA compatibility, 128 GB of unified memory, and the ability to scale into a cluster over standard Ethernet. A Mac Studio cannot run CUDA-only libraries, and a 32 GB RTX 5090 cannot hold a 120B model at all. The Spark is the CUDA-native path that also grows beyond a single node.

Frequently Asked Questions

What is the NVIDIA DGX Spark?
The NVIDIA DGX Spark is a compact, desktop-sized AI supercomputer built on the GB10 Grace Blackwell Superchip. It combines a Blackwell GPU and a 20-core Arm Grace CPU with 128 GB of unified memory to run, fine-tune, and serve large AI models locally.

How much does the NVIDIA DGX Spark cost?
NVIDIA launched the DGX Spark at $3,999 and later raised the official price to $4,699 due to memory supply constraints. Reseller prices vary; FiberMall supplies factory-direct pricing on the Spark plus its 200G/400G interconnect. Request a quote for current availability.

What are the DGX Spark specifications?
The DGX Spark features the GB10 Grace Blackwell Superchip (1 PFLOPS FP4), 128 GB LPDDR5x unified memory, 4 TB NVMe storage, a ConnectX-7 SmartNIC with two 200G QSFP56 ports, a 10 GbE management port, Wi-Fi 7, and a 240 W power supply in a 150 × 150 × 50.5 mm chassis.

How do I connect two DGX Spark units?
Connect the two systems directly over one 200 GbE link using a 200G QSFP56 DAC cable. This pools the two units into 256 GB of unified memory and supports inference up to roughly 405-billion-parameter models, with no switch required.

What cable do I need to link two DGX Spark systems?
You need a 200G QSFP56 DAC cable for short runs, or a 200G QSFP56 active optical cable (AOC) plus QSFP56 transceivers for longer reach. FiberMall supplies both and confirms the exact length and format for your deployment.

Can the DGX Spark run a 405B-parameter model?
A single unit reaches roughly 200-billion-parameter models. Two units paired over a 200G QSFP56 link combine to support approximately 405-billion-parameter models.

How do I cluster more than two DGX Spark units?
For three or more units, connect them through a 200G or 400G Ethernet switch, using 400G-to-2×200G breakout DACs to fan each 400G switch port to two Spark ports. FiberMall supplies the transceivers, breakout cables, fiber, and switch as a single bill of materials.

DGX Spark vs. Mac Studio — which should I choose?
Choose the DGX Spark for CUDA-native development, fine-tuning, and scaling with NVIDIA's software stack and standard Ethernet networking. Choose a Mac Studio if you need maximum token-generation speed on large models and do not require CUDA. The two systems serve entirely different workloads.
More ↓