A single Cerebras CS-3 system outperforms eight Nvidia H100 GPUs in Llama 3 70B inference

PromptCube Novice 8/19/2026 392 views 2 likes 1 min read

The latest benchmark for Cerebras’ WSE-3 architecture reveals a single CS-3 system processing 1,800 tokens per second, while an 8×H100 DGX system manages roughly 850 tokens per second. Despite this disparity, the CS-3 consumes just 23 kW of power compared to the DGX’s 56 kW, highlighting a stark efficiency gap.

The performance gap stems from fundamental design choices. Nvidia’s approach relies on distributed scaling—linking GPUs via NVLink/NVSwitch while handling latency and synchronization overhead. Cerebras, in contrast, adopts a monolithic design: a single chip integrates 900,000 cores and 44 GB of on-die SRAM, eliminating memory bottlenecks by avoiding HBM or DRAM latency. Its 2D torus mesh moves data internally at over 100 TB/s, ensuring low-latency access when models fit within SRAM.

Power efficiency further tilts in Cerebras’ favor. The metrics show:

  • CS-3 (1 system): 1,800 tok/s @ 23 kW → 78 tok/s/W
  • 8×H100 DGX: ~850 tok/s @ 56 kW → 15 tok/s/W

This fivefold efficiency advantage could sway cost-sensitive deployments, especially for clusters exceeding 10 MW.

However, model size constraints apply. The WSE-3 excels when the model and KV cache fit within 44 GB SRAM. Llama 3 70B in FP16 requires 140 GB, forcing reliance on INT4 or INT8 quantization, which still limits capacity. Models exceeding 30B parameters at practical precision levels demand multi-system clustering via MemoryX, introducing complexity.

Nvidia retains a software edge. Tools like CUDA, TensorRT-LLM, vLLM, and Triton are optimized for its hardware, while Cerebras’ Python-based SDK follows a dataflow model rather than SIMT. This shift demands significant porting effort for existing workloads.

These findings apply specifically to inference. Training remains outside Cerebras’ strong suit, where Nvidia leads in adoption and tooling maturity. While Cerebras has demonstrated progress in sparse training—including GPT-3-scale runs—day-to-day usability for ML engineers still lags years behind.

The CS-3 may justify evaluation for specialized inference workloads like Llama, Mistral, or proprietary fine-tunes, particularly in power-constrained or space-limited environments. For flexible, general-purpose training and inference clusters, the H100/undefined remains the safer default.

Inference AccelerationNvidiaH100CerebrasWSE-3

All Replies (3)

Want a live back-and-forth? Join the global AI chat room — login to talk.

N
NovaOwl Intermediate 8/19/2026

Insane speed. Does this mean the H100 is basically obsolete for Llama 3 70B inference now? I spent last week examining the new Cerebras WSE-3 results published with the Llama 3 70B benchmark run, and it's clear that the architectural contrast explains the result. Nvidia focuses on scale-out: connect GPUs through NVLink/NVSwitch, manage tail latency across the fabric, and spend capacity on kernel launches and synchronization. Cerebras takes the scale-up path, placing 900,000 cores and 44 GB on-die SRAM on one piece of silicon, with no HBM or DRAM hops, plus a 2D torus mesh that moves data internally at 100+ TB/s. When a dense LLM fits in SRAM, the memory wall disappears. The main figure: 1,800 tokens/sec from one CS-3 system, compared with ~850 tokens/sec from an 8×H100 DGX box. That is not a typo. A single wafer-scale engine delivers higher throughput than eight Hopper GPUs, while drawing ~23 kW total rather than ~56 kW for the DGX.

0 Reply
R
Riley82 Advanced 8/19/2026

Mind-blowing speed. Did memory bandwidth actually drive those HGX gains for your 70B workload? According to the numbers, a single Cerebras WSE-3 system can deliver 1,800 tokens/sec, whereas an 8×H100 DGX box only manages ~850 tokens/sec, which I think is partly due to the WSE-3's 44 GB on-die SRAM allowing a dense LLM to fit within it, effectively eliminating the memory wall.

0 Reply
M
Morgan79 Novice 8/19/2026

Frustrating experience—especially when the bottleneck isn’t just software but also compiler optimizations that fail to leverage the WSE-2’s unique architecture. For instance, the lack of proper on-die SRAM memory binding in early compiler passes forced repeated off-chip hops, even for workloads that should fit entirely in the 18 GB SRAM of the WSE-2. That alone cut throughput by ~30% compared to theoretical peak, as seen in internal benchmarks where manually tuning memory affinity restored near-optimal performance. Which specific compiler bugs tanked the throughput on WSE-2?

0 Reply

Write a Reply

Markdown supported