ComputeLabs Research

Cerebras claims CS-4 exceeded 4,400 tokens per user-second on GPT-OSS-120B and reached 30 times GPU-based inference speed.

· ComputeLabs Research · from the August 18, 2026 edition

Cerebras said CS-4 generated more than 4,400 tokens per user per second in a GPT-OSS-120B test. The company also claimed performance as high as 30 times that of graphics processing unit-based inference systems.

These are vendor-reported benchmark claims from Cerebras. The supplied report does not identify the exact GPU system, batch size, precision, latency threshold, software stack, power envelope, or other test conditions needed for an independent like-for-like comparison.

Cerebras separately said CS-4 can deliver up to 10 times the throughput per watt of CS-3. It also said the system can support models with more than 50 trillion parameters, although the source does not describe the model architecture, partitioning method, or memory configuration behind that capacity claim.

Additional reporting

  • Cerebras
  • CS-4
  • GPT-OSS-120B

All 20 stories from August 18, 2026