Cerebras unveiled the CS-4, its first multi-wafer rack system, packing three WSE-3 Turbo wafers for 750 petaflops of sparse FP16 compute and claiming up to 30x faster inference than GPU-based systems. The Register's analysis found the WSE-3 Turbo isn't new silicon but the existing WSE-3 die clocked from 1.4GHz to 2.8GHz, with per-wafer compute and bandwidth doubling in the pattern typical of a clock bump rather than a redesign. Power efficiency looks genuinely improved, at an estimated 120-140kW per rack versus roughly double that for comparable AMD/Nvidia racks. Cerebras named OpenAI, G42, MBZUAI, and AWS as launch partners but disclosed no CS-4 customer agreements or pricing. Financially, Q2 revenue was $180.1m (up 74% YoY but down sequentially from $193.4m), with a GAAP net loss of $450.5m and heavy customer concentration in G42 and MBZUAI. A genuinely new chip generation is planned for 2027.

4m read timeFrom thenextweb.com
Post cover image

Questions this post answers

What is the Cerebras CS-4 and how much faster is it than GPUs for AI inference?

The CS-4 is Cerebras's first multi-wafer rack system, packing three WSE-3 Turbo wafers for a combined 750 petaflops of sparse FP16 compute and 129.6 petabytes per second of memory bandwidth, supporting models above 50 trillion parameters. Cerebras claims it runs inference up to 30 times faster than GPU-based systems, measured as tokens per second per user on gpt-oss-120b against unnamed GPU systems. Shipments begin before the end of the quarter it was announced in. daily.dev tracks new AI inference hardware claims like this for engineers evaluating deployment options.

Is the Cerebras WSE-3 Turbo chip in the CS-4 actually new silicon?

No, analysis from The Register concluded the WSE-3 Turbo is not new silicon but the same die as the WSE-3, pushed from roughly 1.4GHz to 2.8GHz. It still has the same four trillion transistors, 900,000 cores, 44GB of on-chip SRAM, and TSMC 5nm node. Per-wafer compute and bandwidth both exactly doubled, a pattern consistent with a clock bump rather than a redesign; a genuinely new generation is scheduled for 2027. developers comparing AI chip vendor claims can follow hardware analysis breakdowns like this on daily.dev.

How much power does the Cerebras CS-4 rack draw compared to Nvidia and AMD systems?

The CS-4 is estimated at 120 to 140 kilowatts per rack, roughly half what comparable AMD and Nvidia rack systems draw, according to The Register's analysis. This power efficiency underpins Cerebras's claimed tenfold gain in throughput per watt, making it the area where the design appears most genuinely differentiated compared to the underlying chip, which is a clocked-up version of the existing WSE-3. teams weighing power efficiency in AI infrastructure choices can track hardware comparisons on daily.dev.

1 Impression