Navigation

Cerebras

An American AI computing company that develops wafer-scale accelerators and offers hardware systems, cloud training, and managed inference.

Report an issue

Cerebras Systems, commonly known as Cerebras, is an American semiconductor and artificial-intelligence computing company. It develops wafer-scale processors, complete computing systems, and cloud services for model training and inference.1 The company is headquartered in Sunnyvale, California.

The Cerebras Wafer-Scale Engine is a specialized AI accelerator rather than a graphics-processing unit. Cerebras cloud services compete with cloud GPU providers for some workloads, but do not provide conventional GPU virtual machines. Its hosted model API, Cerebras Inference, is an inference provider.2

History

Cerebras was founded in 2015 by Andrew Feldman, Gary Lauterbach, Michael James, Sean Lie, and Jean-Philippe Fricker. Feldman became chief executive.13 Several of the founders had previously worked together at SeaMicro, a server company acquired by AMD in 2012.

The company introduced the first Wafer-Scale Engine (WSE-1) and the CS-1 system in 2019. The second-generation WSE-2 and CS-2 followed in 2021. Cerebras announced its third-generation WSE-3 and CS-3 in March 2024.14

Cerebras priced an initial public offering on 13 May 2026. Its shares began trading on the Nasdaq Global Select Market under the symbol CBRS the following day.53

Wafer-Scale Engine

Most semiconductor wafers are cut into many separate dies before packaging. Cerebras instead links a grid of computing cores across most of a wafer and packages it as a single processor. Redundant cores and communication links allow the design to operate around manufacturing defects.

The WSE-3 contains approximately four trillion transistors and 900,000 AI-oriented cores, according to Cerebras. It is installed in the liquid-cooled CS-3 system, which includes power, networking, and supporting software.46 External MemoryX systems store and stream model weights, while the SwarmX interconnect links multiple CS systems.

Cerebras software compiles supported PyTorch models for the WSE and presents a cluster as a single logical system.7 The platform has its own compiler, runtime, and software-development kit. CUDA-specific programs and custom GPU kernels are not directly compatible with the WSE architecture.

Products

Cerebras sells CS systems for installation in customer data centers. These systems provide local control of hardware and data and can be joined into larger clusters.6 The company also makes its hardware available through Cerebras Cloud and partner-operated facilities.

AI Model Studio is a managed training service hosted on dedicated CS-3 clusters at Cirrascale Cloud. Customers submit supported PyTorch workloads and are charged for the model-training service.7 The service abstracts the underlying wafer-scale cluster rather than exposing it as a general-purpose accelerator virtual machine.

The company’s software products are collectively known as CSoft. They include PyTorch integration, a Model Zoo, the Cerebras SDK, and the Cerebras Software Language for lower-level kernels.8 Model compatibility and compiler support are significant parts of the platform because software written specifically for GPU ecosystems may require adaptation.

Cerebras Inference

Cerebras introduced its hosted inference service in August 2024.9 The service operates supported open-weight models on WSE systems and exposes them through an OpenAI-compatible API. A self-service pay-per-token tier was introduced in October 2025.10

Developer and enterprise plans provide different rate limits, priorities, service arrangements, and support. Enterprise offerings also include custom model weights and training or fine-tuning services.11 Cerebras determines the deployed model versions, serving software, and underlying capacity.

The company markets the service primarily on token-generation speed and has published results showing substantially higher output rates than selected GPU-based endpoints.2 These are company benchmarks and depend on the model, quantization, prompt length, concurrency, and measurement method. Network latency and queueing also contribute to the response time experienced by an application.

Cerebras distinguishes production models from preview models in its catalog. Available models and their status can change as new versions are introduced or older deployments are retired.12 The service is also available through partners including OpenRouter, which can act as a separate routing and billing layer.

Business

Cerebras earns revenue from hardware systems, cloud capacity, software, and related services. Its public-offering filings describe a business dependent on specialized semiconductor manufacturing, cloud infrastructure, and a limited number of large customers.13 The company’s architecture provides an alternative to GPU clusters for supported workloads, while its distinct software and hardware environment can create additional work when moving applications to or from GPU platforms.

References

Footnotes

  1. Company, Cerebras. 2 3

  2. Cerebras Inference, Cerebras. 2

  3. Investor FAQs, Cerebras. 2

  4. Cerebras CS-3, Cerebras, 12 March 2024. 2

  5. Cerebras Systems announces pricing of initial public offering, Cerebras, 13 May 2026.

  6. CS-3 system, Cerebras. 2

  7. Cerebras software platform, Cerebras. 2

  8. Cerebras developer documentation, Cerebras.

  9. Introducing Cerebras Inference, Cerebras, 27 August 2024.

  10. Cerebras Inference now available via pay per token, Cerebras, 13 October 2025.

  11. Pricing, Cerebras, accessed 18 July 2026.

  12. Model overview, Cerebras inference documentation.

  13. Form S-1 registration statement, Cerebras Systems, filed with the US Securities and Exchange Commission, 2026.

Search