In brief

Cerebras and Gimlet Labs announced an inference-cloud partnership on September 28. Gimlet will connect its software with Cerebras wafer-scale computers. The companies say the service could deliver up to 3,000 generated tokens per second for agents and real-time apps. The first Cerebras-powered Gimlet data centre is expected later in 2026.

New to this? Read it in simple words
  • Cerebras and Gimlet Labs plan a cloud service for running AI models.
  • The service will join Gimlet software with Cerebras computers built around very large chips.
  • The companies say it could generate up to 3,000 tokens each second.
  • That speed is a target, and the first data centre is still planned for later in 2026.
Words to know
Inference
The stage when a trained AI model produces an answer.
Token
A small piece of text that a language model reads or writes.
Wafer-scale chip
A very large chip that uses most of one silicon wafer.

The planned service

Cerebras and Gimlet Labs announced a partnership on September 28. Gimlet Cloud will add Cerebras wafer-scale systems. A wafer-scale chip uses most of one silicon wafer as a single computing device.

The companies say their service could deliver up to 3,000 generated tokens per second. Tokens are small pieces of text. They are the units that language models read and write.

The first Cerebras-powered Gimlet data centre is expected to open later in 2026. It will be an early site for the CS-4 system. The companies did not give an exact date or location.

Sources12

SPEED CLAIM 01
A planned path from wafer-scale compute to an API.

The 3,000-token speed is a partner target, not an independent result.

Why speed matters

Fast output can help coding assistants, voice agents and other services that must respond while a person waits. It can also let one system serve more users with the same time budget.

Gimlet says its software can separate parts of inference and send them to the best hardware. Inference is the stage when a trained model produces an answer. The partners want to join that routing with Cerebras compute.

The plan covers the path from data-centre machines to developer APIs. That could make the unusual Cerebras hardware easier to use without customers managing it directly.

Sources12

A target, not a measured service

The 3,000-token figure is a company target. The announcement does not name the models, input sizes or test method behind it. Those details can change a speed result by a large amount.

The partners also did not publish prices, power use, customer names or service limits. Fast output is valuable only when the answer quality stays useful and the cost fits the job.

The launch matters because it gives Cerebras another route into cloud inference. The next evidence should be independent tests on real models after the first data centre comes online.

Sources12

Sources

Every fact in this story comes from the sources below. Open them to check our work.

  1. 1
    Primary source · September 28, 2026Gimlet Labs Adds Cerebras to Deliver Ultrafast AI Inference Cerebras
  2. 2
How we checked this story

We read the joint company announcement and compared it with HeadsUpAI’s summary. The speed and timetable are partner claims. We label them as targets and note the missing model, price, power and test details.