Hardware

Cerebras to supply roughly 100 megawatts of CS-4 hardware to Gimlet Labs for ultrafast inference cloud

The multi year agreement pairs Cerebras wafer scale accelerators with Gimlet Labs disaggregated inference software, aiming at up to 3,000 tokens per second for agentic and real time workloads.

T
By TechQuire Daily Staff TechQuire Daily Staff
September 28, 2026 / 7 min read

Artificial intelligence companies are spending heavily on inference, the step in which a trained model such as Anthropic's Claude produces an answer to a user prompt. As chatbots, coding assistants and autonomous agents move into production, the speed at which a model emits tokens has become a competitive metric, shaping user experience, cybersecurity response times, voice applications and financial analysis.

On September 28, 2026, Cerebras Systems and Gimlet Labs said they would work together to build what the two companies describe as a new class of ultrafast AI inference at massive scale. Under the collaboration, Cerebras will supply CS-4 wafer-scale systems with the capacity to consume roughly 100 megawatts of power, and Gimlet will fold that hardware into its disaggregated inference cloud, which spans datacenter infrastructure through to developer APIs.

The stated target is up to 3,000 tokens per second for demanding agentic and real-time applications. The first Cerebras-powered Gimlet Cloud datacenter is expected to come online later this year, while Gimlet plans to make the CS-4 hardware available in its cloud in 2027.

Cerebras Systems, which trades on Nasdaq under the ticker CBRS, builds wafer-scale AI infrastructure and is led by chief executive Andrew Feldman. Gimlet Labs is a San Francisco startup led by co-founder and chief executive Zain Asgar that raised a $300 million Series B round at a $3 billion valuation in early September 2026. The two companies have worked together since last year on joint customer engagements and already run an integrated solution serving tokens in private deployments.

Key Facts

Reuters reported on September 28 that Cerebras plans to supply AI chips and hardware with capacity to consume roughly 100 megawatts to Gimlet Labs to help run its cloud computing business. The wire service noted that the deal is part of rising interest among AI companies in procuring hardware capable of speedy inference. Cerebras plans to supply its CS-4 systems, which it unveiled over the summer, over one to two years, Feldman told Reuters, and Gimlet plans to make the CS-4 hardware available in its cloud in 2027.

GlobeNewswire reported on September 28 that the collaboration combines Cerebras wafer-scale compute with the Gimlet Cloud to deliver a purpose-built disaggregated inference cloud, with speeds of up to 3,000 tokens per second and a first Cerebras-powered Gimlet Cloud datacenter expected later this year. Gimlet Cloud combines the Cerebras Wafer Scale Engine with GPUs into an integrated inference solution and uses inference disaggregation technology to orchestrate model execution so each phase of inference runs on the silicon best suited to it.

GlobeNewswire reported on August 18 that Cerebras introduced the CS-4, describing it as the fastest AI accelerator in the industry. The CS-4 is a rack-scale system built from three Wafer Scale Engine 3 Turbo processors and is up to twice as fast as the CS-3, bringing its advantage in tokens per second per user over GPUs to as much as 30 times. It delivers 750 PFLOPS of AI compute, 7.2 terabits per second of I/O and 129.6 petabytes per second of memory bandwidth, with total compute fabric bandwidth of 160.5 PB/s and wafer-to-wafer latency as low as two microseconds, supporting models with more than 50 trillion parameters.

Each WSE-3T contains four trillion transistors and 900,000 AI-optimized cores across 46,225 square millimeters of silicon with 44GB of SRAM on the wafer, doubling AI compute to 250 PFLOPS per wafer and memory bandwidth to 43.2 PB/s. On GPT-OSS-120B, the CS-4 delivers more than 4,400 tokens per second per user, up to 30 times faster than GPU solutions.

GlobeNewswire reported on September 4 that Gimlet Labs raised $300 million in Series B funding led by Andreessen Horowitz, bringing its valuation to $3 billion and total funding to $392 million. The round included participating investor Sapphire Ventures, new investors M12 and Arm, and previous investors Menlo Ventures and Factory. Gimlet cites an estimated $765 billion in AI capital expenditure this year and $7.6 trillion in cumulative spending from 2026 to 2031.

Analysis

What this really means is that the fight over AI infrastructure is shifting from training to serving. Training still consumes the largest single clusters, but inference happens every time a user or an agent asks a question, and it repeats billions of times a day. Gimlet says inference is now reaching quadrillions of tokens per month. When a workload repeats that often, speed stops being a nice feature and becomes an economic lever: faster tokens per second per user translate into fewer accelerators needed to serve the same traffic, which changes datacenter economics.

The architecture choice is the more interesting part. Rather than betting on a single vendor, Gimlet disaggregates a model so that each phase of inference runs on the silicon best suited to it, mixing Cerebras wafer-scale engines with GPUs. Cerebras chief technology officer Sean Lie framed the split directly: combining the fastest tokens from Cerebras with the highest throughput GPUs delivers the best datacenter economics. That is a pragmatic position for a young company, and it is a notable one for Cerebras, which is selling into an environment it does not control. Feldman called the deal further validation of how easy it is to deploy even in heterogeneous environments.

For Cerebras, the deal is also a commercial proof point for the CS-4, which began shipping in the quarter that started in August 2026. Sean Lie described Gimlet as a launch partner for the CS-4, giving customers a direct path to the company's latest technology. The 100 megawatt figure matters because it is a power commitment, not a unit count, and power is the constraint that datacenter operators actually plan around. Roughly 100 megawatts over one to two years is a meaningful pipeline for a company that has to prove its rack-scale platform can be deployed at scale outside its own facilities.

There is a caveat worth stating plainly. Cerebras and Gimlet declined to disclose the financial terms of the deal, and the delivery schedule stretches over one to two years, which means the hardware volume is a plan rather than booked and shipped revenue. Gimlet will be responsible for maintaining and operating the systems once Cerebras delivers them. The collaboration also builds on engagements that began last year, so the announcement formalizes and expands an existing relationship rather than starting from zero, and Gimlet describes the expansion as covering software, infrastructure design, APIs, developer tooling, optimization, validation and production operations.

Why It Matters

Speed has become a strategic category in AI infrastructure. Nvidia and OpenAI have both signaled the importance of fast inference, and the market has responded: earlier in 2026 Cerebras signed a deal to supply OpenAI with chips, and in the prior year Nvidia signed a licensing deal with Groq for its chips and hardware. That pattern suggests the industry expects latency and throughput to determine which clouds win agentic workloads, where a model may call tools, check code or query data repeatedly before answering.

Gimlet's positioning explains why it needs partners rather than a single accelerator line. It targets companies and startups building entire products around AI, and frontier models that require enormous computing power to run speedily. Asgar has pointed to cybersecurity, voice applications and financial analysis as areas where performing inference quickly changes what a product can do. Gimlet also reports that it tripled its customer base in March 2026 and added one of the top three frontier labs and one of the top three hyperscalers as customers, and says it has since secured billions of dollars in contracted revenue while scaling to hundreds of megawatts in managed heterogeneous infrastructure.

For buyers, the practical implication is that multi-silicon clouds are becoming a normal procurement option rather than an experiment. Gimlet works with chip companies including Nvidia, AMD, Intel, Arm, Cerebras and d-Matrix, and now has a named wafer-scale partner for its fastest tier of service. If the 3,000 tokens per second target holds at production scale, it would sit well above the throughput per user that conventional GPU clusters deliver today, and it would give developers a reason to route latency-sensitive traffic to a specialist cloud instead of a general purpose one.

Next Up

The near-term milestone is the first Cerebras-powered Gimlet Cloud datacenter, which the companies expect to come online later this year, followed by CS-4 availability in Gimlet's cloud in 2027. Cerebras said first CS-4 shipments began in the quarter that started in August 2026, so the coming months will show whether the 100 megawatt pipeline converts into installed racks. Cerebras and Gimlet have not disclosed financial terms, and neither company has published a delivery schedule beyond the one to two year window.

Watch for details on how the disaggregated stack splits work between wafer-scale engines and GPUs, and whether Gimlet's frontier lab and hyperscaler customers adopt the ultrafast tier in production. Andreessen Horowitz managing partner Raghu Raghuram, a Gimlet board member, has argued that the answer to rising demand is not just more infrastructure but a better architecture. The Cerebras collaboration is the first large test of that claim at datacenter scale.

Tagged

Comments (0)

No comments yet. Be the first to share your thoughts.