Gimlet Labs, a San Francisco startup building what it calls the first multi-silicon inference cloud, announced on September 4 that it raised a $300 million Series B led by Andreessen Horowitz, a round that Bloomberg reported values the company at $3 billion. The company's own blog post did not state a valuation, but the figure marks a roughly 33-fold increase over the $80 million Series A it closed just over five months earlier, and it makes Gimlet one of the fastest-valuing companies in the AI infrastructure wave. SiliconANGLE reported on September 4 that the new capital will fund a major expansion of the company's managed compute capacity and its push into custom hardware.
Gimlet's pitch is that the AI industry has spent the last three years optimizing training, while inference, the part of AI that runs every day in production, is still being executed on hardware that was designed for something else. The company's software disaggregates inference workloads and routes different phases of a model's execution to whatever chip architecture does them best, whether that is a GPU, a purpose-built AI accelerator, a CPU or a near-memory dataflow design. The result, the company claims, is 3 to 10 times faster inference for the same cost and power footprint compared with running everything on a homogeneous cluster.
The round is notable for who is writing checks. Arm and Microsoft's M12 venture fund joined as strategic investors, alongside Sapphire Ventures, Menlo Ventures, Tiger Global Management and Samsung Ventures, and Gimlet said it is working with Arm on chip compatibility. The strategic participation from Arm and Microsoft signals that the biggest names in silicon and cloud see Gimlet's chip-agnostic layer as a hedge against any single vendor dominating AI inference hardware.
Key Facts
The financing details are clear even where the valuation is not. Gimlet Labs announced on September 4 that the Series B is led by Andreessen Horowitz, and Bloomberg reported on September 4 that it values the company at $3 billion. The company's total funding now stands at roughly $392 million, including a $12 million seed round and the $80 million Series A that closed just over five months before the Series B announcement.
Gimlet's growth claims are aggressive. In its September 4 announcement, the company said that since March it has added "billions in contracted revenue," built a "gigawatts of datacenter pipeline" and is scaling toward "hundreds of megawatts" of managed capacity. SiliconANGLE reported on September 4 that the company tripled its customer base in March 2026, adding one of the top three frontier AI labs and one of the top three hyperscalers as customers, along with unnamed financial-services firms.
The technology story is about heterogeneous disaggregation. Gimlet, which emerged from stealth in October 2025, works with chipmakers including Nvidia, AMD, Intel, Arm, Cerebras and d-Matrix, and its platform decides in real time which silicon should execute each stage of a model. The company says this approach delivers the 3 to 10 times performance improvement that underpins its contracts, and it argues that the era of renting one kind of GPU for everything is ending as inference demand explodes.
The company was founded in 2023 by CEO Zain Asgar, a former Nvidia GPU architect and Google AI engineering lead, together with Michelle Nguyen, Omid Azizi, Natalie Serrino and James Bartlett, who previously built Pixie Labs, a startup acquired by New Relic. TechFundingNews noted on September 4 that the founding team's background in observability and infrastructure software is central to Gimlet's thesis: managing inference across many kinds of chips is, at heart, a routing and monitoring problem.
Analysis
What this really means is that the AI infrastructure market is pivoting from training to inference faster than the hardware industry can adapt, and Gimlet is betting that the winner will be the software layer that makes every chip usable rather than the chipmaker with the biggest single product. The participation of Arm and Microsoft is the tell. Arm sells CPU intellectual property that competes with GPUs for inference, and Microsoft operates one of the largest cloud fleets in the world; both have an interest in a world where inference is not synonymous with Nvidia GPUs. Their investment is a hedge, but it is also a signal to the market that disaggregation is a real architectural direction, not a startup fantasy.
The bigger picture here is that the unit economics of AI are about to be decided at the inference layer. Training a frontier model is a spectacular, expensive event, but inference is the recurring cost that runs every time a model is used, and it is growing as agents multiply. Gimlet's claim that token demand is up roughly sixfold in 12 months, and that power is now the binding constraint on datacenter expansion, frames the opportunity: whoever can squeeze the most useful inference out of each megawatt will own the economics of the next phase of AI. If Gimlet's 3 to 10 times claim holds up at scale, it is not an incremental improvement; it is the difference between profitable AI services and money-losing ones.
The skeptics will point to the gap between claims and audited reality. "Billions in contracted revenue" and "gigawatts of pipeline" are non-standard metrics, and a private company's self-reported performance numbers should be read with caution until customers confirm them publicly. The valuation, reported by Bloomberg rather than the company, also means Gimlet is now priced like a company that must execute flawlessly. The counterargument is that the strategic investors are doing their own due diligence: Arm and Microsoft do not write checks of this size on marketing slides.
Why It Matters
For AI developers, Gimlet's rise is evidence that the cost of running models is about to become a competitive battleground. If multi-silicon inference clouds genuinely deliver the performance gains Gimlet claims, application builders will be able to run larger models or more agents for the same budget, and the pricing of AI APIs could compress as infrastructure gets more efficient. For chipmakers, the company is a demand aggregator: Gimlet's routing decisions determine which silicon gets bought, which gives it unusual influence over a supply chain that is used to dictating terms.
For the datacenter industry, the hundreds of megawatts Gimlet says it is scaling toward represent real power commitments, and the company's emphasis on power efficiency aligns with a broader market where electricity, not chips, is the scarce resource. For Nvidia, the round is a reminder that its dominance of training does not automatically extend to inference, where specialized accelerators from Arm licensees, Cerebras and others are competing for the same workloads. The next 18 months will show whether Gimlet's heterogeneous bet becomes the standard architecture for AI infrastructure or a cautionary tale about raising money on a thesis that hardware vendors make obsolete.
Next Up
The company says the new capital will fund hundreds of additional megawatts of managed capacity and an expansion into custom hardware, including an inference-optimized server design that does not use a traditional motherboard. Watch for the first public confirmation from Gimlet's named hyperscaler and frontier-lab customers, which would convert the company's self-reported traction into something auditors and competitors can verify. The competitive response also matters: if Nvidia and the major clouds move to bundle inference routing into their own platforms, Gimlet will need to stay ahead on performance, and the second half of 2026 should reveal whether its contracts convert into the kind of revenue that justifies a $3 billion price tag.
Comments (0)
Log in or sign up to leave a comment.
No comments yet. Be the first to share your thoughts.