- SambaNova told Fierce Network its unique architecture can do the job faster, cooler and with less energy than traditional GPU-based systems
- The company's air-cooled inferencing chip can slot into existing data center facilities
- SambaNova is targeting neoclouds, enterprises, telecom providers and sovereign AI initiatives
SambaNova isn’t trying to win the AI chip race by outmuscling Nvidia at training. It’s betting the bigger prize is inference and says its unique architecture can do the job faster, cooler and with less energy than traditional GPU-based systems.
GPUs ended up becoming the initial workhorse for AI because the parallel computation they used to render graphics pixels happened to be perfect for model training. But inference – which involves running queries over the information graphs created during training – demands a different solution, SambaNova Chief Product and Strategy Officer Abhi Ingle told Fierce.
“We are solving a different problem, which is how do you minimize data movement when you’re actually flowing the graph over the chip to get the answers that you want,” Ingle said. When you do that, he added, you end up with a chip that is more energy efficient, generates less heat and can run really fast.
That’s what SambaNova has created with its reconfigurable dataflow unit (RDU). The result is an air-cooled inferencing chip that can slot into existing data center facilities. In a world where many top-of the line chips now require liquid cooling, that’s actually a huge deal.
“Our current generation of racks only consumes about 10 kilowatts on average, 19 kilowatts peak, which means it can go into an existing cloud data center as opposed to the whole darn thing being retrofitted,” Ingle said. He added racks based around its latest SN50 chips – which ships later this year – will consume an average of 20 kilowatts on average.
Additionally, SambaNova’s three-tier memory structure — spanning on-chip SRAM, board-level DDR and high-bandwidth memory — lets a single rack hold multiple copies of models like DeepSeek, Ingle said, making on-prem AI deployments more feasible.
“SambaNova is building a processor with specific capabilities in inference by putting extra memory on the chip so it doesn’t have to call external memory as often, which can lead to a major bottleneck,” J.Gold Associates Founder Jack Gold told Fierce. “Latency is critical in agentic AI and especially in Physical AI. And SambaNova can make the RDU configurable to enhance multi-model and agentic processes. In theory that is a major advantage.”
Carving out a market niche
The company is peddling its chips to neoclouds (like Europe’s OVHCloud) and AI labs, but is also explicitly targeting enterprises, telecom providers and sovereign AI initiatives.
The goal, Ingle said, is not to supplant Nvidia and other GPU vendors, but complement them by offloading inferencing tasks and freeing up GPUs for more heavy lifting on the training side. He also talked up the potential of hybrid disaggregated inferencing, in which inferencing tasks (prompt, decoding, action) are spread across the best chip for each job. In this scenario, which he demoed for Fierce, Ingle said SambaNova’s chips could actually improve the utilization and performance of Nvidia GPUs.
Ingle said the company already has one telecom customer outside the U.S. and is engaged with several others across the globe, particularly in countries focused on energy usage and sovereignty. “I believe it’s a very natural role for them to be able to play, to step into being a provider of sovereign inference,” he said.
The coming enterprise inference wave
Analysts have long expected inference to overtake training as AI’s dominant workload — and IDC data suggests that shift is underway. In December, IDC’s Dave McCarthy noted inference now accounts for 47% of AI operations, the largest workload segment, and 37% of AI workloads are expected to run on-premises as hybrid deployments gain steam.
SambaNova isn’t alone. Rebellions, Cerebras and Groq (recently acquired by Nvidia) have all chased the same inference opportunity, underscoring how quickly chip startups are moving beyond the training-heavy GPU market.
“SambaNova is about low latency inference, a la Cerebras or Groq. But unlike Groq in particular, it does this with SRAM, HBM, and DDR,” Moor Insights and Strategy VP and Principal Analyst Matt Kimball told Fierce. “So, while Groq (its closest competitor) is super-fast at token delivery, SambaNova can support those bigger models because of its tiered memory. And the upcoming SN50 should be even more performant with multi-model memory. It’s an interesting play.”
The company’s tech is already gaining traction among early AI adopters, with financial giant JPMorgan Chase notably among those deploying its RDU-based servers on-premises to run AI workloads.
Ingle said AI adoption is “happening in waves” with AI natives leading the way and enterprises now moving from proof-of-concepts to production. As enterprises make this move, they’re evaluating whether to run workloads in the cloud or on-prem. The shift toward the hybrid on-prem+cloud strategy highlighted by McCarthy is where SambaNova sees an opening.
Ingle predicted that “most enterprises will run a hybrid strategy,” adding “I think that curve is starting.” Ingle said.
Where things are headed
SambaNova has raised around $2.5 billion and is working toward an IPO next year. Its latest $1 billion round included backing from Intel (which owns a 9% stake in the company) and private equity investors
The company debuted its latest generation chip, the SN50, in February, following up on the launch of its SN40L in 2023. Ingle said its chips are manufactured by TSMC in Taiwan and confirmed the SN50 is in production and on schedule to become commercially available toward the end of this year.
Though memory and component shortages are impacting all vendors in this space, Ingle said SambaNova’s connections with Intel and its other investors are helping it “secure some of that supply chain” so it can meet the “extremely strong” demand it’s seeing for its chips.
Ingle added that the money from its latest fundraising round is going toward three primary areas: securing the aforementioned supply chain so it can produce its chips at scale, continuing its innovation with the development of its forthcoming SN60 and SN70 chips, and working with customers on deployments.
As far as the SN60 and SN70 are concerned, Ingle said SambaNova plans to focus on improvements in speed and efficiency.
Read more about the burgeoning AI chip market here:
What is an AI accelerator?
Podcast: Data Center 2.0—Compute comes knocking
Red Hat exec: Sovereign AI could drive CPU demand for inference
Here’s why SK Telecom is building its own AI stack
Here’s why Nvidia is dropping $20B on Groq’s AI tech
IBM thinks Groq’s chips could make all the difference for enterprise AI
CPUs are the unsung heroes of AI
Watch out Nvidia: AI startup bets big on reversible computing