- Together AI says its IBM-backed Nvidia compute could sell out before launch, signaling surging enterprise AI demand
- Open-source AI models are gaining enterprise traction as companies chase lower inference costs, customization and data control
- Together AI sees hybrid AI strategies taking hold, with open, custom and closed models used where they fit best
Together AI just inked a $240 million compute deal with IBM, but onlookers shouldn’t expect that capacity to last long. Company Chief Revenue Officer Kai Mak told Fierce it expects all of its Nvidia-based compute will be “pre-sold well before it comes online in January.”
The comment is a strong statement about the kind of demand Together AI is seeing as it pioneers development of and access to open-source AI models, especially for enterprises.
“While AI-native companies have always been core to our business, enterprises are adopting AI much faster than anticipated,” Mak said.
Cost is often the initial reason enterprises turn to open models, Mak said, in large part because they can “deliver substantially lower token costs than closed alternatives” depending on the workload. But he noted enterprises are also drawn to open-source options because they’re seeking the ability to customize and optimize AI for their specific performance and latency requirements. Mak added enterprises are also looking to retain control of their data and avoid vendor lock-in.
“Today, we are seeing strong interest in GLM, Kimi, MiniMax, [Nvidia’s] Nemotron family, DeepSeek, [Google’s] Gemma and Alibaba’s Qwen family. We also serve a growing class of customer-owned models that began with an open foundation model but have been extensively post-trained on proprietary data and tasks,” he said.
That’s why it turned to IBM for its latest deal, to deliver not just the Nvidia compute everyone wants but also the “reliability, security and operational maturity enterprises expect from IBM Cloud.”
Cutting inference costs
Using open models is one way to help cut inferencing costs as enterprises scale AI, but far from the only way.
Mak said one of the biggest opportunities is simply matching the right model to the right task. Enterprises are realizing most workloads do not require the largest frontier model, or, as Mak put it, “you don’t need a god to compose your emails.” Routing routine work to smaller open models while reserving heavyweight models for harder problems can cut bills by a factor of 10 without users noticing, he said.
Indeed, Storyblok exec Sebastian Gierlinger recently told Fierce the same thing. “If there is no education on that, people will automatically – most of the time – opt for the higher model and then automatically [use] more tokens than they need,” Gierlinger said at the time. And while it’s true token costs are expected to fall between now and 2030, Gartner warned token costs likely “will not be fully passed on to enterprise customers.”
Beyond model choice, Mak pointed to smarter harness engineering, arguing the harness can matter as much or more than the model itself in determining token efficiency. And advances in Nvidia chips, kernel optimization, quantization, attention mechanisms, speculative decoding and infrastructure utilization should further reduce the processing and memory required to generate responses, driving down the cost per token over time.
Mak added that while open-source AI is gaining traction, it’s not likely to replace proprietary models – at least not entirely. “For many enterprises, the likely end state is a hybrid strategy in which open, custom and closed models are each used where they make the most economic and technical sense,” he said.
Read more about Together AI and open AI:
IBM lends Together AI its enterprise street cred in $240M compute deal
Open weight AI vs open-source AI: What’s the difference?
How should enterprises and telcos think about open weight AI?