The AI compute provider picks IBM over the default hyperscalers, citing lower per token inference cost and a non overlapping business model.
Together AI is putting $240 million into IBM Cloud to rent a large cluster of Nvidia HGX B300 systems, the chipmaker's high-end GPU server platform for AI training and inference. The multi-year commitment, first reported by The Register, targets Q1 2027 deployment and is the first dedicated, large-scale inference cluster IBM Cloud has built around B300 silicon.
The choice of host is the story. Together AI runs an AI compute platform serving over a million developers and reports handling 400 trillion tokens monthly on its inference product, both figures the company disclosed in the IBM announcement. It could have parked that workload at AWS, Azure, or a pure-play GPU neocloud. It picked IBM because, in CEO Vipul Ved Prakash's words, the deal offered the lowest token cost at the GPU capacity needed. IBM Cloud does not compete with Together AI at the AI service layer, so the two are not bidding for the same enterprise customers.
IBM frames the B300 plus Spectrum-X Ethernet stack as built to deliver 30x more AI factory output than prior generations, an Nvidia-attributed marketing figure rather than an independent benchmark, and the announcement calls the hardware "last-gen" relative to Nvidia's roadmap. Q1 2027 availability is forward-looking. What is on the record today is one signed commitment and a non-vertical-integration pitch that the AI cloud market is not a closed three-horse race.