
IBM and startup Together AI have signed a $240 million multi-year agreement to build a large-scale artificial intelligence cluster on IBM Cloud using Nvidia systems, as announced by the companies on Tuesday. The partnership reflects the rising demand for computing capacity to run open-source AI models as businesses seek to optimize AI costs and address cybersecurity concerns. The cluster will specifically provide inference capabilities for open-source AI models, which have gained significant traction in the current market environment. Together AI plans to use the environment to deliver inference services for open-source models, expanding its AI Native Cloud platform for enterprise customers.
The AI cluster utilizes Nvidia's HGX B300 systems with the chipmaker's newer Blackwell processors and Spectrum-X Ethernet networking gear, creating an AI factory architecture intended to support high-throughput, low-latency inference workloads. According to Reuters, the initial cluster will feature about 2,000 Nvidia Blackwell 300 chips and will be located in the U.S. IBM said it is the first dedicated, large-scale cluster built for inference on IBM Cloud using those systems. Nvidia has said the Blackwell chips are optimized for AI inference, positioning the infrastructure to handle the growing demand for inference processing in the AI ecosystem. This technical configuration is built to deliver 30x more AI factory output compared to prior generations, though workload-level performance will depend on model architecture, precision, batch size, and serving configuration. The cluster is expected to come online in Q1 2027, providing concrete deployment tied to IBM's broader Nvidia partnership announced in March.
Inference, the process of running trained AI models to generate responses, has become one of the largest drivers of demand for computing capacity, as reported by Business Standard. This surge in inference demand has prompted cloud providers and chipmakers to spend billions of dollars expanding AI infrastructure. The partnership addresses the growing need for scalable AI computing solutions as businesses increasingly rely on open-source AI models for their operations. Together AI reports that its inference platform now serves 400 trillion tokens monthly, highlighting the scale of demand for production-grade inference infrastructure. Together AI CEO Vipul Ved Prakash stated that "Enterprises want the performance of the best frontier models without the closed-model price tag, and that only works if the infrastructure underneath is fast and reliable at scale." Together AI's chief revenue officer Kai Mak told Reuters that "We think this will be sold out at least two to three months ahead of time" and "We'll have full offtake well before it's ready for service."
The agreement comes at a time when open-source AI models have gained traction as businesses seek to rein in AI costs and weigh concerns about cybersecurity incidents involving models from Anthropic, OpenAI and Meta, as reported by Business Standard. The cluster's focus on open-source AI model inference positions it to serve the evolving needs of businesses looking for cost-effective and secure AI solutions. Together AI CEO Vipul Ved Prakash emphasized that "Working alongside IBM with NVIDIA gives us that foundation. This cluster lets us bring production-grade inference to more companies, faster, and it's a big step in our push to make open-source AI the obvious choice for enterprises." IBM General Manager Alan Peacock made a particularly revealing comment about IBM's approach to AI infrastructure: "I'm not going to go and sign up for massive infrastructure deals to build a lake and then hope somebody comes to drink from it." This distinction highlights that IBM is building AI infrastructure when somebody has already come to drink from the lake, rather than speculating on future demand. IBM's participation isn't that it has suddenly produced the world's most powerful LLM - IBM brings enterprise relationships, hybrid infrastructure, security, governance and integration capabilities. Together AI raised an $800 million Series C financing round at an $8.3 billion valuation, and the company selected IBM and Nvidia based on their product roadmaps and ability to deliver GPU capacity at the pace required for AI scaling at low token cost.
The IBM and Together AI deal is part of a broader collaboration between IBM and Nvidia that also spans GPU-native data analytics, unstructured data extraction, on-premises and cloud infrastructure, and consulting services, IBM said. IBM's push into open-source AI infrastructure follows a separate commitment to the open-source ecosystem, with the company announcing Project Lightwell, a $5 billion commitment with Red Hat to help enterprises secure open-source software using AI tools and more than 20,000 engineers. That initiative centers on a trusted enterprise clearinghouse that uses AI to vet and verify patches across large portions of the open-source ecosystem, with pilot participants including Bank of America, Goldman Sachs, and JPMorganChase. For IBM, the deal provides a concrete deployment tied to its broader Nvidia partnership announced in March, which includes plans to bring Nvidia Blackwell Ultra GPUs to IBM Cloud and integrate Nvidia's data and inference tools into IBM's Red Hat AI Factory stack. IBM Research says enterprises moving from LLM experimentation into production are increasingly choosing on-premises deployment to retain control of their infrastructure, data and costs, with edge inference particularly suited to applications such as factory sensors and hospital monitoring where latency and privacy matter.