IBM and Together AI signed a $240 million, multi-year agreement on August 11, 2026, to build a large AI inference cluster on IBM Cloud. The project is aimed at a recurring enterprise expense: running models after they've been trained.

The planned cluster will use Nvidia HGX B300 systems, built on Blackwell 300-generation hardware, connected through Nvidia's Spectrum-X Ethernet fabric. It will initially contain roughly 2,000 Blackwell chips and is targeted to go live in Q1 2027.

Together AI will operate its inference platform on that hardware. The company, valued at $8.3 billion as of July 2026, serves models including DeepSeek, MiniMax, and Kimi. Its pricing is positioned below what AWS, Azure, and Google charge for equivalent workloads through proprietary model APIs. Whether that advantage holds for a particular application depends on the model and the workload.

Together AI's chief revenue officer expects the cluster to sell out two to three months before it opens. As of August 13, that remains a forecast, but it suggests the company expects substantial demand before the new capacity is available.

What open inference offers

Inference is the work a model does when it processes a request and produces an answer. A hosted inference service handles the hardware and software needed to do that at production scale.

The useful distinction here is access to model weights, the learned parameters that determine how a model behaves. Together AI's platform serves models with publicly available weights. That gives customers more options for where to run a model and how to adapt it, rather than tying them entirely to one provider's API and fine-tuning system.

Calling all of this “open-source AI” can blur the distinction. Publicly available weights are the relevant feature in this comparison. They give enterprises a deployment option that a closed, API-only model doesn't offer.

With a proprietary API, customers rent access to a capability they cannot independently run on-premises or reproduce if the provider withdraws a version. A pricing change or model update can therefore affect an application that depends on it. That tradeoff can be reasonable, especially when the proprietary model performs better or the managed service saves substantial work.

For a bank processing documents or a hospital generating clinical summaries, however, deployment control has practical value. Model changes, data residency, and the ability to maintain a particular version can all become operating and audit concerns.

Together AI reportedly serves approximately 400 trillion tokens per month across its platform. Tokens are the units of text models process and generate. The reported volume supports the argument that hosted open-model inference is operating at substantial scale. Enterprise uses include document ingestion, code generation, internal search, and support automation, although the figure alone doesn't establish how much traffic comes from each type of customer.

IBM's role is in running the models

The strategic logic is straightforward. Building a frontier model requires enormous investment, and only a small group of companies can sustain that competition. Running models efficiently creates a different opportunity for cloud providers and infrastructure vendors.

The Next Web's coverage describes IBM's position as a bet that AI revenue increasingly comes from operating models cheaply, as well as developing more capable ones. IBM doesn't need this project to produce a model that beats Anthropic or Google. It needs the infrastructure and Together AI's serving software to deliver competitive performance at an attractive price.

The HGX B300 cluster supplies the hardware. Together AI supplies an existing platform and its inference optimization software. The partnership gives IBM a way to sell capacity for models customers have already chosen, rather than requiring those customers to adopt an IBM model.

That is a defensible approach if the partners can maintain utilization, service quality, and pricing discipline. The agreement establishes a substantial investment in that approach. It doesn't yet establish the economics of the finished cluster.

The recurring cost of managed APIs

AWS Bedrock, Azure OpenAI Service, and Google Vertex AI charge for more than the underlying compute. Customers also pay for a managed API, service-level commitments, integrations, and relief from running the infrastructure themselves. For an organization already deeply invested in one cloud, those benefits can justify a premium.

Inference costs nevertheless grow with use. Every additional automated workflow adds requests, input tokens, and generated output. A hypothetical customer-support summarization pipeline processing 50 million tokens a day would create a meaningful recurring bill. A dozen AI-supported workflows across a mid-size enterprise could make even a modest difference in per-token pricing worth examining.

The comparison has to account for useful output, not just the advertised token rate. A cheaper model may be a poor choice if it performs worse on the task. A more expensive API may remain economical when its integrations reduce operating work. The deal makes another option worth evaluating; it doesn't make every proprietary service overpriced.

Banks, hospitals, and other enterprises with sensitive data also have reasons to compare deployment arrangements. Open weights can support greater control because the model can be run in a chosen environment. The proposed arrangement puts that inference on Together AI's platform within IBM Cloud, potentially in a region that meets the customer's compliance requirements.

That deployment flexibility is useful, but access to weights alone doesn't establish how a hosted service handles customer data. Concerns about retention or use of data for training still require a review of the provider's terms and controls. A hosted open-model service shouldn't be assumed to keep data within a customer's required boundary solely because the weights are public.

Why the Ethernet choice deserves attention

The cluster will connect its HGX B300 systems using Spectrum-X, Nvidia's Ethernet-based networking for AI clusters, rather than InfiniBand. That is a relevant infrastructure choice at a planned scale of roughly 2,000 chips.

InfiniBand clusters can be expensive to build and operate, and tuning them requires specialist knowledge. Ethernet-based infrastructure is closer to the networking many enterprise infrastructure teams already manage. If it delivers the required inference performance, that familiarity could help reduce operating complexity.

Spectrum-X is purpose-built for AI networking, so the announcement shouldn't be read as proof that ordinary Ethernet equipment delivers the same results. Nor does the choice alone quantify savings against InfiniBand. It does show that the partners have selected an Ethernet-based design for a large inference deployment.

The broader expectation is that inference will take an increasing share of GPU work as more applications enter production. Together AI's reported token volume is consistent with strong demand for inference, but it doesn't establish whether inference already dominates GPU usage. The new cluster's cost and performance will matter more than the networking label alone.

What to compare for 2027 budgets

For enterprises planning 2027 AI spending, the project provides a concrete alternative to include in an infrastructure review: current Blackwell hardware, an established open-model serving platform, and IBM Cloud infrastructure, with launch targeted for the first quarter.

A useful comparison should cover:

  • Model performance on the actual task. An open-weight model needs to meet the application's quality requirements before a lower token price becomes useful.
  • Workload-level cost. The relevant figure is the cost of running the production workflow, including the operating work and integrations that a managed cloud service may already provide.
  • Service commitments and capacity. Pricing needs to be considered alongside the service-level agreement and access to enough capacity. Together AI's expected sellout is a reason to examine availability, not evidence that capacity is already contracted.
  • Deployment and data requirements. Region selection, model-version control, and data-handling terms need to satisfy the application's compliance obligations.

For existing proprietary API customers, a comparison with Together AI's current pricing could be useful before the next renewal. The result may still favor the incumbent. A stronger model or valuable cloud integration can justify the difference, but that decision is easier to defend with workload-specific numbers.

The Linux comparison is relevant as an analogy, not a guarantee. Open-weight models may follow a similar path from lower-cost alternatives with additional operating complexity to widely used infrastructure supported by a deeper ecosystem. IBM and Together AI are committing $240 million to capacity that could help that happen. For an enterprise choosing a platform, the deciding evidence will be model quality, delivered cost, and reliable service once the cluster is running.