One email each morning — yesterday's AI, sortedGet it in your inbox

METAL LAB

Together AI to Deploy 10,000 B300 GPUs for India's Largest AI Factory

Partnership with Larsen & Toubro to build infrastructure for open-source inference, fine-tuning, and training

이미지: METAL LAB 생성

Summary

  • Together AI announced it is building a cluster of 10,000 NVIDIA B300 GPUs together with Indian construction and engineering firm Larsen & Toubro
  • The cluster was introduced as the largest AI factory in India, intended to support inference, fine-tuning, and large-scale training of open-source models
  • The move continues a recent string of infrastructure expansion announcements from Together AI, including a B300 cluster build with IBM and NVIDIA and an $800 million Series C raise
발표 계정
Together AI (X 공식 계정, 2026-08-13)
파트너사
Larsen & Toubro (인도 건설·엔지니어링 기업)
GPU 규모
NVIDIA B300 1만 장
지역 규모 표현
인도 최대(India's largest) AI 팩토리로 소개
용도
오픈소스 모델 추론·파인튜닝·대규모 훈련 지원
최근 관련 발표
Together AI-IBM-NVIDIA, IBM Cloud B300 전용 클러스터 구축 (2026-08-12)
최근 투자 유치
Together AI, 8억 달러 규모 시리즈 C 유치 (지난달)

10,000 GPUs, Billed as India's Largest

Together AI announced it is building a cluster of 10,000 NVIDIA B300 GPUs together with Indian construction and engineering firm Larsen & Toubro. Together AI described it as "India's largest AI factory," explaining that the infrastructure will be used to support inference, fine-tuning, and large-scale training of open-source models.

The B300 is NVIDIA's latest-generation data center GPU. A scale of 10,000 GPUs is large enough to train large language models from scratch or handle inference requests from multiple companies simultaneously — an investment level that individual startups typically cannot afford on their own. Such clusters are commonly referred to as "AI factories," meaning large-scale computing facilities designed end-to-end, including power, cooling, and networking.

Together AI's String of Infrastructure Expansions

This announcement extends a series of infrastructure expansions Together AI has made over the past month. On August 12, Together AI said it was partnering with IBM and NVIDIA to build a configuration combining a dedicated B300 cluster with Spectrum-X networking on IBM Cloud — reportedly the first deployment of its kind on IBM Cloud. Prior to that, Together AI had raised $800 million in a Series C round last month.

In a short span, the company has announced major cluster builds in both the United States (IBM Cloud) and India (Larsen & Toubro). For Together AI, which hosts open-source models and provides inference services, securing data centers across regions directly translates into shorter service latency and stronger customer acquisition.

Compared to Other Regional AI Factories

Around the same period, announcements of large-scale GPU cluster builds have been emerging from multiple regions. On August 8, NVIDIA said AI cloud company Firebird brought online what it called the largest AI factory in the CIS (Commonwealth of Independent States) region, located in Hrazdan, Armenia. Firebird plans to deploy more than 70,000 NVIDIA Rubin and Blackwell GPUs and build out 300 megawatts of infrastructure by the end of 2027.

ProjectPartnersLocationGPU Scale
Firebird AI FactoryFirebirdHrazdan, ArmeniaRubin/Blackwell 70,000+ (targeted for 2027)
Together AI-IBM ClusterTogether AI, IBM, NVIDIAIBM CloudDedicated B300 cluster
Together AI-L&T ClusterTogether AI, Larsen & ToubroIndia10,000 B300 GPUs

What each region is claiming is the title of "largest within its own country or cloud." Though the stages differ — India, Armenia, and IBM Cloud in the US — they share a common goal: processing inference and training for open-source models on domestic or in-house infrastructure.

So What Does This Change

The arrival of large-scale computing infrastructure for open-source models in India means more startups and companies there will have the option to run large-scale fine-tuning or inference services without going through overseas clouds. For Together AI, having major clusters in both the US and India simultaneously expands both its processing capacity and its regional reach in the open-source model hosting market. Specific operational details, such as exact GPU counts or launch timing, have not yet been further confirmed.