Together AI and Hypertec Cloud Partner to Build Massive GPU Cluster with NVIDIA GB200
Together AI has announced a strategic partnership with Hypertec Cloud to construct one of the world's largest GPU clusters, featuring 36,000 NVIDIA GB200 NVL72 GPUs. This collaboration aims to provide next-generation infrastructure designed to accelerate AI development, according to Together AI.
Enhancing AI Infrastructure
The partnership leverages Together AI’s expertise in high-performance GPU clusters and AI research with Hypertec Cloud’s data center capabilities. This initiative is set to optimize resource usage, reduce operational costs, and empower companies to push the boundaries of AI without compromising performance, cost, or reliability.
The infrastructure, designed by AI researchers, incorporates advanced technology such as Blackwell Tensor Core GPUs, Grace CPUs, NVIDIA’s NVLink, and the Together Kernel Collection. These innovations aim to enhance performance, scalability, and cost efficiency by optimizing GPU usage, thus enabling efficient scaling for AI workloads.
Immediate and Long-Term Benefits
Slated to commence in Q1 2025, the GPU cluster will offer immediate access to thousands of H100 and H200 GPUs across North America. The infrastructure provides secured data center capacity for over 100,000 GPUs throughout 2025, delivering industry-leading deployment times and reliability.
Jonathan Ahdoot, President of Hypertec Cloud, expressed excitement about the collaboration, emphasizing the delivery of high-performance AI solutions that are both efficient and powerful, while also minimizing environmental impact.
Together AI's Proprietary Optimization
Central to this infrastructure is the Together Kernel Collection, a proprietary suite of optimizations that significantly enhance AI computation. This suite delivers up to 24% faster training operations and a 75% boost in FP8 inference tasks, translating into reduced GPU hours and operational costs for users.
Tests conducted on the Black Forest Labs’ Flux model demonstrate substantial efficiency gains, with text-to-image generation time reduced significantly when utilizing the Together Kernel Collection.
Global Scale and AI Expertise
Together GPU Clusters offer a flexible AI infrastructure solution that can be configured with a range of NVIDIA GPUs, catering to the needs of both startups and large enterprises. High-speed interconnects like InfiniBand and NVLink ensure low-latency communication, accelerating model training and convergence.
The platform also provides AI-native storage solutions and advisory services to help maximize model performance, offering dedicated support for kernel customization and system optimization.
NVIDIA’s GB200 NVL72 Power
The NVIDIA GB200 NVL72 adds significant capability to the Together GPU Clusters, promising 30X faster real-time inference for trillion-parameter models and up to 4X accelerated training. This performance is attributed to its advanced Transformer Engine and FP4 precision capabilities, alongside fifth-generation NVLink, creating a cohesive GPU powerhouse ideal for high-demand models.
Through this partnership, Together AI and Hypertec Cloud are poised to meet the growing demands of frontier AI, offering reliability and scalability for industry innovators and large-scale enterprises.