NVIDIA NIM Boosts Nemotron 3 Ultra Throughput by 2.5x
NVIDIA has unveiled a significant upgrade to its AI inference microservices, claiming its NIM 2.0.12 optimization stack delivers a 2.5x increase in throughput on the Nemotron 3 Ultra system. This improvement allows enterprises to serve more concurrent users on the same GPU infrastructure without compromising latency.
The optimization, detailed in a technical blog published on September 10, 2026, focuses on agentic AI workloads—applications that require extended context reuse and high interactivity. Using NVIDIA's NIM (NVIDIA Inference Microservices), which packages model- and GPU-aware serving configurations, developers can deploy AI applications with validated configurations that meet production-grade performance targets.
Key Performance Gains
Benchmarks of the Nemotron 3 Ultra system, running on four NVIDIA B200 GPUs, highlight the impact of NIM optimizations:
- NIM-off baseline throughput: 718 tokens per second (tok/s).
- NIM 2.0.12 optimized stack: 1,997 tok/s, a 2.5x improvement.
This performance boost is achieved through a combination of precision tuning, parallel execution, memory optimization, and speculative decoding strategies. For instance, MTP (Mixture-of-Experts Tensor Parallelism) speculative decoding enhances token generation speed while maintaining low latency targets, critical for applications like chat interfaces and AI copilots.
Implications for Enterprises
The enhanced throughput directly translates to increased scalability, enabling enterprises to serve up to 2.5 times more users at 50 transactions per second (TPS) per user. This is particularly valuable as demand for generative AI applications continues to grow, with businesses deploying solutions across customer service, content generation, and advanced analytics.
By integrating these prepackaged inference services, companies can reduce deployment complexity and lower operational costs. NVIDIA has also included commercial support and hardware validation through its AI Enterprise suite, ensuring reliability for production environments.
How to Benchmark NIM
NVIDIA encourages developers to test NIM’s performance on their workloads using AIPerf, its proprietary benchmarking tool. By replaying representative traffic, enterprises can determine optimal configurations tailored to their latency and throughput requirements. The company emphasizes that while the published benchmarks provide a strong starting point, real-world performance will vary based on workload specifics.
Market Context
NVIDIA’s push for full-stack optimization aligns with its broader strategy to dominate the AI infrastructure market. As of September 2026, NVIDIA maintains a market cap of $5.299 trillion, underscoring its leadership in AI hardware and software solutions. Recent moves, such as its $2 billion partnership with Coherent to upgrade its AI infrastructure, further cement its position.
This announcement comes as enterprises increasingly prioritize scalable, cost-efficient AI deployments. With new liquid-cooling innovations announced earlier this year to cut data-center energy costs, NVIDIA is addressing both performance and sustainability concerns—a dual focus that resonates with enterprise buyers.
Looking Ahead
For developers and enterprises looking to leverage these advancements, the Nemotron 3 Ultra NIM 2.0.12 is now available for download. NVIDIA plans to expand its library of performance-optimized configurations across a broader range of models, offering more tailored solutions for diverse use cases.
In a market where performance and scalability are critical to staying competitive, NVIDIA’s NIM 2.0.12 positions the company—and its customers—one step ahead in deploying production-ready AI systems.