Copied


Microsoft and NVIDIA Enhance AI Capabilities on RTX AI PCs

Caroline Bishop   Nov 19, 2024 09:04 0 Min Read


Microsoft and NVIDIA are taking significant strides in AI development with their latest collaboration, focusing on RTX AI PCs. This partnership aims to accelerate Windows app development by leveraging the power of NVIDIA's GeForce RTX GPUs, according to NVIDIA's blog. The advancements promise enhanced performance for over 600 Windows apps and games running AI locally on more than 100 million RTX AI PCs globally.

AI Tools for Enhanced App Development

At the recent Microsoft Ignite conference, NVIDIA and Microsoft unveiled new tools designed to facilitate the development of AI-powered applications on RTX AI PCs. These tools are set to aid developers in optimizing applications and games by utilizing the robust capabilities of RTX GPUs. The enhanced tools are particularly beneficial for creating AI agents, app assistants, and digital humans, making local AI more accessible to developers.

Digital Humans Powered by Multimodal Language Models

NVIDIA's digital human technologies, including the NVIDIA ACE suite, are advancing the realism of digital agents and avatars. These technologies enable digital humans to perceive and interact with their environment in a more human-like manner. Central to this development are multimodal small language models that process both text and imagery, thereby enhancing contextual understanding and response accuracy.

The upcoming NVIDIA Nemovision-4B-Instruct model, part of this initiative, utilizes the NVIDIA VILA and NeMo frameworks. These innovations enable the model to operate efficiently on RTX GPUs, providing developers with the precision needed for realistic digital human interactions. Additionally, the Mistral NeMo Minitron 128k Instruct family offers scalable options for efficient digital human interactions, supporting large datasets without data segmentation.

NVIDIA TensorRT Model Optimizer for Windows

Addressing the challenge of limited memory and compute resources in PCs, NVIDIA has updated its TensorRT Model Optimizer. This tool now allows Windows developers to optimize models for deployment in ONNX runtime environments using GPU execution providers like CUDA, TensorRT, and DirectML. The updates include advanced quantization algorithms that reduce the memory footprint and improve throughput performance on RTX GPUs.

These optimizations result in a significant reduction in memory usage, up to 2.6 times less than FP16 models, without sacrificing accuracy. This enhancement allows AI models to run effectively on a broader range of PCs, thereby expanding accessibility.

The collaboration between Microsoft and NVIDIA marks a significant milestone in making AI-driven applications more efficient and accessible, paving the way for future innovations in digital human interactions and AI technology.


Read More