Copied


NVIDIA Expands AI Capabilities with NIM Integration on KServe

Rebeca Moen   Jun 01, 2024 20:22 3 Min Read


Deploying generative AI in the enterprise is about to get easier than ever. NVIDIA NIM, a set of generative AI inference microservices, will work with KServe, open-source software that automates putting AI models to work at the scale of a cloud computing application, according to NVIDIA Blog.

The combination ensures generative AI can be deployed like any other large enterprise application. It also makes NIM widely available through platforms from dozens of companies, such as Canonical, Nutanix, and Red Hat.

Serving AI on Kubernetes

KServe originated from Kubeflow, a machine learning toolkit based on Kubernetes, the open-source system for deploying and managing software containers that hold all the components of large distributed applications. As Kubeflow expanded its work on AI inference, KServe evolved into its own open-source project.

Many companies have contributed to and adopted KServe, including AWS, Bloomberg, Canonical, Cisco, Hewlett Packard Enterprise, IBM, Red Hat, Zillow, and NVIDIA.

Under the Hood With KServe

KServe is essentially an extension of Kubernetes that runs AI inference like a powerful cloud application. It uses a standard protocol, runs with optimized performance, and supports PyTorch, Scikit-learn, TensorFlow, and XGBoost without users needing to know the details of those AI frameworks.

The software is particularly useful as new large language models (LLMs) emerge rapidly. KServe allows users to switch between models easily, testing which one best suits their needs. A feature called “canary rollouts” automates the gradual deployment of updated model versions into production.

Another feature, GPU autoscaling, efficiently manages how models are deployed as demand for a service changes, ensuring optimal performance for customers and service providers.

An API Call to Generative AI

The integration of KServe with NVIDIA NIM makes generative AI deployment as simple as an API call. Enterprise IT administrators get the metrics they need to ensure their applications run optimally, whether in their data centers or on remote cloud services, even if they change the AI models they use.

NIM lets IT professionals become generative AI pros, transforming their company’s operations. Enterprises such as Foxconn and ServiceNow are deploying NIM microservices.

NIM Rides Dozens of Kubernetes Platforms

Thanks to its integration with KServe, users will be able to access NIM on dozens of enterprise platforms such as Canonical’s Charmed KubeFlow and Charmed Kubernetes, Nutanix GPT-in-a-Box 2.0, Red Hat’s OpenShift AI, and many others.

“Red Hat has been working with NVIDIA to make it easier for enterprises to deploy AI using open-source technologies,” said KServe contributor Yuan Tang, a principal software engineer at Red Hat. “By enhancing KServe and adding support for NIM in Red Hat OpenShift AI, we’re able to provide streamlined access to NVIDIA’s generative AI platform for Red Hat customers.”

Debojyoti Dutta, vice president of engineering at Nutanix, stated, “Through the integration of NVIDIA NIM inference microservices with Nutanix GPT-in-a-Box 2.0, customers can build scalable, secure, high-performance generative AI applications in a consistent way, from the cloud to the edge.”

Andreea Munteanu, MLOps product manager at Canonical, also commented, “As a company that contributes significantly to KServe, we’re pleased to offer NIM through Charmed Kubernetes and Charmed Kubeflow. Users will be able to access the full power of generative AI, with the highest performance, efficiency, and ease thanks to the combination of our efforts.”

Dozens of other software providers can benefit from NIM simply because they include KServe in their offerings.

Serving the Open-Source Community

NVIDIA has a long track record on the KServe project. KServe’s Open Inference Protocol is used in NVIDIA Triton Inference Server, which helps users run many AI models simultaneously across many GPUs, frameworks, and operating modes.

With KServe, NVIDIA focuses on use cases that involve running one AI model at a time across many GPUs. As part of the NIM integration, NVIDIA plans to be an active contributor to KServe, building on its portfolio of contributions to open-source software, which includes Triton and TensorRT-LLM. NVIDIA is also an active member of the Cloud Native Computing Foundation, which supports open-source code for generative AI and other projects.

Try the NIM API on the NVIDIA API Catalog using the Llama 3 8B or Llama 3 70B LLM models today. Hundreds of NVIDIA partners worldwide are using NIM to deploy generative AI.

Watch NVIDIA founder and CEO Jensen Huang’s COMPUTEX keynote to get the latest on AI and more.


Read More