Copied


Together AI Unveils Serverless Multi-LoRA for Scalable Model Customization

Tony Kim   Dec 18, 2024 18:10 0 Min Read


Together AI has announced the launch of Serverless Multi-LoRA, a new feature that enables the fine-tuning and deployment of hundreds of custom Low-Rank Adaptation (LoRA) adapters. This development promises to revolutionize model customization by reducing costs and complexity associated with multiple fine-tuned models, according to together.ai.

Innovative Serverless LoRA Inference

The Serverless Multi-LoRA feature allows users to upload their own LoRA adapters and run inference on them with compatible models such as Llama 3.1 and Qwen 2.5. This is done using a pay-per-token pricing model, which significantly cuts down on infrastructure costs. The platform also supports dynamic adapter switching, enabling the operation of numerous models for the same cost as a single base model.

Fine-Tuning API for Seamless Integration

The new LoRA fine-tuning API allows users to test and deploy fine-tuned adapters through Together AI's playground or APIs. This functionality supports several base models and provides the flexibility to download adapters, making it easier for companies to bring their models from experimentation to production efficiently.

LoRA: Efficient Fine-Tuning

LoRA is recognized as an efficient fine-tuning method that creates lightweight adapters without modifying the entire model's weights. This reduces memory requirements and infrastructure costs, allowing for the use of a single base model with task-specific adapters. This flexibility is particularly beneficial for industries requiring diverse use cases, such as marketing and enterprise operations.

The Potential of Multi-LoRA

Multi-LoRA technology enables the serving of multiple AI adapters with a single base model, significantly reducing the need for separate infrastructure for each model. This leads to substantial cost savings and facilitates rapid experimentation. By automating memory and batching configurations, Together Serverless simplifies the process of managing GPU resources and adapter swapping.

Advantages of Using Together AI

Cost-Efficiency

With Together AI's serverless infrastructure, users can run hundreds of custom adapters at the same cost as a single base model, only paying per token used. This eliminates unnecessary spending on idle infrastructure.

Optimized Performance

Together AI's optimized serving system maintains high performance while offering flexible per-token pricing. Innovations like Together FlashAttention 3 and Cross-LoRA Continuous Batching enhance GPU utilization and adapter prefetching, ensuring efficient adapter serving.

Getting Started with LoRA on Together AI

To leverage the benefits of Multi-LoRA, users can begin by fine-tuning and deploying custom model adapters, taking advantage of the platform's cost-efficient and high-performance capabilities.


Read More