Next Article in Journal
InfoMax Classification-Enhanced Learnable Network for Few-Shot Node Classification
Previous Article in Journal
A Deep Learning-Based Phishing Detection System Using CNN, LSTM, and LSTM-CNN
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

HetSev: Exploiting Heterogeneity-Aware Autoscaling and Resource-Efficient Scheduling for Cost-Effective Machine-Learning Model Serving

1
State Key Laboratory of Media Convergence and Communication, Communication University of China, Beijing 100024, China
2
Beijng Key Laboratory of Big Data in Security & Protection Industry, Beijing 100024, China
*
Authors to whom correspondence should be addressed.
Electronics 2023, 12(1), 240; https://doi.org/10.3390/electronics12010240
Submission received: 25 November 2022 / Revised: 16 December 2022 / Accepted: 26 December 2022 / Published: 3 January 2023

Abstract

To accelerate the inference of machine-learning (ML) model serving, clusters of machines require the use of expensive hardware accelerators (e.g., GPUs) to reduce execution time. Advanced inference serving systems are needed to satisfy latency service-level objectives (SLOs) in a cost-effective manner. Novel autoscaling mechanisms that greedily minimize the number of service instances while ensuring SLO compliance are helpful. However, we find that it is not adequate to guarantee cost effectiveness across heterogeneous GPU hardware, and this does not maximize resource utilization. In this paper, we propose HetSev to address these challenges by incorporating heterogeneity-aware autoscaling and resource-efficient scheduling to achieve cost effectiveness. We develop an autoscaling mechanism which accounts for SLO compliance and GPU heterogeneity, thus provisioning the appropriate type and number of instances to guarantee cost effectiveness. We leverage multi-tenant inference to improve GPU resource utilization, while alleviating inter-tenant interference by avoiding the co-location of identical ML instances on the same GPU during placement decisions. HetSev is integrated into Kubernetes and deployed onto a heterogeneous GPU cluster. We evaluated the performance of HetSev using several representative ML models. Compared with default Kubernetes, HetSev reduces resource cost by up to 2.15× while meeting SLO requirements.
Keywords: inference serving; autoscaling; cost effectiveness; multi-tenant inference inference serving; autoscaling; cost effectiveness; multi-tenant inference

Share and Cite

MDPI and ACS Style

Mo, H.; Zhu, L.; Shi, L.; Tan, S.; Wang, S. HetSev: Exploiting Heterogeneity-Aware Autoscaling and Resource-Efficient Scheduling for Cost-Effective Machine-Learning Model Serving. Electronics 2023, 12, 240. https://doi.org/10.3390/electronics12010240

AMA Style

Mo H, Zhu L, Shi L, Tan S, Wang S. HetSev: Exploiting Heterogeneity-Aware Autoscaling and Resource-Efficient Scheduling for Cost-Effective Machine-Learning Model Serving. Electronics. 2023; 12(1):240. https://doi.org/10.3390/electronics12010240

Chicago/Turabian Style

Mo, Hao, Ligu Zhu, Lei Shi, Songfu Tan, and Suping Wang. 2023. "HetSev: Exploiting Heterogeneity-Aware Autoscaling and Resource-Efficient Scheduling for Cost-Effective Machine-Learning Model Serving" Electronics 12, no. 1: 240. https://doi.org/10.3390/electronics12010240

APA Style

Mo, H., Zhu, L., Shi, L., Tan, S., & Wang, S. (2023). HetSev: Exploiting Heterogeneity-Aware Autoscaling and Resource-Efficient Scheduling for Cost-Effective Machine-Learning Model Serving. Electronics, 12(1), 240. https://doi.org/10.3390/electronics12010240

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop