The Study of Machine Learning Inference Tasks Based on Serverless Computing Platforms
DOI:
https://doi.org/10.61173/7mzsgf91Keywords:
Serverless, ResNet50, Model Partitioning, Parallel Execution, Inference EfficiencyAbstract
This study proposes a distributed inference method based on the ResNet50 model, aiming to improve inference efficiency and resource utilization by dividing the model into multiple sub-models. Specifically, the model is divided into the initial convolutional layer, four stages of residual blocks, and the subsequent global average pooling and fully connected layers. Each sub-model independently handles specific tasks, allowing for parallel execution on different computing devices, thereby accelerating the overall inference process. This partitioning strategy effectively addresses high-concurrency requests, enhancing the system‘s response speed. Additionally, it enables dynamic scaling of resources based on workload demands, which is crucial for real-time applications. The implementation of distributed inference also makes the model more flexible, adapting to various computing resources and application scenarios. Experimental results indicate that this method significantly enhances inference efficiency while maintaining model performance, providing new ideas and solutions for the practical application of deep learning models. These findings underscore the potential of distributed architectures in advancing the deployment of complex neural networks across diverse environments.References
[1] Zhang C, Yu M, Wang W, et al. Enabling cost-effective, slow machine learning inference serving on public cloud. IEEE Transactions on Cloud Computing, 2020, 10(3): 1765-1779.
[2] Bhattacharjee A. Algorithms and Techniques for Automated Deployment and Efficient Management of Large-Scale Distributed Data Analytics Services. Vanderbilt University, 2020.
[3] Wen X, Zeng T, Li C, et al. Research on Model Inference Service Switching Method for Serverless Computing. Computer Engineering and Science, 2024, 46(07): 1210.
[4] Chaitanya K T. EXPLORING SERVER-LESS COMPUTING FOR EFFICIENT RESOURCE MANAGEMENT IN CLOUD ARCHITECTURES. Journal of Science Technology and Research (JSTAR), 2023, 4 (1):77-83.
[5] Mampage A, Karunasekera S, Buyya R. A holistic view on resource management in serverless computing environments: Dean&Francis ISSN 2959-6157 Taxonomy and future directions. ACM Computing Surveys (CSUR), 2022, 54(11s): 1-36.
[6] Shafiei H, Khonsari A, Mousavi P. Serverless computing: a survey of opportunities, challenges, and applications. ACM Computing Surveys, 2022, 54(11s): 1-32.
[7] Yu M, Jiang Z, Ng H C, et al. Gillis: Serving large neural networks in serverless functions with automatic model partitioning//2021 IEEE 41st International Conference on Distributed Computing Systems (ICDCS). IEEE, 2021: 138-148.
[8] Hassan H B, Barakat S A, Sarhan Q I. Survey on serverless computing. Journal of Cloud Computing, 2021, 10: 1-29.
[9] Koonce B. ResNet 50. In: Convolutional Neural Networks with Swift for Tensorflow. Apress, Berkeley, CA. https://doi. org/10.1007/978-1-4842-6168-2_6, 2021.
Downloads
Published
Issue
Section
License
Copyright (c) 2024 by the authors.

This work is licensed under a Creative Commons Attribution 4.0 International License.
