Computational Optimization Nash-MTL: An Efficient Multi-Task Learning Method Based on Gradient Intelligent Sampling
DOI:
https://doi.org/10.61173/2f0wr760Keywords:
Multi-task learning, Gradient optimization, Nash bargaining solution, Computation OptimizationAbstract
Aiming at the exponential computational complexity problem of Nash-MTL in multi-task learning, this paper proposes a computationally optimized Nash-MTL framework. This method introduces three core points. First, there is a phased gradient update mechanism, which combines cyclic sampling and dynamic random sampling strategies, and can maintain the optimized performance while To minimize redundant gradient computing, secondly, there is a dynamic importance scheduling model, which assesses task priority by means of loss change rate and gradient size, thereby intelligently allocating computing resources, and is also supplemented by a security recovery strategy. Thirdly, there is a stability guarantee mechanism with periodic global updates and abnormal trigger rollback operations. Experiments were conducted on QM9, NYUv2 and cityscape datasets, which confirmed the effectiveness of the framework. The framework can maintain task performance (with a deviation within 5%) while significantly reducing computing time by 55.4%. These advancements greatly enhance the feasibility of deploying complex multi-task learning systems in resource-constrained edge computing environments.
References
[1] Yu, T., Kumar, S., Gupta, A., Levine, S., Hausman, K., & Finn, C. (2020). Gradient surgery for multi-task learning. Advances in Neural Information Processing Systems, *33*, 5824-5836.
[2] Navon, A., et al. (2022). Multi-task learning as a bargaining game. Proceedings of the 39th International Conference on Machine Learning (ICML), 162, 112-125.
[3] Zhou, J., Zhang, Y., Wang, X., & Liu, Y. (2023). Joint multi-task offloading and resource allocation for mobile edge computing systems in satellite IoT. IEEE Transactions on Vehicular Technology, 72(5), 6789-6802.
[4] Liu, C., Hoi, S. C. H., Zhao, P., & Sun, J. (2022). Stable gradient directions for robust multi-task learning. Neural Networks, 156, 1-12.
[5] Kendall, A., Gal, Y., & Cipolla, R. (2018). Multi-task learning using uncertainty to weigh losses for scene geometry and semantics. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 7482-7491.
[6] Zhang, J., He, T., Sra, S., & Jadbabaie, A. (2020). Why gradient clipping accelerates training: A theoretical justification for adaptivity. Proceedings of the International Conference on Learning Representations (ICLR 2020).
[7] Zhao, J., Wu, D., Wu, J.-J., Ye, W., Huang, F., Wang, J., & See-To, E. W. K. (2024). Consistency approximation: Incremental feature selection based on fuzzy rough set theory. Pattern Recognition, 155, 110652.
[8] Zhou, J., Ye, K., Liu, J., Ma, T., Wang, Z., Qiu, R., Lin, K.-Y., Zhao, Z., & Liang, J. (2025). Exploring the limits of vision-language-action manipulations in cross-task generalization. ICLR 2025 Conference Proceedings.
[9] Khanna, D., Guru, A., et al. (2025). QuickSilver: Speeding up LLM Inference through Dynamic Token Halting, KV Skipping, Contextual Token Fusion, and Adaptive Matryoshka Quantization. Proceedings of the ACM Web Conference 2025 (WWW ‚25). ACM.
[10] Alistarh, D., Grubic, D., Li, J., Tomioka, R., & Vojnovic, M. (2017). Gradient sparsification for communication-efficient distributed optimization. Proceedings of the 34th International Conference on Machine Learning (ICML 2017), 597–606.
[11] Cao, T., Liu, M., & Zhang, T. (2021). The scheduling problem with cyclic time windows on machines. Advances in Applied Mathematics, *10*(2), 42-46.
[12] He, Y., Feng, X., Cheng, C., Ji, G., Guo, Y., & Caverlee, J. (2022). MetaBalance: Improving Multi-Task Recommendations via Adapting Gradient Magnitudes of Auxiliary Tasks. In Proceedings of the ACM Web Conference 2022 (pp. 2205– 2214). ACM.
[13] Wang, Z., Tsvetkov, Y., Firat, O., & Cao, Y. (2021). Gradient vaccine: Investigating and improving multi-task optimization in massively multilingual models. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 14609-14619.
[14] Wu, H., Luo, H., Ma, Y., Wang, J., & Long, M. (2024). RoPINN: Region Optimized Physics-Informed Neural Networks. Advances in Neural Information Processing Systems 37 (NeurIPS 2024).
Downloads
Published
Issue
Section
License
Copyright (c) 2025 by the authors.

This work is licensed under a Creative Commons Attribution 4.0 International License.
