Enhancing Large Language Model Performance through Sketching Techniques

Authors

  • Junyin Zhang
  • Zecen Ding
  • Yitian Wan
  • Yuheng Shen

DOI:

https://doi.org/10.61173/yg8q1674

Keywords:

Large Language Model (LLM), Sketching, PolySketchFormer, Prompt Sketching

Abstract

This article explores the theoretical application of sketching techniques to Large Language Models (LLMs), which use deep learning and extensive datasets for natural language processing tasks. The study summarizes two approaches: PolySketchFormer and Prompt Sketching. PolySketchFormer accelerates transformer models using sketching for performance optimization, while Prompt Sketching aims to enhance model accuracy. The article delves into the theory and processes of these methods, highlighting their advantages and potential implications for advancing LLM capabilities.

References

[1] Naveed, H., Khan, A. U., Qiu, S., Saqib, M., Anwar, S., Usman, M., Akhtar, N., Barnes, N., & Mian, A. (2024). A Comprehensive Overview of Large Language Models (arXiv:2307.06435). arXiv. http://arxiv.org/abs/2307.06435

[2] Woodruff, D. P. (2014). Sketching as a Tool for Numerical Linear Algebra. Foundations and Trends® in Theoretical Computer Science, 10(1–2), 1–157. https://doi. org/10.1561/0400000060

[3] Kacham, P., Mirrokni, V., & Zhong, P. (2024). PolySketchFormer: Fast Transformers via Sketching Polynomial Kernels (arXiv:2310.01655). arXiv. http://arxiv.org/ abs/2310.01655

[4] Beurer-Kellner, L., Fischer, M., & Vechev, M. (2023). Prompting Is Programming: A Query Language for Large Language Models. Proceedings of the ACM on Programming Languages, 7(PLDI), 1946–1969. https://doi. org/10.1145/3591300

[5] Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D., Wu, J., Winter, C., … Amodei, D. (2020). Language Models are Few-Shot Learners. Advances in Neural Information Processing Systems, 33, 1877–1901. https:// proceedings.neurips.cc/paper/2020/hash/1457c0d6bfcb4967418 bfb8ac142f64a-Abstract.html

[6] Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., & Amodei, D. (2020). Scaling Laws for Neural Language Models (arXiv:2001.08361). arXiv. http://arxiv.org/abs/2001.08361

[7] Charikar, M., Chen, K., & Farach-Colton, M. (n.d.). Finding Frequent Items in Data Streams.

[8] Reynolds L., McDonell K. Prompt Programming for Large Language Models: Beyond the Few-Shot Paradigm | Extended Abstracts of the 2021 CHI Conference on Human Factors in Computing Systems. (n.d.). Retrieved July 13, 2024, from https:// dl.acm.org/doi/10.1145/3411763.3451760#core-collateralpurchase-access

[9] Post, M., & Vilar, D. (2018). Fast Lexically Constrained Decoding with Dynamic Beam Allocation for Neural Machine Translation. In M. Walker, H. Ji, & A. Stent (Eds.), Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers) (pp. 1314–1324). Association for Computational Linguistics. https://doi. org/10.18653/v1/N18-1119

[10] Ling, W., Yogatama, D., Dyer, C., & Blunsom, P. (2017). Program Induction by Rationale Generation: Learning to Solve and Explain Algebraic Word Problems. In R. Barzilay & M.-Y. Kan (Eds.), Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (pp. 158–167). Association for Computational Linguistics. https://doi.org/10.18653/v1/P17-1015

Downloads

Published

2025-07-06