An Experiment of Implementing Reservoir Sampling to StreamingLLM

Authors

  • Jiyang Pan
  • Jiayi Weng
  • Yansen Huang
  • Patrick Pan
  • Qianbiao Zhao

DOI:

https://doi.org/10.61173/ecst9g72

Keywords:

streamingLLM, Reservoir Sampling, Large Language Models, KV-cache, tokens

Abstract

When the text is longer than the training sequence length, the streamingLLM can successfully improve the computational speed and guarantee a certain degree of accuracy. But this method is only suitable for short term memory questions and answers. Because the StreamingLLM doesn’t improve a computer’s Long-term memory ability. We tried to combine the reservoir sampling with streamingLLM. Since StreamingLLM, the reservoir sampling will randomly take samples from what were meant to be discarded. Through adding reservoir sampling, we find the results are more accurate and representative.

References

[1] Praneeth Kacham, Vahab Mirrokni, Peilin Zhong, (2024, Mar17). PolySketchFormer: Fast Transformers via Sketching Polynomial Kernels, arXiv:2310.01655v3 [cs.LG]

[2] Ali Jamali, Swalpa Kumar Roy, Avik Bhattacharya, Pedram Ghamisi, 2023, Local Window Attention Transformer for Polarimetric SAR Image Classification, IEEE Geoscience and Remote Sensing Letters, DOI: 10.1109/LGRS.2023.3239263

[3] Muhammad Adnan, Akhil Arunkumar, Gaurav Jain, Prashant Nair, Ilya Soloveychik, Purushotham Kamath, 2024, Keyformer: KV Cache reduction through key tokens selection for Efficient Generative Inference, Proceedings of Machine Learning and Systems 6(MLSys2024) Conference

[4] Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya,Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, Tom Hennigan, Eric Noland, Katie Millican, George van den Driessche, Bogdan Damoc, Aurelia Guy, Simon Osindero, Karen Simonyan, Erich Elsen, Jack W. Rae, Oriol Vinyals, Laurent Sifre, 2022, Training Compute-Optimal Large Language Models, arXiv: 2203.15556 [cs.CL]

[5] Richard Startin, 2020, Reservoir Sampling, Reservoir Sampling | Richard Startin’s Blog

[6] Rajesh Jayaram, Gokarna Sharma, Srikanta Tirthapura, David P. Woodruf, 2019, Weighted Reservoir Sampling from Distributed Streams, Proceedings of the 38th ACM SIGMOD- SIGACT-SIGAI Symposium on Principles of Database Systems, Pages 218 - 235

[7] Han, I., Jayaram, R., Karbasi, A., Mirrokni, V., Woodruff, D. P., & Zandieh, A. (2023, October 9). HyperAttention: Longcontext attention in Near-Linear time. arXiv.org. https://arxiv. org/abs/2310.05869

[8] Xiao, G., Tian, Y., Chen, B., Han, S., & Lewis, M. (2023, September 29). Efficient Streaming Language Models with Attention Sinks. arXiv.org. https://arxiv.org/abs/2309.17453

[9] Mohammed Al-Kateb, Byung Suk Lee, X. Sean Wang, 2007, Adaptive-Size Reservoir Sampling over Data Streams, 19th International Conference on Scientific and Statistical Database Management (SSDBM 2007)

[10] Jagbir Kaur, 2023, STREAMING DATA ANALYTICS: CHALLENGES AND OPPORTUNITIES, International Journal of Applied Engineering & Technology, ISSN: 2633-4828

Downloads

Published

2025-07-06