Challenges and Improvements in Mathematical Reasoning for Large Models

Authors

  • Ziming Wang

DOI:

https://doi.org/10.61173/gbgrkk38

Keywords:

Mathematical reasoning, LLMs, Chain-of-Thought, self-verification, AI Benchmarking

Abstract

Mathematical reasoning is a fundamental criterion for AI, and this is also a problem to be solved before arriving at AGI. Large Language Models(LLMs) have become prominent in this domain.. LLM’s benchmark performance (GSM8K, MATH): However, a large number of investigation indicate that they only mimic what the training data are, so they do not properly understand logical problems, and that means their reasoning ability is extraordinarily flaky, meaning that even a small perturbation in how problems are phrased could completely alter the reasoning which is really the difference between being smart statistically and understanding. So, if you are going to try and reason around these sorts of limitations, to find a way to a more robust application of math intelligence, you need to consider all research directions being explored right now. Analysis is broken down into 3 directions:(1) The first direction involves methods that include tweaking methods from a processing intervention standpoint towards reasoning chains. (2). The second direction is placeholder self-verification and reflection, where the models are urged to think about and improve their own reasoning. (3). The third direction is scientific verification and a measure of efficiency, focusing on thorough benchmarking and ways to test like tool integration. This survey is meaningful, as it is on a topic that is ever-changing, and will provide a baseline to help a lay audience, in addition to showcasing what the state-of-the-art approaches are, while still critically examining whether it can overcome any of the inherent challenges associated with LLMs. I hope it can be useful in conducting research in the future for more logical/more trustworthy AIs.

References

[1] Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez A N, et al. Attention is all you need. Advances in Neural Information Processing Systems, 2017, 30.

[2] Kahneman D. Thinking, fast and slow. Macmillan, 2011.

[3] Bubeck S, Chandrasekaran V, Eldan R, Gehrke J, Horvitz E, Kamar E, et al. Sparks of artificial general intelligence: Early experiments with GPT-4. arXiv:2303.12712, 2023.

[4] Wei J, Wang X, Schuurmans D, Bosma M, Xia F, Chi E, et al. Chain-of-thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems, 2022, 35: 24824–24837.

[5] Lightman H, Kosaraju V, Burda Y, Edwards H, Baker B, Lee T, et al. Let’s verify step by step. The Twelfth International Conference on Learning Representations, May 2023.

[6] Ouyang L, Wu J, Jiang X, Almeida D, Wainwright C, Mishkin P, et al. Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 2022, 35: 27730–27744.

[7] Wang X, Wei J, Schuurmans D, Le Q, Chi E, Narang S, et al. Self-consistency improves chain of thought reasoning in language models. arXiv:2203.11171, 2022.

[8] Yao S, Yu D, Zhao J, Shafran I, Griffiths T, Cao Y, Narasimhan K. Tree of thoughts: Deliberate problem solving with large language models. Advances in Neural Information Processing Systems, 2023, 36: 11809–11822.

[9] Cobbe K, Kosaraju V, Bavarian M, Chen M, Jun H, Kaiser Dean&Francis Ziming Wang L, et al. Training verifiers to solve math word problems. arXiv:2110.14168, 2021.

[10] Hendrycks D, Burns C, Kadavath S, Arora A, Basart S, Tang E, et al. Measuring mathematical problem solving with the math dataset. arXiv:2103.03874, 2021.

[11] Zheng Z, Ning K, Wang Y, Zhang J, Zheng D, Ye M, Chen J. A survey of large language models for code: Evolution, benchmarking, and future trends. arXiv:2311.10372, 2023.

[12] Schick T, Dwivedi-Yu J, Dessì R, Raileanu R, Lomeli M, Hambro E, et al. Toolformer: Language models can teach themselves to use tools. Advances in Neural Information Processing Systems, 2023, 36: 68539–68551.

[13] Gao L, Madaan A, Zhou S, Alon U, Liu P, Yang Y, et al. PAL: Program-aided language models. International Conference on Machine Learning, July 2023: 10764–10799.

[14] Zhou Z, Ning X, Hong K, Fu T, Xu J, Li S, et al. A survey on efficient inference for large language models. arXiv:2404.14294, 2024.

[15] Olausson T X, Gu A, Lipkin B, Zhang C E, Solar-Lezama A, Tenenbaum J B, Levy R. LINC: A neurosymbolic approach for logical reasoning by combining language models with first-order logic provers. arXiv:2310.15164, 2023.

Downloads

Published

2026-02-28