Enhancing Formalization for LLM’s Mathematical Reasoning

Authors

  • Wanchen Jiang

DOI:

https://doi.org/10.61173/cp7xmq61

Keywords:

Large Language Models (LLMs), Mathe-matical Reasoning, Self-Verification, Formalization, Ro-bustness

Abstract

Large Language Models (LLMs) perform very well in natural language processing tasks, but they still have problems in complex mathematical reasoning. One important difficulty is the formalization step, which means to correctly and precisely formalize natural language math problems into mathematical expressions. Current methods rely too much on this formalization, so they are often easy to make misun derstanding and inconsistency. In this paper, we propose an enhanced formalization framework, which combines multi-round formalization and selfverification. In detail, the LLM will formalize the same problem several times into different formal representation, and then a verification module is used to select and correct the results, to make sure consistency and correctness. We do experiments with advanced models like GPT-4 and DeepSeek, on benchmark datasets such as GSM8K and MATH. The results show that our method can improve accuracy about 0.05 compared with methods like Chainof-Thought and Self-Consistency, and it is also more stable when facing rephrased adversarial examples.

References

Janice Ahn, Rishu Verma, Renze Lou, Di Liu, Rui Zhang, and Wenpeng Yin. Large language models for mathematical reasoning: Progresses and challenges. In Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics: Student Research Workshop, pages 225–237, 2024. Mingyu Zong and Bhaskar Krishnamachari. Solving math word problems concerning systems of equations with gpt- 3. In Proceedings of the AAAI conference on artificial intelligence, volume 37, 129 pages 15972–15979, 2023. Wenhu Chen, Xueguang Ma, Xinyi Wang, and William W Cohen. Program of thoughts prompting: Disentangling computation from reasoning for numerical reasoning tasks. arXiv preprint arXiv:2211.12588, 2022. Shunyu Yao, Dian Yu, Jeffrey Zhao, et al. Tree of thoughts: Deliberate problem solving with large language models. arXiv preprint arXiv:2305.10601, 2023. Jacob Lightman, Xuezhi Wang, Jason Wei, et al. Let’s verify step by step. arXiv preprint arXiv:2305.20050, 2023. Yikang Shen, Dian Yu, Shuyan Zhou, et al. Large lan-

Dean&Francis Wanchen Jiang guage models are better reasoners with self-verification. arXiv preprint arXiv:2309.07973, 2023. OpenAI. Gpt-4 technical report. arXiv preprint arX- iv:2303.08774, 2023. DeepSeek-AI Team. Deepseek: Scaling open-source language models with efficient training and reasoning. arXiv preprint arXiv:2401.10968, 2024. Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, et al. Gsm8k: A dataset for grade school math word problems. arXiv preprint arXiv:2110.14168, 2021. Introduced as part of the Training Verifiers work.

Downloads

Published

2025-12-19