Search arXivSearch

arXiv · 2401.05384

From Good to Great: Improving Math Reasoning with Tool-Augmented Interleaf Prompting

Abstract

This paper investigates the performance of Large Language Models (LLMs) and Tool-augmented LLMs in tackling complex mathematical reasoning tasks. We introduce IMP-TIP: Improving Math Reasoning with Tool-augmented Interleaf Prompting, a framework that combines the strengths of both LLMs and Tool-augmented LLMs. IMP-TIP follows the ``From Good to Great" concept, collecting multiple potential solutions from both LLMs and their Tool-Augmented counterparts for the same math problem, and then selecting or re-generating the most accurate answer after cross-checking these solutions via tool-augmented interleaf prompting. The framework incorporates two key aspects: self-prompt and tool-augmented interleaf prompting (TIP). The former allows LLMs to autonomously refine and improve an initial prompt related to tool usage, while the latter enables LLMs to derive the final answer by dynamically analyzing the problem, cross-checking potential solutions, and revising previous reasoning hints in an interleaved manner. Experimental analysis shows that IMP-TIP achieves enhanced mathematical capabilities and outperforms traditional LLMs and tool-augmented LLMs in accuracy and reasoning diversity on math reasoning tasks. For instance, IMP-TIP can improve Tool-augmented ChatGPT on GSM8K-Hard from 56.0% to 65.2%.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Nuo Chen, Hongguang Li, Baoyuan Wang, Jia Li. 2023-12-18. From Good to Great: Improving Math Reasoning with Tool-Augmented Interleaf Prompting. https://arxiv.org/abs/2401.05384

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Come for the vibe, stay for the math

This article describes our experiences in mathematical outreach over the past decade. We talk about specific activities, but also general principles that we've learned along the way.

math.HO

Graduate Mathematics in the Age of AI: Forming Mathematicians for Original, Independent, and Responsible Inquiry

Artificial intelligence can increasingly produce plausible, sophisticated mathematical material faster than a developing graduate student can understand or verify it. A sophisticated result or paper draft therefore becomes weaker evidence of the student's own mathematical development. This creates a formation gap between output and personal capacity, and a trust gap between a convincing argument and warranted acceptance. The formation gap can persist even when the student understands the output: understanding a supplied argument does not by itself establish the capacity to initiate and direct inquiry. These gaps are not the whole story. AI can also help students explore examples, compare approaches, enter unfamiliar areas, and undertake ambitious research. The task is to design an apprenticeship that realizes these possibilities while developing substantive mathematical command. The central purpose of a mathematics PhD is to form mathematicians capable of original, independent, and responsible inquiry, including inquiry conducted with AI. This document develops that objective through four connected capacities: competence, judgment, independence, and responsibility. It distinguishes a work's contribution to mathematics from the evidence it provides of a student's formation; explains how a known answer can initiate rather than end creative inquiry; and proposes changes in learning activities, assessment, doctoral originality, advising, and institutional support. Purposeful independent work and ambitious AI-assisted research are complementary parts of the model. Its recommendations include proportionate contribution statements, recognition of advising costs, and staged pilots evaluating both mathematical ability and effective human--AI collaboration. The aim is not to preserve an inherited sequence of training, but to improve mathematical formation as mathematical practice changes.

math.HO