Teaching LLMs to Refine with Tools

Yu, Dian; Zhang, Yuheng; Xu, Jiahao; Liang, Tian; Song, Linfeng; Tu, Zhaopeng; Mi, Haitao; Yu, Dong

Computer Science > Computation and Language

arXiv:2412.16871 (cs)

[Submitted on 22 Dec 2024]

Title:Teaching LLMs to Refine with Tools

Authors:Dian Yu, Yuheng Zhang, Jiahao Xu, Tian Liang, Linfeng Song, Zhaopeng Tu, Haitao Mi, Dong Yu

View PDF HTML (experimental)

Abstract:Large language models (LLMs) can refine their responses based on feedback, enabling self-improvement through iterative training or test-time refinement. However, existing methods predominantly focus on refinement within the same reasoning format, which may lead to non-correcting behaviors. We propose CaP, a novel approach that uses external tools to refine chain-of-thought (CoT) responses generated by the same or other LLMs. CaP employs a two-stage training process: supervised fine-tuning followed by preference optimization with DPO variants. Our observations highlight the critical role of preference optimization in enabling effective refinement. Additionally, we compare several sampling strategies to leverage CoT and tools at inference time. Experimental results demonstrate CaP's potential for effective cross-reasoning refinement and efficient inference.

Subjects:	Computation and Language (cs.CL)
Cite as:	arXiv:2412.16871 [cs.CL]
	(or arXiv:2412.16871v1 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2412.16871

Submission history

From: Dian Yu [view email]
[v1] Sun, 22 Dec 2024 05:43:50 UTC (132 KB)

Full-text links:

Access Paper:

view license

Current browse context:

cs.CL

< prev | next >

new | recent | 2024-12

Change to browse by:

References & Citations

export BibTeX citation

Computer Science > Computation and Language

Title:Teaching LLMs to Refine with Tools

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:Teaching LLMs to Refine with Tools

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators