Towards better Human-Agent Alignment: Assessing Task Utility in LLM-Powered Applications

Arabzadeh, Negar; Kiseleva, Julia; Wu, Qingyun; Wang, Chi; Awadallah, Ahmed; Dibia, Victor; Fourney, Adam; Clarke, Charles

Computer Science > Computation and Language

arXiv:2402.09015v3 (cs)

[Submitted on 14 Feb 2024 (v1), last revised 22 Feb 2024 (this version, v3)]

Title:Towards better Human-Agent Alignment: Assessing Task Utility in LLM-Powered Applications

Authors:Negar Arabzadeh, Julia Kiseleva, Qingyun Wu, Chi Wang, Ahmed Awadallah, Victor Dibia, Adam Fourney, Charles Clarke

View PDF HTML (experimental)

Abstract:The rapid development in the field of Large Language Models (LLMs) has led to a surge in applications that facilitate collaboration among multiple agents to assist humans in their daily tasks. However, a significant gap remains in assessing whether LLM-powered applications genuinely enhance user experience and task execution efficiency. This highlights the pressing need for methods to verify utility of LLM-powered applications, particularly by ensuring alignment between the application's functionality and end-user needs. We introduce AgentEval provides an implementation for the math problems, a novel framework designed to simplify the utility verification process by automatically proposing a set of criteria tailored to the unique purpose of any given application. This allows for a comprehensive assessment, quantifying the utility of an application against the suggested criteria. We present a comprehensive analysis of the robustness of quantifier's work.

Subjects:	Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
Cite as:	arXiv:2402.09015 [cs.CL]
	(or arXiv:2402.09015v3 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2402.09015

Submission history

From: Julia Kiseleva [view email]
[v1] Wed, 14 Feb 2024 08:46:15 UTC (14,525 KB)
[v2] Thu, 15 Feb 2024 18:24:03 UTC (15,116 KB)
[v3] Thu, 22 Feb 2024 23:49:10 UTC (15,117 KB)

Computer Science > Computation and Language

Title:Towards better Human-Agent Alignment: Assessing Task Utility in LLM-Powered Applications

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:Towards better Human-Agent Alignment: Assessing Task Utility in LLM-Powered Applications

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators