POEM: Interactive Prompt Optimization for Enhancing Multimodal Reasoning of Large Language Models

He, Jianben; Wang, Xingbo; Liu, Shiyi; Wu, Guande; Silva, Claudio; Qu, Huamin

Computer Science > Human-Computer Interaction

arXiv:2406.03843 (cs)

[Submitted on 6 Jun 2024 (v1), last revised 14 Jun 2024 (this version, v2)]

Title:POEM: Interactive Prompt Optimization for Enhancing Multimodal Reasoning of Large Language Models

Authors:Jianben He, Xingbo Wang, Shiyi Liu, Guande Wu, Claudio Silva, Huamin Qu

View PDF HTML (experimental)

Abstract:Large language models (LLMs) have exhibited impressive abilities for multimodal content comprehension and reasoning with proper prompting in zero- or few-shot settings. Despite the proliferation of interactive systems developed to support prompt engineering for LLMs across various tasks, most have primarily focused on textual or visual inputs, thus neglecting the complex interplay between modalities within multimodal inputs. This oversight hinders the development of effective prompts that guide model multimodal reasoning processes by fully exploiting the rich context provided by multiple modalities. In this paper, we present POEM, a visual analytics system to facilitate efficient prompt engineering for enhancing the multimodal reasoning performance of LLMs. The system enables users to explore the interaction patterns across modalities at varying levels of detail for a comprehensive understanding of the multimodal knowledge elicited by various prompts. Through diverse recommendations of demonstration examples and instructional principles, POEM supports users in iteratively crafting and refining prompts to better align and enhance model knowledge with human insights. The effectiveness and efficiency of our system are validated through two case studies and interviews with experts.

Comments:	11 pages, 5 figures
Subjects:	Human-Computer Interaction (cs.HC); Artificial Intelligence (cs.AI)
MSC classes:	68
ACM classes:	H.5; I.2.1
Cite as:	arXiv:2406.03843 [cs.HC]
	(or arXiv:2406.03843v2 [cs.HC] for this version)
	https://doi.org/10.48550/arXiv.2406.03843

Submission history

From: Jianben He [view email]
[v1] Thu, 6 Jun 2024 08:21:30 UTC (21,881 KB)
[v2] Fri, 14 Jun 2024 14:36:58 UTC (21,881 KB)

Computer Science > Human-Computer Interaction

Title:POEM: Interactive Prompt Optimization for Enhancing Multimodal Reasoning of Large Language Models

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Human-Computer Interaction

Title:POEM: Interactive Prompt Optimization for Enhancing Multimodal Reasoning of Large Language Models

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators