Identifying Multi-modal Knowledge Neurons in Pretrained Transformers via Two-stage Filtering

Sato, Yugen; Takagi, Tomohiro

Computer Science > Artificial Intelligence

arXiv:2503.22941 (cs)

[Submitted on 29 Mar 2025]

Title:Identifying Multi-modal Knowledge Neurons in Pretrained Transformers via Two-stage Filtering

Authors:Yugen Sato, Tomohiro Takagi

View PDF HTML (experimental)

Abstract:Recent advances in large language models (LLMs) have led to the development of multimodal LLMs (MLLMs) in the fields of natural language processing (NLP) and computer vision. Although these models allow for integrated visual and language understanding, they present challenges such as opaque internal processing and the generation of hallucinations and misinformation. Therefore, there is a need for a method to clarify the location of knowledge in MLLMs.
In this study, we propose a method to identify neurons associated with specific knowledge using MiniGPT-4, a Transformer-based MLLM. Specifically, we extract knowledge neurons through two stages: activation differences filtering using inpainting and gradient-based filtering using GradCAM. Experiments on the image caption generation task using the MS COCO 2017 dataset, BLEU, ROUGE, and BERTScore quantitative evaluation, and qualitative evaluation using an activation heatmap showed that our method is able to locate knowledge with higher accuracy than existing methods.
This study contributes to the visualization and explainability of knowledge in MLLMs and shows the potential for future knowledge editing and control.

Subjects:	Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Multimedia (cs.MM)
Cite as:	arXiv:2503.22941 [cs.AI]
	(or arXiv:2503.22941v1 [cs.AI] for this version)
	https://doi.org/10.48550/arXiv.2503.22941

Submission history

From: Yugen Sato [view email]
[v1] Sat, 29 Mar 2025 02:16:15 UTC (8,602 KB)

Computer Science > Artificial Intelligence

Title:Identifying Multi-modal Knowledge Neurons in Pretrained Transformers via Two-stage Filtering

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Artificial Intelligence

Title:Identifying Multi-modal Knowledge Neurons in Pretrained Transformers via Two-stage Filtering

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators