Assessing the trade-off between prediction accuracy and interpretability for topic modeling on energetic materials corpora

Puerto, Monica; Kellett, Mason; Nikopoulou, Rodanthi; Fuge, Mark D.; Doherty, Ruth; Chung, Peter W.; Boukouvalas, Zois

Computer Science > Computation and Language

arXiv:2206.00773 (cs)

[Submitted on 1 Jun 2022]

Title:Assessing the trade-off between prediction accuracy and interpretability for topic modeling on energetic materials corpora

Authors:Monica Puerto, Mason Kellett, Rodanthi Nikopoulou, Mark D. Fuge, Ruth Doherty, Peter W. Chung, Zois Boukouvalas

View PDF

Abstract:As the amount and variety of energetics research increases, machine aware topic identification is necessary to streamline future research pipelines. The makeup of an automatic topic identification process consists of creating document representations and performing classification. However, the implementation of these processes on energetics research imposes new challenges. Energetics datasets contain many scientific terms that are necessary to understand the context of a document but may require more complex document representations. Secondly, the predictions from classification must be understandable and trusted by the chemists within the pipeline. In this work, we study the trade-off between prediction accuracy and interpretability by implementing three document embedding methods that vary in computational complexity. With our accuracy results, we also introduce local interpretability model-agnostic explanations (LIME) of each prediction to provide a localized understanding of each prediction and to validate classifier decisions with our team of energetics experts. This study was carried out on a novel labeled energetics dataset created and validated by our team of energetics experts.

Comments:	Accepted for publication in the 25th International Seminar New Trends in Research of Energetic Materials (NTREM 2022 proceedings)
Subjects:	Computation and Language (cs.CL); Machine Learning (cs.LG)
Cite as:	arXiv:2206.00773 [cs.CL]
	(or arXiv:2206.00773v1 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2206.00773

Submission history

From: Zois Boukouvalas [view email]
[v1] Wed, 1 Jun 2022 21:28:21 UTC (1,636 KB)

Computer Science > Computation and Language

Title:Assessing the trade-off between prediction accuracy and interpretability for topic modeling on energetic materials corpora

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:Assessing the trade-off between prediction accuracy and interpretability for topic modeling on energetic materials corpora

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators