IdBench: Evaluating Semantic Representations of Identifier Names in Source Code

Wainakh, Yaza; Rauf, Moiz; Pradel, Michael

Computer Science > Machine Learning

arXiv:1910.05177 (cs)

[Submitted on 11 Oct 2019 (v1), last revised 14 Jan 2021 (this version, v2)]

Title:IdBench: Evaluating Semantic Representations of Identifier Names in Source Code

Authors:Yaza Wainakh, Moiz Rauf, Michael Pradel

View PDF

Abstract:Identifier names convey useful information about the intended semantics of code. Name-based program analyses use this information, e.g., to detect bugs, to predict types, and to improve the readability of code. At the core of name-based analyses are semantic representations of identifiers, e.g., in the form of learned embeddings. The high-level goal of such a representation is to encode whether two identifiers, e.g., len and size, are semantically similar. Unfortunately, it is currently unclear to what extent semantic representations match the semantic relatedness and similarity perceived by developers. This paper presents IdBench, the first benchmark for evaluating semantic representations against a ground truth created from thousands of ratings by 500 software developers. We use IdBench to study state-of-the-art embedding techniques proposed for natural language, an embedding technique specifically designed for source code, and lexical string distance functions. Our results show that the effectiveness of semantic representations varies significantly and that the best available embeddings successfully represent semantic relatedness. On the downside, no existing technique provides a satisfactory representation of semantic similarities, among other reasons because identifiers with opposing meanings are incorrectly considered to be similar, which may lead to fatal mistakes, e.g., in a refactoring tool. Studying the strengths and weaknesses of the different techniques shows that they complement each other. As a first step toward exploiting this complementarity, we present an ensemble model that combines existing techniques and that clearly outperforms the best available semantic representation.

Comments:	Accepted as full research paper at International Conference on Software Engineering (ICSE) 2021
Subjects:	Machine Learning (cs.LG); Programming Languages (cs.PL); Software Engineering (cs.SE); Machine Learning (stat.ML)
Cite as:	arXiv:1910.05177 [cs.LG]
	(or arXiv:1910.05177v2 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.1910.05177

Submission history

From: Michael Pradel [view email]
[v1] Fri, 11 Oct 2019 13:34:30 UTC (559 KB)
[v2] Thu, 14 Jan 2021 10:07:16 UTC (1,481 KB)

Computer Science > Machine Learning

Title:IdBench: Evaluating Semantic Representations of Identifier Names in Source Code

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:IdBench: Evaluating Semantic Representations of Identifier Names in Source Code

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators