Learning Goal-Oriented Visual Dialog via Tempered Policy Gradient

Zhao, Rui; Tresp, Volker

doi:10.1109/SLT.2018.8639546

Computer Science > Machine Learning

arXiv:1807.00737 (cs)

[Submitted on 2 Jul 2018 (v1), last revised 24 May 2020 (this version, v5)]

Title:Learning Goal-Oriented Visual Dialog via Tempered Policy Gradient

Authors:Rui Zhao, Volker Tresp

View PDF

Abstract:Learning goal-oriented dialogues by means of deep reinforcement learning has recently become a popular research topic. However, commonly used policy-based dialogue agents often end up focusing on simple utterances and suboptimal policies. To mitigate this problem, we propose a class of novel temperature-based extensions for policy gradient methods, which are referred to as Tempered Policy Gradients (TPGs). On a recent AI-testbed, i.e., the GuessWhat?! game, we achieve significant improvements with two innovations. The first one is an extension of the state-of-the-art solutions with Seq2Seq and Memory Network structures that leads to an improvement of 7%. The second one is the application of our newly developed TPG methods, which improves the performance additionally by around 5% and, even more importantly, helps produce more convincing utterances.

Comments:	Published in IEEE Spoken Language Technology (SLT 2018), Athens, Greece
Subjects:	Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Machine Learning (stat.ML)
Cite as:	arXiv:1807.00737 [cs.LG]
	(or arXiv:1807.00737v5 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.1807.00737
Related DOI:	https://doi.org/10.1109/SLT.2018.8639546

Submission history

From: Rui Zhao [view email]
[v1] Mon, 2 Jul 2018 15:14:43 UTC (1,501 KB)
[v2] Tue, 3 Jul 2018 05:35:32 UTC (1,501 KB)
[v3] Thu, 4 Oct 2018 08:24:41 UTC (1,505 KB)
[v4] Wed, 20 Feb 2019 10:22:01 UTC (1,505 KB)
[v5] Sun, 24 May 2020 08:03:58 UTC (1,505 KB)

Full-text links:

Access Paper:

view license

Current browse context:

cs.LG

< prev | next >

new | recent | 2018-07

Change to browse by:

cs
cs.AI
stat
stat.ML

References & Citations

DBLP - CS Bibliography

listing | bibtex

Rui Zhao
Volker Tresp

export BibTeX citation

Computer Science > Machine Learning

Title:Learning Goal-Oriented Visual Dialog via Tempered Policy Gradient

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:Learning Goal-Oriented Visual Dialog via Tempered Policy Gradient

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators