A Method of Automatic Data Construction for Unseen Values in Multi-domain Dialogue State Tracking

AAAI 2022 DSTC10 workshop

Joongbo Shin, Gyeonghun Kim, Hyunjik Jo, Hosung Song, Janghoon Han, Yireun Kim, Stanley Jungkyu Choi

Abstract

The track of DSTC10 Track2 Task1 aims to benchmark the robustness of dialogue state tracking (DST) models against the gaps between written and spoken conversation. To this end, we develop an approach of automatic data construction to synthesize similar data imitating the validation set of DSTC10 Track2 Task1. Among ontology-based, generation-based, and span-based DST models, we decide to use ontology-based DST models since it performs best than others in our synthetic data. In addition, we design hand-crafted rules for utilizing the database for DST, and those rules play an important role in improving the recall of slots. For the final evaluation of DSTC10 Track2 Task1, our method achieves 27.03% and 26.7% joint goal accuracy on the validation set and test set, respectively.