Chinese Idiom Paraphrasing

Jipeng Qiang; Yang Li; Chaowei Zhang; Yun Li; Yi Zhu; Yunhao Yuan; Xindong Wu

Vol. 11 (2023)

TACL approved

Chinese Idiom Paraphrasing

Published 2023-07-24

Jipeng Qiang
Yang Li
Chaowei Zhang
Yun Li
Yi Zhu
Yunhao Yuan
Xindong Wu

Jipeng Qiang
Yangzhou University

Yang Li
Yangzhou University

Chaowei Zhang
Yangzhou University

Yun Li
Yangzhou University

Yi Zhu
Yangzhou University

Yunhao Yuan
Yangzhou University

Xindong Wu
Hefei University of Technology

Abstract

Idioms, are a kind of idiomatic expression in Chinese, most of which consist of four Chinese characters. Due to the properties of non-compositionality and metaphorical meaning, Chinese Idioms are hard to be understood by children and non-native speakers. This study proposes a novel task, denoted as Chinese Idiom Paraphrasing (CIP). CIP aims to rephrase idioms-contained sentences to non-idiomatic ones under the premise of preserving the original sentence's meaning. Since the sentences without idioms are more easily handled by Chinese NLP systems, CIP can be used to pre-process Chinese datasets, thereby facilitating and improving the performance of Chinese NLP tasks, e.g., machine translation system, Chinese idiom cloze, and Chinese idiom embeddings. In this study, we can treat CIP task as a special paraphrase generation task. To circumvent difficulties in acquiring annotations, we first establish a large-scale CIP dataset based on human and machine collaboration, which consists of 115,529 sentence pairs. In addition to three sequence-to-sequence methods as the baselines, we further propose a novel infill-based approach based on text infilling. The results show that the proposed method has better performance than the baselines based on the established CIP dataset.

Article at MIT Press