Membership Inference Attacks on Sequence-to-Sequence Models: Is My Data In Your Machine Translation System?

Sorami Hisamoto; Matt Post; Kevin Duh

Vol. 8 (2020)

TACL approved

Membership Inference Attacks on Sequence-to-Sequence Models: Is My Data In Your Machine Translation System?

Published 2020-03-13

Sorami Hisamoto
Matt Post
Kevin Duh

Sorami Hisamoto
Johns Hopkins University / Works Applications

Matt Post
Johns Hopkins University

Kevin Duh
Johns Hopkins University

Abstract

Data privacy is an important issue for "machine learning as a service" providers. We focus on the problem of membership inference attacks: given a data sample and black-box access to a model's API, determine whether the sample existed in the model's training data. Our contribution is an investigation of this problem in the context of sequence-to-sequence models, which are important in applications such as machine translation and video captioning. We define the membership inference problem for sequence generation, provide an open dataset based on state-of-the-art machine translation models, and report initial results on whether these models leak private information against several kinds of membership inference attacks.

Article at MIT Press (presented at ACL 2020)