Trustworthy DD
Private Set Generation with Discriminative Information
Dingfan Chen, Raouf Kerkouche, Mario Fritz
NeurIPS 2022 · first public 2022-11-07 · arXiv 2211.04446
In one paragraph
Rather than fitting a full private generative model to the data distribution, directly optimizes a small set of representative samples under differential privacy, supervised by discriminative information from the downstream task, which the paper argues is an easier and more DP-training-friendly target than full-distribution generative modeling. Reports greatly improved sample utility over prior state-of-the-art differentially private generation approaches for high-dimensional data.
Where it sits
- Setting: Image classification
Abstract (verbatim from arXiv)
Differentially private data generation techniques have become a promising solution to the data privacy challenge -- it enables sharing of data while complying with rigorous privacy guarantees, which is essential for scientific progress in sensitive domains. Unfortunately, restricted by the inherent complexity of modeling high-dimensional distributions, existing private generative models are struggling with the utility of synthetic samples. In contrast to existing works that aim at fitting the complete data distribution, we directly optimize for a small set of samples that are representative of the distribution under the supervision of discriminative information from downstream tasks, which is generally an easier task and more suitable for private training. Our work provides an alternative view for differentially private generation of high-dimensional data and introduces a simple yet effective method that greatly improves the sample utility of state-of-the-art approaches.
BibTeX (generated; prefer the venue's official entry)
@article{chen2022private,
title = {Private Set Generation with Discriminative Information},
author = {Dingfan Chen and Raouf Kerkouche and Mario Fritz},
journal = {NeurIPS 2022},
year = {2022}
}