Dataset Condensation Atlas

Survey

Sachdeva & McAuley survey

Data Distillation: A Survey

Noveen Sachdeva, Julian McAuley

TMLR 2023 · first public 2023-01-11 · arXiv 2301.04272

paper ↗core✓ abstract read

In one paragraph

This survey presents a formal framework for data distillation with a detailed taxonomy of existing approaches, and covers the method across three data modalities: images, graphs, and user-item interactions (recommender systems), identifying current challenges and future research directions for each.

Where it sits

Abstract (verbatim from arXiv)

The popularity of deep learning has led to the curation of a vast number of massive and multifarious datasets. Despite having close-to-human performance on individual tasks, training parameter-hungry models on large datasets poses multi-faceted problems such as (a) high model-training time; (b) slow research iteration; and (c) poor eco-sustainability. As an alternative, data distillation approaches aim to synthesize terse data summaries, which can serve as effective drop-in replacements of the original dataset for scenarios like model training, inference, architecture search, etc. In this survey, we present a formal framework for data distillation, along with providing a detailed taxonomy of existing approaches. Additionally, we cover data distillation approaches for different data modalities, namely images, graphs, and user-item interactions (recommender systems), while also identifying current challenges and future research directions.

BibTeX (generated; prefer the venue's official entry)
@article{sachdeva2023data,
  title   = {Data Distillation: A Survey},
  author  = {Noveen Sachdeva and Julian McAuley},
  journal = {TMLR 2023},
  year    = {2023}
}