Dataset Condensation Atlas

Method · Label distillation and soft labels

DRUPI

DRUPI: Dataset Reduction Using Privileged Information

Shaobo Wang, Youxin Jiang, Tianle Niu, Yantai Yang, Ruiji Zhang, Shuhao Hu, Shuaiyu Zhang, Chenghao Sun, Weiya Li, Conghui He, Xuming Hu, Linfeng Zhang

arXiv 2024 · first public 2024-10-02 · arXiv 2410.01611

paper ↗catalogued✓ abstract read

In one paragraph

Introduces Dataset Condensation using Privileged Information (DRUPI/DCPI): alongside condensed images and labels, synthesizes auxiliary feature or attention labels as an additional training target; finds that moderately (not maximally) discriminative and diverse feature labels work best, and shows the technique plugs into existing condensation methods for consistent gains on ImageNet-1K, CIFAR-10/100 and Tiny-ImageNet.

Where it sits

Abstract (verbatim from arXiv)

Dataset Condensation (DC) seeks to select or distill samples from large datasets into smaller subsets while preserving performance on target tasks. Existing methods primarily focus on pruning or synthesizing data in the same format as the original dataset, typically being the input data and corresponding labels. However, in DC settings, we find it is possible to synthesize more information beyond the data-label pair as an additional learning target to facilitate model training. In this paper, we introduce Dataset Condensation using Privileged Information (DCPI), which enriches DC by synthesizing privileged information alongside the reduced dataset. This privileged information can take the form of feature labels or attention labels, providing auxiliary supervision to improve model learning. Our findings reveal that effective feature labels must balance between being overly discriminative and excessively diverse, with a moderate level proves optimal for improving the reduced dataset's efficacy. Extensive experiments on ImageNet-1K, CIFAR-10/100 and Tiny ImageNet demonstrate that DCPI integrates seamlessly with existing dataset condensation methods, offering significant performance gains.

BibTeX (generated; prefer the venue's official entry)
@article{wang2024drupi,
  title   = {DRUPI: Dataset Reduction Using Privileged Information},
  author  = {Shaobo Wang and Youxin Jiang and Tianle Niu and Yantai Yang and Ruiji Zhang and Shuhao Hu and Shuaiyu Zhang and Chenghao Sun and Weiya Li and Conghui He and Xuming Hu and Linfeng Zhang},
  journal = {arXiv preprint arXiv:2410.01611},
  year    = {2024}
}

Nearby in Label distillation and soft labels

2026-04

Soft Label Pruning and Quantization for Large-Scale Dataset Distillation

Xiao Lingao, Yang He · TPAMI 2026notablepaper ↗code ↗

2025-12

HALD — Hard Labels In! Rethinking the Role of Hard Labels in Mitigating Local Semantic Drift

Jiacheng Cui, Bingkui Tong, Xinyue Bi et al. · ICML 2026notablepaper ↗code ↗

2025-11

Optimizing Distributional Geometry Alignment with Optimal Transport for Generative Dataset Distillation

Xiao Cui, Yulei Qin, Wengang Zhou et al. · NeurIPS 2025notablepaper ↗

2025-11

RLDD — Rethinking Long-tailed Dataset Distillation: A Uni-Level Framework with Unbiased Recovery and Relabeling

Xiao Cui, Yulei Qin, Xinyue Li et al. · AAAI 2026notablepaper ↗code ↗

2025-11

ADSA — Rectifying Soft-Label Entangled Bias in Long-Tailed Dataset Distillation

Chenyang Jiang, Hang Zhao, Xinyu Zhang et al. · NeurIPS 2025notablepaper ↗code ↗