Dataset Condensation Atlas

Application · Distribution and feature matching

IEM

Towards Efficient Deep Hashing Retrieval: Condensing Your Data via Feature-Embedding Matching

Tao Feng, Jie Zhang, Huashan Liu, Zhijie Wang, Shengyuan Pang

arXiv 2023 · first public 2023-05-29 · arXiv 2305.18076

paper ↗catalogued✓ abstract read

In one paragraph

Adapts dataset condensation to deep hashing retrieval with IEM (Information-intensive feature Embedding Matching), a distribution-matching-centered method that adds model and data augmentation to strengthen the condensed hashing-space features, since retrieval training does not benefit directly from condensation methods designed for classification accuracy. Reports superior performance and efficiency relative to applying mainstream condensation methods to deep hashing retrieval.

Where it sits

Abstract (verbatim from arXiv)

Deep hashing retrieval has gained widespread use in big data retrieval due to its robust feature extraction and efficient hashing process. However, training advanced deep hashing models has become more expensive due to complex optimizations and large datasets. Coreset selection and Dataset Condensation lower overall training costs by reducing the volume of training data without significantly compromising model accuracy for classification task. In this paper, we explore the effect of mainstream dataset condensation methods for deep hashing retrieval and propose IEM (Information-intensive feature Embedding Matching), which is centered on distribution matching and incorporates model and data augmentation techniques to further enhance the feature of hashing space. Extensive experiments demonstrate the superior performance and efficiency of our approach.

BibTeX (generated; prefer the venue's official entry)
@article{feng2023towards,
  title   = {Towards Efficient Deep Hashing Retrieval: Condensing Your Data via Feature-Embedding Matching},
  author  = {Tao Feng and Jie Zhang and Huashan Liu and Zhijie Wang and Shengyuan Pang},
  journal = {arXiv preprint arXiv:2305.18076},
  year    = {2023}
}

Nearby in Distribution and feature matching

2026-06

RAHA — Rank-Aware Hyperbolic Alignment for Vision-Language Dataset Distillation

Jongoh Jeong, Sun-Kyung Lee, Kuk-Jin Yoon · ECCV 2026notableVision–languagepaper ↗code ↗

2026-05

MDM — Multimodal Distribution Matching for Vision-Language Dataset Distillation

Jongoh Jeong, Hoyong Kwon, Minseok Kim et al. · CVPR 2026notableVision–languagepaper ↗code ↗

2026-03

Sneakdoor — SNEAKDOOR: Stealthy Backdoor Attacks against Distribution Matching-based Dataset Condensation

He Yang, Dongyi Lv, Song Ma et al. · NeurIPS 2025notablepaper ↗code ↗

2026-03

Harmonic Dataset Distillation for Time Series Forecasting

Seungha Hong, Sanghwan Jang, Wonbin Kweon et al. · AAAI 2026notableTime seriespaper ↗

2025-11

Algorithmic Guarantees for Distilling Supervised and Offline RL Datasets

Aaryan Gupta, Rishi Saket, Aravindan Raghuveer · ICLR 2026notableOther datapaper ↗