Trustworthy DD
Attention Hijacking
Attention Hijacking: Backdooring Text Dataset Distillation via Semantic Anchors
Hang Ren
ICML 2026 · first public 2026-01-01
In one paragraph
Proposes a backdoor attack on text dataset distillation built on a "Semantic Anchoring Hypothesis": the attack reshapes gradients into input embeddings so the synthetic data evolves to turn a trigger word into an adversarial feature, integrated into the bi-level distillation loop so the attack satisfies both the clean-task and backdoor objectives at once.
Where it sits
- Setting: Text and language models
BibTeX (generated; prefer the venue's official entry)
@article{ren2026attention,
title = {Attention Hijacking: Backdooring Text Dataset Distillation via Semantic Anchors},
author = {Hang Ren},
journal = {ICML 2026},
year = {2026}
}