Research line · ascendant
Evaluation and benchmarks
Measure whether generators actually do what their scores claim, and fix the benchmarks when they stop tracking human judgment.
What defines membership
Progress claims are only as good as the measurement behind them, so benchmark design is a research contribution in its own right rather than infrastructure.
How the line developed
Read top to bottom: what came before, the idea itself, the evidence for it, what improved, and where it breaks.
The idea
ICE-Bench · 2025-03strong-followup
Evaluates creation and editing jointly, matching the architectural shift toward single models that do both.
Improvement
RISEBench · 2025-04core
First benchmark for edits requiring reasoning rather than recognition; the best model tested scored under 30 percent.
ImgEdit · 2025-05core
Pairs a large editing dataset with a benchmark covering instruction adherence and background retention.
KRIS-Bench · 2025-05strong-followup
Critiques the reasoning benchmarks as too coarse and grounds its own in an explicit knowledge framework.
UniEval · 2025-05strong-followup
Evaluates understanding and generation together, the only way to catch a unified model gaining one while losing the other.
ContextBias · 2026-08emerging
Tests whether stereotyped attributes persist when prompt context changes, across 92 roles and 66,000 images.
Limitation
GenEval 2 · 2025-12core
Measures up to 17.7 percent drift between the field's most-cited compositional benchmark and human judgment as models improved.
What it gets right
- Every claim elsewhere in this atlas depends on these holding up
- Actively self-correcting: the strongest results here are about benchmarks failing, published by the people who use them
- Cheap to run relative to training, so it scales with the field
Where it is weak
- Benchmarks drift as models improve, and detection lags the drift
- Model-as-judge evaluation inherits the judge's biases
- The same checkpoint scores differently across papers, so absolute numbers carry less information than they appear to