Generative Vision Atlas

Research line · ascendant

Evaluation and benchmarks

Measure whether generators actually do what their scores claim, and fix the benchmarks when they stop tracking human judgment.

What defines membership

Progress claims are only as good as the measurement behind them, so benchmark design is a research contribution in its own right rather than infrastructure.

How the line developed

Read top to bottom: what came before, the idea itself, the evidence for it, what improved, and where it breaks.

The idea

ICE-Bench · 2025-03strong-followup

Evaluates creation and editing jointly, matching the architectural shift toward single models that do both.

Improvement

RISEBench · 2025-04core

First benchmark for edits requiring reasoning rather than recognition; the best model tested scored under 30 percent.

ImgEdit · 2025-05core

Pairs a large editing dataset with a benchmark covering instruction adherence and background retention.

KRIS-Bench · 2025-05strong-followup

Critiques the reasoning benchmarks as too coarse and grounds its own in an explicit knowledge framework.

UniEval · 2025-05strong-followup

Evaluates understanding and generation together, the only way to catch a unified model gaining one while losing the other.

ContextBias · 2026-08emerging

Tests whether stereotyped attributes persist when prompt context changes, across 92 roles and 66,000 images.

Limitation

GenEval 2 · 2025-12core

Measures up to 17.7 percent drift between the field's most-cited compositional benchmark and human judgment as models improved.

What it gets right

  • Every claim elsewhere in this atlas depends on these holding up
  • Actively self-correcting: the strongest results here are about benchmarks failing, published by the people who use them
  • Cheap to run relative to training, so it scales with the field

Where it is weak

  • Benchmarks drift as models improve, and detection lags the drift
  • Model-as-judge evaluation inherits the judge's biases
  • The same checkpoint scores differently across papers, so absolute numbers carry less information than they appear to