A Pulse Field Note · Compass

Which AI image model do people actually trust with their photos?

We let three AI models retouch the same real photos, turn by turn, and watched exactly where each one lost the user's trust.

Panel
8 persona journeys · multi-turn
Protocol
5 photo scenarios × multi-turn × 3 models
Records
1,200 judge-score records · 3 judge families
Field date
May 15, 2026
The anatomy of this study
Test Cases
×
Persona Agents
×
Metrics
×
Business Context
=
Contextual
Evaluation
01 · Test Cases

Realistic scenarios & goals

5 mobile photo scenarios × multi-turn
The goalFind where trust holds and where it breaks as a user edits the same photo repeatedly — not a one-shot aesthetic score.
02 · Persona Agents

Who we put on the panel

stateful, OCEAN-driven

Eight persona journeys, each carrying its own context, goals, and tolerance forward across every turn.

8
journeys
3
models
3
judge families
03 · Metrics

What we measured, & why

trust signals · acceptance / drift / friction

Failure is measured as behavior over time — acceptance, workaround, switching — not a single aesthetic score.

How they scored

GPT Image 2 · most session wins
24/40 sessions
04 · Business Context

Why this matters

what makes the finding meaningful

For an image-model team, a single aesthetic score hides where the model actually breaks. This study separates the fast-moving score from the slow-moving discovery of failures.

Identity drift or text corruption that only appears on turn four is exactly the failure a one-shot eval never sees — but a real user does, and it's where they lose trust.

Summary
Verdict
Across five mobile photo-editing journeys, GPT Image 2 wins 24 of 40 sessions. But the real finding isn't the winner — it's that model scores stabilize early while new failure modes keep surfacing turn after turn.
Method
Eight persona journeys each start from the same seed photo and edit it turn by turn across three image models; only a persona's own reaction feeds its next edit. Three independent judge families score 1,200 per-turn records.
Knockout
Scores converge before failures are discovered — model score and friction-coverage are two separate signals, and the score moves first.
Surprise
Five distinct failure modes only emerge over repeated edits: identity drift, text corruption, object instability, over-polish, and trust loss.
Limitations
Eight personas, five scenarios, single run — a longitudinal debugging lens, not a benchmark rate. Mobile photo-editing journeys only.
Go deeper — the full study, every persona, and all the evidence.
See the full study