We let three AI models retouch the same real photos, turn by turn, and watched exactly where each one lost the user's trust.
Eight persona journeys, each carrying its own context, goals, and tolerance forward across every turn.
Failure is measured as behavior over time — acceptance, workaround, switching — not a single aesthetic score.
How they scored
For an image-model team, a single aesthetic score hides where the model actually breaks. This study separates the fast-moving score from the slow-moving discovery of failures.
Identity drift or text corruption that only appears on turn four is exactly the failure a one-shot eval never sees — but a real user does, and it's where they lose trust.