Every AI art generator claims character consistency. None of them publish a way to check it, so the claim is unfalsifiable and every comparison collapses into vibes. This is a protocol to fix that: six fixed scenes, five scored attributes, 60 points, about fifteen minutes per tool. It is deliberately simple enough that two people running it separately should land within a few points of each other — and it is published here so you can run it against us.
What this test measures
Character consistency is the ability to place the same character in different scenes without the character changing. This benchmark measures it as a loss: how much identity is destroyed by six specific stresses, each chosen because it breaks a different thing.
What it is not is an image-quality test. A generator can produce beautiful frames and fail this badly, and a plain-looking generator can score well. Quality and consistency are separate axes, and conflating them is why most comparisons are useless.
Each scene is chosen to break a different attribute.
The six scenes
Generate all six from one reference, one seed, and one wording template. Change only what each scene requires.
How to run the 6-scene consistency test
- 1
Scene 1 — neutral portrait (the control)
Head and shoulders, facing forward, neutral expression, even light, plain backdrop. This frame defines the character for scoring purposes. Everything else is compared against it.
- 2
Scene 2 — side profile
A strict side view, same framing. This tests face geometry, because a profile forces the model to infer a view your reference probably never showed.
- 3
Scene 3 — seated, full body
Sitting on a simple chair, full body in frame. Pulling the camera back shrinks the face and forces the model to commit to proportions.
- 4
Scene 4 — complex outfit
Add layers: a hooded cloak, gloves, belts. This tests whether identity survives a changed silhouette, the most common real-world edit.
- 5
Scene 5 — hard lighting
A single hard light source with deep shadow. Tests palette and skin tone stability when the model can no longer rely on flat even light.
- 6
Scene 6 — style swap
Same everything, different render style. The hardest frame by construction: styles do not share a face grammar, so score this one gently.
Two rules make results comparable between people. Fix the seed across all six scenes, so differences come from your edits rather than from the dice. And keep the unchanged wording byte-identical — retyping a phrase is a second edit you did not intend to make.
The scoring rubric
Score each of the six frames on five attributes. Zero, one or two points each — no half points, because they only ever mean "I could not decide".
| Attribute | 0 points | 1 point | 2 points |
|---|---|---|---|
| Face | A different person | Same person, visibly re-drawn features | Same face, same proportions |
| Hair | Different colour or length | Same hair, restyled or reshaped | Same cut, colour and parting |
| Outfit | Signature outfit or accessories gone | Kept but altered, or items added | As specified, accessories intact |
| Proportions | Wrong build, broken anatomy | Slightly off build or limb length | Consistent build and height |
| Recognisability | Would not identify as the same character | Recognisable after a second look | Instantly the same character |
Sixty points is the maximum. Anything above 50 is a genuinely steerable workflow; 35–50 means usable with re-rolls; below 35 means you are not generating a character, you are generating a family resemblance.
Our first run: 47/60
Here is the protocol applied to our own generator, with the frames
published so you can disagree with the scoring. One reference
(a dark-haired half-elf cleric in silver armour), seed 770125, all six
scenes generated in one sitting.

| Scene | Face | Hair | Outfit | Proportions | Recognisable | Total |
|---|---|---|---|---|---|---|
| 1 · Neutral portrait (control) | 2 | 2 | 2 | 2 | 2 | 10 |
| 2 · Side profile | 2 | 1 | 1 | 2 | 2 | 8 |
| 3 · Seated, full body | 1 | 1 | 1 | 0 | 1 | 4 |
| 4 · Complex outfit | 2 | 2 | 2 | 2 | 2 | 10 |
| 5 · Hard lighting | 2 | 1 | 1 | 2 | 2 | 8 |
| 6 · Style swap | 1 | 1 | 1 | 2 | 2 | 7 |
| Total | 47 / 60 |
The interesting part is not the number, it is where the points went:
- The complex outfit scene scored full marks, which is the opposite of what most people expect. A hooded cloak with a fur collar over the armour kept the face, the hair and the armour underneath intact. Adding layers is apparently easier than replacing them.
- The seated full-body frame was the disaster at 4/10. The face survived as recognisable but was re-drawn smaller and rounder, the armour grew into a long robe that was never requested, and the build came back visibly compressed — short arms, a torso that reads as a much younger figure against the chair. Pulling the camera back cost more than changing the clothes did.
- The profile lost the hair, not the face. Face geometry held up cleanly at 2, while shoulder-length hair became a long ponytail.
- Hard lighting held the palette but restyled the hair again, which suggests hair is simply the least anchored attribute in the pipeline.
- The style swap scored 7 and that is a pass, not a failure. Anime and semi-realistic do not share a face grammar; the character stayed recognisable across a change that redraws every pixel of the face.

Two honest caveats about this run. The backdrop was specified as a plain neutral grey studio in five of six scenes, and several frames came back with candlelit gothic architecture instead — an instruction that was simply overridden. And in two frames a small camera-like prop appeared at the belt, traceable to the phrase "facing the camera directly" in the pose wording, which the model read as an object. Both are documented in detail in the piece on why characters change when you change their clothes.
What the test does not measure
Publishing a benchmark means publishing its limits, otherwise it becomes a marketing number:
- It does not measure image quality, style range or speed. A tool can score 55 and still make images you dislike.
- It does not cover multi-character scenes. Two characters in one frame is a harder and different problem.
- It is one character. Distinctive designs with strong canonical outfits behave differently from plain ones — run the test on the kind of character you actually generate.
- Scoring is human. Two people will differ by a few points. That is acceptable for comparing 47 to 30, and useless for comparing 47 to 49.
- It rewards the anchored, not the imaginative. A generator that copies your reference exactly and refuses to interpret would score brilliantly here and be useless in practice.
- It scores five attributes, and identity has more. Apparent age, distinguishing marks and emotional register are not in the rubric — see character consistency is not face consistency for the full set.
Used with those caveats, it does the one thing vibes cannot: it makes "consistent characters" a claim someone can check. If you run it on another tool, the frames and the per-scene table are what make your result worth reading — publish those, not just the total.
For the workflow side of the problem rather than the measurement side, see the guides on consistent characters across scenes and reference-based generation. If you are choosing between tools on price, the attempts you spend fighting drift belong in your cost per keeper.


Frequently asked questions
What is the AI character consistency test?
A six-scene benchmark: generate one character as a neutral portrait, side profile, seated full body, complex outfit, hard lighting and style swap, then score each frame 0–2 on face, hair, outfit, proportions and recognisability for a total out of 60.
How do I test whether an AI generator keeps characters consistent?
Fix one reference and one seed, generate the six scenes above changing only what each scene requires, then score every frame against the neutral portrait rather than against your own idea of the character.
What is a good consistency score?
Above 50 out of 60 is a genuinely steerable workflow. 35–50 is usable with re-rolls. Below 35 means the tool is producing a family resemblance rather than the same character.
Why fix the seed during the test?
A fixed seed removes random variation between runs, so a difference you see is caused by the edit you made. It does not preserve identity by itself — that is what the reference does.
Which scene breaks character consistency the most?
In our run, the seated full-body frame. Pulling the camera back renders the face at a smaller scale and forces the model to re-decide body proportions, which cost more points than changing the outfit.
Can I use this benchmark to compare two generators?
Yes, if you keep the character, the seed policy, the wording template and the scoring rubric identical across both, and publish the frames alongside the scores so others can check your judgement.
Benchmark it yourself
Six scenes, one seed, 60 points. Then tell us where we lost points.
Start generating