The AI Character Consistency Test: A 6-Scene Benchmark

9 min read
The six scenes of the consistency test with what each one tests

Every AI art generator claims character consistency. None of them publish a way to check it, so the claim is unfalsifiable and every comparison collapses into vibes. This is a protocol to fix that: six fixed scenes, five scored attributes, 60 points, about fifteen minutes per tool. It is deliberately simple enough that two people running it separately should land within a few points of each other — and it is published here so you can run it against us.

What this test measures

Character consistency is the ability to place the same character in different scenes without the character changing. This benchmark measures it as a loss: how much identity is destroyed by six specific stresses, each chosen because it breaks a different thing.

What it is not is an image-quality test. A generator can produce beautiful frames and fail this badly, and a plain-looking generator can score well. Quality and consistency are separate axes, and conflating them is why most comparisons are useless.

The six scenes of the consistency test and what each one testsEach scene is chosen to break a different attribute.

The six scenes

Generate all six from one reference, one seed, and one wording template. Change only what each scene requires.

How to run the 6-scene consistency test

  1. 1

    Scene 1 — neutral portrait (the control)

    Head and shoulders, facing forward, neutral expression, even light, plain backdrop. This frame defines the character for scoring purposes. Everything else is compared against it.

  2. 2

    Scene 2 — side profile

    A strict side view, same framing. This tests face geometry, because a profile forces the model to infer a view your reference probably never showed.

  3. 3

    Scene 3 — seated, full body

    Sitting on a simple chair, full body in frame. Pulling the camera back shrinks the face and forces the model to commit to proportions.

  4. 4

    Scene 4 — complex outfit

    Add layers: a hooded cloak, gloves, belts. This tests whether identity survives a changed silhouette, the most common real-world edit.

  5. 5

    Scene 5 — hard lighting

    A single hard light source with deep shadow. Tests palette and skin tone stability when the model can no longer rely on flat even light.

  6. 6

    Scene 6 — style swap

    Same everything, different render style. The hardest frame by construction: styles do not share a face grammar, so score this one gently.

Two rules make results comparable between people. Fix the seed across all six scenes, so differences come from your edits rather than from the dice. And keep the unchanged wording byte-identical — retyping a phrase is a second edit you did not intend to make.

The scoring rubric

Score each of the six frames on five attributes. Zero, one or two points each — no half points, because they only ever mean "I could not decide".

Scoring rubric — applied to each scene against Scene 1
Attribute0 points1 point2 points
FaceA different personSame person, visibly re-drawn featuresSame face, same proportions
HairDifferent colour or lengthSame hair, restyled or reshapedSame cut, colour and parting
OutfitSignature outfit or accessories goneKept but altered, or items addedAs specified, accessories intact
ProportionsWrong build, broken anatomySlightly off build or limb lengthConsistent build and height
RecognisabilityWould not identify as the same characterRecognisable after a second lookInstantly the same character

Sixty points is the maximum. Anything above 50 is a genuinely steerable workflow; 35–50 means usable with re-rolls; below 35 means you are not generating a character, you are generating a family resemblance.

Our first run: 47/60

Here is the protocol applied to our own generator, with the frames published so you can disagree with the scoring. One reference (a dark-haired half-elf cleric in silver armour), seed 770125, all six scenes generated in one sitting.

Contact sheet of the six benchmark scenes generated from one character reference
The full run. Scene 1 top left is the control; scoring is relative to it.
Scored run — one reference, six scenes, 60 points maximum
SceneFaceHairOutfitProportionsRecognisableTotal
1 · Neutral portrait (control)2222210
2 · Side profile211228
3 · Seated, full body111014
4 · Complex outfit2222210
5 · Hard lighting211228
6 · Style swap111227
Total47 / 60

The interesting part is not the number, it is where the points went:

  • The complex outfit scene scored full marks, which is the opposite of what most people expect. A hooded cloak with a fur collar over the armour kept the face, the hair and the armour underneath intact. Adding layers is apparently easier than replacing them.
  • The seated full-body frame was the disaster at 4/10. The face survived as recognisable but was re-drawn smaller and rounder, the armour grew into a long robe that was never requested, and the build came back visibly compressed — short arms, a torso that reads as a much younger figure against the chair. Pulling the camera back cost more than changing the clothes did.
  • The profile lost the hair, not the face. Face geometry held up cleanly at 2, while shoulder-length hair became a long ponytail.
  • Hard lighting held the palette but restyled the hair again, which suggests hair is simply the least anchored attribute in the pipeline.
  • The style swap scored 7 and that is a pass, not a failure. Anime and semi-realistic do not share a face grammar; the character stayed recognisable across a change that redraws every pixel of the face.
The seated full-body benchmark frame with compressed proportions and an unrequested robe
Scene 3, the worst frame of the run: recognisable face, invented robe, compressed build.

Two honest caveats about this run. The backdrop was specified as a plain neutral grey studio in five of six scenes, and several frames came back with candlelit gothic architecture instead — an instruction that was simply overridden. And in two frames a small camera-like prop appeared at the belt, traceable to the phrase "facing the camera directly" in the pose wording, which the model read as an object. Both are documented in detail in the piece on why characters change when you change their clothes.

Score your own run

Six scenes, one reference, one seed. Fifteen minutes.

Open the generator

What the test does not measure

Publishing a benchmark means publishing its limits, otherwise it becomes a marketing number:

  • It does not measure image quality, style range or speed. A tool can score 55 and still make images you dislike.
  • It does not cover multi-character scenes. Two characters in one frame is a harder and different problem.
  • It is one character. Distinctive designs with strong canonical outfits behave differently from plain ones — run the test on the kind of character you actually generate.
  • Scoring is human. Two people will differ by a few points. That is acceptable for comparing 47 to 30, and useless for comparing 47 to 49.
  • It rewards the anchored, not the imaginative. A generator that copies your reference exactly and refuses to interpret would score brilliantly here and be useless in practice.
  • It scores five attributes, and identity has more. Apparent age, distinguishing marks and emotional register are not in the rubric — see character consistency is not face consistency for the full set.

Used with those caveats, it does the one thing vibes cannot: it makes "consistent characters" a claim someone can check. If you run it on another tool, the frames and the per-scene table are what make your result worth reading — publish those, not just the total.

For the workflow side of the problem rather than the measurement side, see the guides on consistent characters across scenes and reference-based generation. If you are choosing between tools on price, the attempts you spend fighting drift belong in your cost per keeper.

Prompt field containing the scene wording for one benchmark run
Keep one wording template in the prompt field and edit only the clause each scene requires.
The generator on a phone with the prompt, quantity and aspect ratio controls
The whole protocol runs on a phone too — quantity 1, aspect ratio fixed, one clause changing per scene.

Frequently asked questions

What is the AI character consistency test?

A six-scene benchmark: generate one character as a neutral portrait, side profile, seated full body, complex outfit, hard lighting and style swap, then score each frame 0–2 on face, hair, outfit, proportions and recognisability for a total out of 60.

How do I test whether an AI generator keeps characters consistent?

Fix one reference and one seed, generate the six scenes above changing only what each scene requires, then score every frame against the neutral portrait rather than against your own idea of the character.

What is a good consistency score?

Above 50 out of 60 is a genuinely steerable workflow. 35–50 is usable with re-rolls. Below 35 means the tool is producing a family resemblance rather than the same character.

Why fix the seed during the test?

A fixed seed removes random variation between runs, so a difference you see is caused by the edit you made. It does not preserve identity by itself — that is what the reference does.

Which scene breaks character consistency the most?

In our run, the seated full-body frame. Pulling the camera back renders the face at a smaller scale and forces the model to re-decide body proportions, which cost more points than changing the outfit.

Can I use this benchmark to compare two generators?

Yes, if you keep the character, the seed policy, the wording template and the scoring rubric identical across both, and publish the frames alongside the scores so others can check your judgement.

Benchmark it yourself

Six scenes, one seed, 60 points. Then tell us where we lost points.

Start generating