How We Test: The Method Behind These Articles

Updated August 11, 20265 min read

The short answer

Every number in this blog comes from one of three places: a generation run we did ourselves and published the frames from, our own product's code and pricing configuration, or a public document we read and dated. Where we have not tested something, the article says so. Where a result is unflattering to us, it stays in.

This page exists so the rest of the blog is checkable, and so you can hold us to a standard rather than a tone.

Method last reviewed: 11 August 2026.

The rules we generate under

How a test run is set up

  1. 1

    One question per run

    Every run answers a single question — does an outfit change move the face, does a build clause apply, does a location instruction hold. Multi-question runs produce anecdotes.

  2. 2

    Fixed seed across the comparison

    All frames in a comparison share a seed, so a difference between them is caused by the input we changed rather than by the dice.

  3. 3

    One changed input, everything else byte-identical

    Copied wording, not re-typed wording. 'Neutral grey studio' and 'plain grey studio' are two different inputs.

  4. 4

    Frames published as generated

    No retouching, no cherry-picking the best of a batch when the point is a hit rate. Where we show a contact sheet, it is the whole run.

  5. 5

    Standards written before counting

    For anything scored, the rubric is written down before the frames are judged. A keeper rate defined after seeing the outputs is a mood, not a measurement.

  6. 6

    Failures reported at the same weight as successes

    The scenes and briefs that failed are in the articles, with the frames, because a method that only produces good news is marketing.

What we will not claim

  • We do not claim to have tested products we did not run. No "I tested 40 apps" numbers. When we discuss another platform, we read its published policy and quote its substance with a date — and say that is what we did.
  • We do not claim performance we have not measured. Our own generation time is stated as measured wall-clock, not as a marketing figure.
  • We do not describe enforcement we cannot observe. For mainstream platforms we report what the policy says, not what filters do internally.
  • We do not attempt to circumvent anyone's safety systems, and we do not publish techniques for it.
  • We do not publish figures without a date where the figure can change.

Conflicts of interest

Gensomnia is our product, this blog belongs to it, and the CTAs go to it. That is a conflict, not a secret, and the only useful response to it is evidence you can check. Concretely:

Contact sheet of six benchmark scenes published with their scores
Receipts: the frames behind a published score, including the one that failed.

Content rules for the blog itself

The product is 18+. These articles are deliberately not: they are indexed, they have no age gate, and they are written to be readable by anyone who lands on them from a search. So:

  • Article text and every illustration stay PG-13.
  • No real people, no celebrity likenesses, no "make X look like Y".
  • Nothing involving minors, in text, in prompts, in keywords, or in imagery — and references whose source characters are minors are excluded from illustrations entirely. We dropped a planned comparison row from one article for exactly this reason and said so in the article.
  • Explicit categories are described as categories, never enumerated.

When we get it wrong

Two things happen. The article gets corrected with an updatedAt date, and the underlying note gets corrected in our internal evidence log so the error does not propagate into the next article. If you find a number here that does not match what you observe, that is worth telling us — the frames and seeds are published precisely so disagreement is possible.

Check it yourself

Same seed, one changed input. Two runs will tell you.

Open the generator

Frequently asked questions

How do you test AI generators for these articles?

One question per run, a fixed seed across every comparison, one changed input with all other wording byte-identical, and the frames published as generated — including the failures.

Do you test competitors' products?

Not by proxy. If we did not run a product, we do not describe how it behaves. For other platforms we read the published policy and report its substance with the effective date.

Why do your articles publish your own bad results?

Because a method that only produces favourable numbers is not a method. Our consistency benchmark scores our own generator 47 out of 60, and the worst scene is shown at full size.

Why is every number dated?

Prices, catalogues and policies change monthly in this category. An undated figure is how a stale claim survives for a year — including claims other people make about us.

Is this blog impartial?

No — it belongs to the product it recommends. The response to that is checkable evidence: published frames, stated rubrics, a self-audit that fails an item, and a workflow guide that sends several use cases elsewhere.

Why is the blog PG-13 if the product is adult?

Because the articles are indexed and have no age gate, so they must be appropriate for anyone who lands on them from a search. Techniques are discussed; explicit imagery stays inside the product.