Search for the best AI art generator and you will find dozens of reviews that were never run. Not fabricated exactly — assembled: pricing pages paraphrased, feature lists reworded, a ranking arranged around whichever affiliate programme pays. They are not obviously wrong, which is what makes them expensive. Here are the seven tells, and the artefact that would settle each one.
1. There are no outputs
A review of image generators that contains no generated images tells you the author read marketing pages. This is the fastest disqualifier and the most common one, because producing real outputs takes credits, time and the risk of your favourite tool losing.
The receipt: images, with the prompt that produced them, and the same brief across every tool in the list. Screenshots of a pricing page are not outputs.
Every claim has an artefact that would prove it. Look for the artefact.
2. There is no method
"I tested 40 generators" is not a method, it is a boast — and at 40 tools it is arithmetically unlikely. Ten generations per tool at three minutes each is twenty hours of pure waiting. Without a stated procedure, "tested" could mean anything from a rigorous protocol to signing up and typing one prompt.
The receipt: what was generated, how many times, judged against what standard, with the scale stated. A method you could rerun and get a comparable number from is the difference between a review and an opinion.
3. The prices have no date
Pricing in this market changes monthly, and credit systems make headline figures nearly meaningless on their own — $9.99 buys wildly different amounts of output across tools.
The receipt: the price, the date it was checked, and what it buys in images per month. If a review says "from $9/month" and nothing else, it cannot help you compare anything, and it will be quietly wrong within a quarter. The metric that survives price changes is cost per keeper.
4. Nobody opened the privacy policy
In adult AI especially, "private and secure" appears in reviews whose authors never opened a policy. The default visibility of your generations and the retention window are both knowable in five minutes, and both matter more than any feature on the comparison table.
The receipt: a quote from the actual policy about default visibility and retention. Both are knowable in five minutes, and a review that skipped them skipped the part that matters most in this category.
5. Chatbots and image generators are in one list
Watch for a ranking that puts a companion chat app, an image generator and a video tool in the same numbered list. These are three different products with three different failure modes, and the only thing that unifies them is that they share affiliate programmes.
The receipt: a stated scope. "The seven best NSFW image generators judged on reference workflow" is a scope. "The seven best NSFW AI apps" is a category laundering exercise.
6. "Free" means "an account"
Free tiers are where reviews are least careful. "Free to use" can mean unlimited generation, three watermarked images, or a signup form followed immediately by a paywall — and reviews rarely distinguish them because distinguishing them requires actually registering.
The receipt: how many usable images a fresh account produces before paying, tested from a clean browser, plus whether the output is watermarked or blurred.
7. Competitors are judged on different tasks
The subtlest tell. The favoured tool is evaluated on what it is good at, while competitors are evaluated on tasks they were never built for — a character-training platform judged on time-to-first-image, or a preset tool judged on inpainting.
The receipt: one brief, applied to everything in the list. If tool A is praised for speed and tool B criticised for a different failure entirely, they were not compared — they were narrated.
Auditing our own listicle
Applying this to first-party content is the only way the checklist means anything, so here is our own reference-first generator ranking against its own seven items.
| Item | Verdict |
|---|---|
| Outputs shown | Fails. It shows product screenshots, not a comparable generated frame from each tool on one shared brief. |
| Method stated | Partial. The criteria are listed — reference workflow, time to first image, style range, content rules, price honesty — but there is no rerunnable scoring procedure. |
| Prices dated | Partial. Prices are named; they are not stamped with a check date, which means they will age badly. |
| Policies read | Passes. Content rules are part of the criteria, and real-person and minor prohibitions are treated as a ranking factor. |
| Scope is clean | Passes. Image generators only, judged specifically on reference workflow. |
| Free tiers described honestly | Partial. 'What the free tier actually gives you' is a stated criterion, but not reported as a tested number per tool. |
| Same task for everyone | Partial. Every tool is discussed through the reference lens, with trade-offs listed — but without a shared brief the comparison is qualitative. |
| Conflict disclosed | Passes. It states plainly that Gensomnia is our product and is ranked first. |
Three passes, four partials, one clear fail. That is a fair verdict on a first-party listicle, and it is more useful to you than a paragraph insisting we are the honest ones. Ranking your own product first is not automatically dishonest — but it does mean the burden of receipts is on us, and on that one row we have not met it.
Our own standing method — fixed seeds, one changed input, published frames, dated numbers — is written down in how we test.
What a review with receipts looks like
The fix for all seven items is the same: publish the artefacts. When we wanted to make a claim about character consistency, the credible form was not an adjective but a protocol — six fixed scenes, a stated rubric, the frames published, and our own score of 47/60 including the frame that scored 4 out of 10.
That is the shape to look for in anyone's review: a number you could reproduce, next to the evidence that produced it. You are free to disagree with our scoring, which is precisely the point — an unfalsifiable review does not give you that option.

Read the protocol in the 6-scene consistency test, and if you want to check a tool's steerability rather than its marketing, the diagnostic guide is the faster hands-on version.
Frequently asked questions
How can I tell if an AI generator review is trustworthy?
Look for artefacts rather than adjectives: generated outputs with the prompts used, a rerunnable method, dated prices, quotes from privacy policies, and one shared brief applied to every tool.
Why do so many AI tool reviews rank the same products first?
Because rankings often follow affiliate payouts rather than testing. A quick check: does the review show real outputs from the top pick and its competitors, judged on the same task?
Is a first-party review automatically biased?
It has a conflict of interest, which is different from being wrong. What matters is whether the conflict is disclosed and whether claims come with evidence you can check independently.
What does 'free' usually mean for AI image generators?
Often an account plus a small number of watermarked or blurred images. Treat 'free to use' as unverified unless the review states how many usable images a fresh account produces before payment.
Why is mixing chatbots and image generators in one ranking a problem?
They are different products with different failure modes, so a single ranking cannot be measuring anything coherent. Mixed lists usually reflect shared affiliate programmes rather than a shared evaluation.
What should a good AI generator comparison include?
One fixed brief run on every tool, published outputs, a stated scoring rubric, prices with a check date and what they buy, a note on default privacy settings, and a disclosure of any commercial relationship.