Two routes lead to short AI video of a character you care about. One is WAN 2.7 — an open model you download and run on your own hardware. The other is Gensomnia — a browser product where you start from a character reference and pick a scene. This is a head-to-head on the four things that actually decide it: speed, cost, effort, and whether the result still looks like your character.
One framing note, because it keeps the rest honest: this compares a model you operate with a product you use. It is not two models benchmarked against each other — nobody wins a pixel contest here, and anyone who tells you otherwise is selling something.
The short answer
Pick WAN 2.7 if you have a capable GPU, want total control over resolution and length, and treat the pipeline as part of the craft. Pick Gensomnia if you want a specific character in a new scene today, from whatever device you have, with identity handled for you. Faster, cheaper and easier describe Gensomnia for most people most of the time — and "more powerful" describes WAN, which is exactly why the choice is not obvious.
Effort: what each route asks of you before the first frame
WAN 2.7 asks for hardware first. Community setup guides converge on roughly 24 GB of VRAM as the practical floor for a usable workflow, with the 14B-class model wanting far more at higher resolutions — figures in the 40–48 GB range at 480p with FP8 quantisation, and higher again at 720p. Quantised GGUF builds let smaller cards join in, at the cost of longer runs and softer output. Then come the checkpoints: 30–40 GB per variant, closer to 100 GB if you want text-to-video, image-to-video and image models side by side. Then the graph — loaders, encoders, samplers, offload toggles — where most early failures actually live.
Gensomnia asks for a browser. There is no VRAM table to read, nothing to download, and no graph. It works on a phone, which matters if your main machine is a laptop.
Speed: time to a usable frame
Self-hosting has two clocks. The first is one-off — install, download, wire the graph, discover which node is unhappy. Budget an evening if nothing breaks. The second is per clip, and it runs in minutes rather than seconds even on datacentre hardware: a widely repeated figure is around four minutes for a five-second 480p clip on an H100. On consumer cards you plan around a run instead of iterating against it.
In Gensomnia images come back in 10–15 seconds. That difference is not just convenience — it changes method. You use stills as a viewfinder, throw away the framings that read wrong, and only commit to video once a frame already works.
Cost: two different shapes, not two numbers
WAN 2.7 is capital plus electricity. You buy the card — or rent one by the hour — and after that runs feel free. If you generate constantly, that is the cheaper end state, and nothing hosted will beat owning silicon at volume.
Gensomnia inverts it: nothing up front, free to start, and you only spend once you want more than the free start gives you. For a handful of clips a month, that is cheaper by a wide margin, because you never paid for a GPU that idles.
The honest version of "cheaper" is therefore a question about you: how many clips will you actually make? Few — hosted wins. Constant heavy output — hardware wins.
Consistency: the axis that actually differs
Everything above is logistics. This is the part that changes the work.
In Gensomnia identity comes from a reference image, not from adjectives in a prompt. You pick a character from the collection, upload your own reference, or describe one in words — and that reference is what carries the face, the hair and the distinctive details into the next scene. Changing the location, the outfit or the light does not restart the character.
The four frames below were all generated in Gensomnia for this article — a bedroom doorway, a balcony above the city, a window seat and a rose garden. Same face, same hair, same shirt, four different sets. That is the behaviour you are buying: the character stops being something you have to re-describe every time.

And the same character in motion, because a still cannot show whether identity survives movement — the honest test for anything calling itself video:
A local WAN workflow can reach comparable identity lock, but not out of the box: you bolt on an adapter, train a LoRA, or build a character model. That is more capable in the end and more work in the middle — which is the trade-off in one sentence.
And the honest limit on our side: a reference is instant but weaker than a trained character over dozens of frames. Our own six-scene test on images scored 47 out of 60, with the seated full-body scene at 4 out of 10 — proportions are where a reference gives up first. The frames and the rubric are published.
Scenes: written, or chosen
WAN 2.7 takes instructions as text, and its improved instruction following is one of the genuine advances of this generation. The flipside is that your result quality tracks your prompt-writing skill.
In Gensomnia the scene is a set of choices: 50 locations, 50 poses, 50 outfits and four render styles, with new ones added every day. Locations are built sets rather than blurred backdrops — the three above came from stock library entries, not bespoke work for this article. If you want the free-text route it is there too, up to 1,000 characters with a negative prompt.
Side by side
| WAN 2.7 | Gensomnia | |
|---|---|---|
| What you need | A capable GPU — guides suggest 24 GB VRAM and up | A browser, including on a phone |
| Disk | 30–40 GB per variant, ~100 GB for a full set | None |
| Setup | Install, checkpoints, node graph, quantisation choices | None |
| Time to first frame | An evening, if nothing breaks | Seconds after you open the page |
| Per-image time | Minutes, hardware-dependent | 10–15 seconds |
| Character identity | Add adapters, LoRAs or training | Built in — anchored to a reference |
| Scene setup | Written as prompt text | 50 locations, 50 poses, 50 outfits, 4 styles |
| Video | Yours to configure — length, resolution | Short-form: 5 seconds at 720p |
| Control | Total: models, samplers, resolution, length | Product-level choices only |
| Cost shape | Hardware up front, runs effectively free | Nothing up front, free to start |
| Content limits | Whatever your machine allows | Fictional characters, adults only |
Getting a character into a scene without a rig
How to generate consistent character video from a reference
- 1
Start from the character
Pick one from the collection, upload a permitted fictional reference, or describe one in words. This step decides whether the result still looks like your character three scenes later.
- 2
Choose the scene, do not describe it
Set pose, outfit and location from the preset libraries, then a render style. Leave the fields you do not care about empty — fewer competing instructions usually reads better than more.
- 3
Use stills as a viewfinder
Images return in 10–15 seconds. Find the framing that reads correctly before you commit it to a clip.
- 4
Move the keeper into video
Switch to video with the same reference and scene. Short-form output means you learn immediately whether the character survives motion.
Where WAN 2.7 is simply the better answer
Not a courtesy paragraph — these are real reasons to walk away from a hosted product:
- You want the pipeline. Custom checkpoints, samplers, control nets, your own LoRAs. None of that is available in a product, by design.
- You want ceilings you set. Clip length and resolution are yours. Gensomnia's video is deliberately short-form.
- You want to train identity, not anchor it. A trained character beats a reference over dozens of frames.
- You want no policy layer. Gensomnia is fictional characters and adults only, and no real, identifiable people. A local rig has no such gate — and that is precisely the line we will not help anyone cross.
Frequently asked questions
What are WAN 2.7's requirements?
Community setup guides put the practical floor at roughly 24 GB of VRAM for a usable workflow, with higher-resolution runs on the 14B-class model wanting considerably more. Checkpoints run 30–40 GB per variant, so budget around 100 GB of disk for a full set.
Is Gensomnia faster than running WAN 2.7 yourself?
For getting to a usable frame, yes. Images return in 10–15 seconds with no setup, while a local run is measured in minutes after an evening of installation. For raw configurability, the comparison does not apply — that is what self-hosting is for.
Which is cheaper?
It depends on volume. Self-hosting means paying for hardware once and then generating freely; Gensomnia costs nothing up front and is free to start. A handful of clips favours hosted, constant heavy output favours owning a GPU.
Can I run WAN 2.7 on a laptop?
Usually not comfortably. Quantised builds let smaller cards produce shorter, softer clips, but a typical laptop GPU sits below the practical floor. A browser-based generator is the realistic route on laptop hardware.
How do I keep the same character across several clips?
Anchor identity to a reference image rather than to a text description. In Gensomnia that is the default behaviour; in a local WAN workflow it means adding an adapter, a LoRA or a trained character on top of the model.
The verdict
If you enjoy owning infrastructure, WAN 2.7 is the more powerful answer and the control is real. If what you want is your character, in a new scene, today — from a laptop, a phone, or whatever is in front of you — Gensomnia gets you there while the checkpoints would still be downloading. Faster, cheaper for normal volumes, and easier are all true for the hosted route; more powerful is true for the rig. Pick the one that matches what you actually enjoy doing.
Skip the setup
Start from a character, choose a scene, and get frames in seconds — free to start.
Try GensomniaExternal figures for WAN come from public community setup guides, not from our own benchmarking: open-source setup guide, local GPU guide. Our own numbers and method live in how we test.