How to pick a photo that won't come out weird
Same model, different photo, wildly different result. This page is the set of rules we actually measured, so your first attempt is usable instead of your fifth.
1. What works best
A big face — but "sharp" matters less than you'd think
The face should fill at least a quarter of the frame height so the model knows what it is animating. For group photos, crop the subject out and upload that.
About image quality: we used to write "the source must be sharp or the face will drift". Then we tested it — we squashed a portrait down to 96 pixels wide and blew it back up, and the result was nearly indistinguishable from the full-resolution original (side-by-side clips on the failure cases page). So phone snapshots of old prints and small images saved from messaging apps are far more usable than you'd expect. Better quality is still better — it just isn't the thing that decides success.
Front-on or three-quarter view
A pure profile (one eye visible) is the riskiest, because the model has to invent the other half of the face and the invention often doesn't look like the person. The one exception is "Turn and look", which is designed for exactly that.
Even lighting
Backlight, hard light from one side and noisy night shots all give the model shadows it mistakes for facial features, and then it deforms them. Window light indoors is the most reliable.
One subject per photo
With two or more people in frame, the model doesn't know who should move. You typically get everyone twitching slightly, or features from two faces blending together.
Avoid these
- Anything covering the face — masks, microphones, a hand. Whatever is behind the obstruction is invented by the model
- Heavy beauty filters — over-smoothed skin has no texture, and the result moves like a wax figure
- Photos of screens, or anything with moiré — the pattern shimmers once it moves
- Illustrations and anime — not impossible, but the motions were designed around real faces
2. Choosing among the six portrait motions
| Motion | Amount | Best for | Risk |
|---|---|---|---|
| Quiet breathing | Tiny | Anything, especially poor-quality photos | Almost none |
| Smile and nod | Small | Flat, expressionless ID photos | Low |
| Raised eyebrow | Small | Close-ups with clear features | Invisible if the brows are covered |
| Turn and look | Medium | Profile / three-quarter view | The revealed half may not match |
| Burst out laughing | Medium | Mouth already slightly smiling | Teeth can smear |
| Wave hello | Large | Half-body shots with room below the shoulders | Hands can come out wrong; on a headshot the hand covers the face |
When unsure, run "Quiet breathing" first. It almost never fails — once you've confirmed the face holds up, move to a bigger motion.
3. When the photo isn't a person
The motion list is split into People and Animals & other. That split is not cosmetic.
A real case. We originally only had portrait motions. Someone uploaded a photo of a bird on a branch and chose "Quiet breathing". The first frames are still the bird — then a stranger's face grows in from the edge of the frame, the shot pushes in on him, and he smiles at the camera for three seconds.
The model didn't break. We gave it a contradictory instruction: there is no person in the photo, but the prompt says "The person stands still and looks straight ahead…". The model does not ignore that
person— it produces one.
So the animal prompts never contain the word person at all. They say "the animal", and they end with
No people appear. The same bird photo with an animal motion just quietly turns its head.
On top of that, uploads run through a face detector. If you pick a portrait motion and no face is found, you get a warning suggesting the animal motions — you can still override it, you just get told first. When detection is unsure (a hat, a profile) it lets you through rather than blocking you.
4. What failure looks like, and what to do
The face drifts away from the person
Check first that there is only one subject and nothing covers the face — those two matter far more than resolution. If both are fine, drop to "Quiet breathing" so the model has less to invent.
The shot slowly pushes in
An old habit of the model. The counter-intuitive part: writing "don't zoom", "locked camera", "static shot" in the prompt makes it worse. We measured background displacement: 6.6 with no mention of the camera, 20.6 with "static shot", 39.2 with a full locked-camera sentence. So the prompt now never mentions the camera at all. If you still see it, a looser composition helps.
Bent hands, extra fingers
Mostly on "Wave hello", because the hand isn't in the source at all — the model invents it. Use a motion that doesn't need hands, or a photo where the hands are already visible.
Things in the background move too
Leaves, water and fabric will be animated along with the subject, which usually looks good. Text on signs, however, will warp — use a cleaner background.
You can see the loop turn around
The output is forward + reverse, so the seam itself is perfectly continuous. But directional motions (waving, calling) do read as "rewinding" on the way back. Pick symmetrical motions — breathing, blinking, smiling — if that bothers you.
5. About the frame
The gold / wood / silver frame around the result is drawn by the web page in CSS, not by the model. We tried describing a picture frame in the prompt: the model duly painted one into the image, crooked, wobbling, corners not meeting. Overlaying it in the page instead means it is always straight and you can change it without re-running anything.
Which also means the mp4 you download has no frame in it — it's the clean video.
6. A checklist
- Face fills at least a quarter of the frame height, features readable
- Front-on or three-quarter view, both eyes visible
- Even light, no half-black face
- Only one subject
- Nothing covering the face
- Resolution isn't critical, but the original file always beats a screenshot
- Run "Quiet breathing" first to see how the photo behaves