Failure cases: same photo, very different results

Every comparison below was actually generated. Both sides use the exact same photo and differ by one setting. One of them disproved an assumption we had written into our own guide — that's on this page too, unedited.

Case 1: no person in the photo, but a portrait motion

Both sides are the same CC0 photo of a grey catbird on a branch. Left: People → Quiet breathing. Right: Animals → Quiet breathing.

Portrait motion (wrong)
Animal motion (correct)

Watch the left one to the end: a human face rises from the bottom of the frame, grows until it fills the shot, and the bird is still perched on top of his face. The right one, same photo, is just a bird breathing.

Why. The portrait prompt reads "The person stands still and looks straight ahead. They blink naturally…". The model finds no person in the image — and rather than give up, it makes the image match the sentence by producing one.

We first hit this when a user uploaded a bird photo with a portrait motion: the first frames were the bird, then a man's face grew in from the right and smiled at the camera for three seconds. The clip above is a reproduction with properly licensed material — and it reproduced on the first try. This is not a rare glitch.

The fix. The animal prompts never contain the word person; they say "the animal" and end with No people appear. Same photo, same model, different subject noun, correct result. Uploads also run a face check now, so you get warned before you spend forty seconds on it.

Case 2: we assumed low resolution would break it. It doesn't.

Practically every guide for tools like this tells you the source must be sharp. We wrote that too — then tested it. We squashed the 1863 Lincoln portrait down to 96 pixels wide (about the size of an app icon), blew it back up to the original dimensions, and ran the same "Smile and nod" on both.

Source crushed to 96px wide
Original high-resolution file

The result: you can barely tell them apart. Both smile naturally, no drifting features, no breakdown. This model is far more tolerant of input resolution than we assumed — it clearly understands "this is a face and here is roughly where the features are" and then regenerates the detail, rather than extrapolating pixel by pixel.

So we changed the advice in our guide. What actually breaks a result is a semantic mismatch — no person but a portrait motion, two subjects in frame, a covered face — not image quality. Phone snapshots of old prints and small images from messaging apps are more usable than people think.

We left this on the site because it is worth more than "use a clear photo", which is true but useless. You only find out which pieces of common knowledge hold up by testing them.

Things that genuinely do go wrong

The shot slowly pushes in

An old habit of the model. Counter-intuitively, writing "don't zoom" or "static shot" in the prompt makes it worse — measured background displacement was 6.6 with no mention of the camera, 20.6 with "static shot", and 39.2 with a full locked-camera sentence. So the prompt never mentions the camera. A looser composition helps.

Bent hands, extra fingers

Mostly "Wave hello", because the hand doesn't exist in the source — it's invented. Avoid it on headshots.

Two people twitching together

With more than one person in frame the model doesn't know who to animate. Crop the subject out.

A covered face

Masks, hands and microphones get treated as part of the face and deformed. One of the few genuinely unfixable cases — use another photo.

Warped text in the background

Signs and lettering get animated as texture. Use a cleaner background.

All source material on this page is public domain or CC0; full credits are on the examples page. Your uploads are never used for demos.

Read the full photo guide