How a still photograph becomes a 6-second loop: implementation notes and the traps
This is about implementation, not a primer on how diffusion works. What each step actually does between a photo going in and a video coming out, and which plausible-looking choices turned out to be wrong.
2026-08-20 作者 William Hsu
The pipeline
- Receive and check: verify the format and size, run a content-safety check. Portrait motions also run a face detector. If the photograph is of a bird and a portrait motion is chosen, the model invents a human out of nowhere.
- Encode: the photograph and the motion instruction are encoded together. This is the slowest step, nearly 20 seconds.
- Sample: only 4 steps, thanks to the turbo LoRA.
- Decode: turn the result back into pictures — 73 frames.
- Post: 73 frames at 24 fps is 3.04 seconds; played forward then backward it becomes a 6.08-second seamless loop.
Why 73 frames? This model's frame count must be 17n+5, and 73 is the only usable value near three seconds. That is not in the paper; you find out by hitting it.
The ping-pong loop
The head and tail of a generated three-second clip rarely match, so looping it directly shows an obvious cut. Playing it forward and then backward makes the two ends identical and the seam disappears.
The cost is a back-and-forth feel, so the motions have to avoid anything with a clear direction. "Wave" is fine — the hand comes back anyway. "Turn to look at you" has to be designed to stop once it has turned, or playing it backwards turns the head away again.
Three approaches that measurement proved wrong
1. Telling the model not to move the camera makes it push in harder
The instinct is to add "camera does not move" to the prompt. Measurement said the opposite: with that line the background displacement went from 6.6 to 39.2. Most likely the phrase puts the model's attention on the idea of a camera in the first place. The final prompt never mentions the camera and describes only what the person does.
2. Let the model draw a picture frame and the frame comes out crooked
The plan was for the result to come with a vintage frame built in. The model can draw one, but it is different every time and never symmetrical. Overlaying it with CSS in the page instead means the frame is always right, the viewer can change or remove it whenever they like, and the downloaded file stays clean.
3. Low-resolution photographs do not necessarily break
I assumed a very small photograph would fail, and even wrote that on the guide page. Then I tested a 96-pixel image: it animated perfectly well, just with soft detail. I overturned my own claim and rewrote the page.
Handling failure took more work than the model did
Getting the model running was a small part of the work. The rest was handling what the real world does:
- A phone freezes the tab, the front end assumes the connection dropped and gives up — while the back end has run all the way to the end. Job state is now written on both the browser and the server, so returning reconnects automatically
- Somebody presses cancel while the job is still on the GPU. The cancel has to reach the generator and actually interrupt it, otherwise the card keeps spinning and everybody in the queue waits for nothing
- The output needs a content-safety check too. Whatever gets blocked on upload has to be blocked on the way out
None of these is a model problem, and every one of them makes the site look broken.