← 網誌 · How a still photograph becomes a 6-second loop: implementation notes and the traps

How a still photograph becomes a 6-second loop: implementation notes and the traps

This is about implementation, not a primer on how diffusion works. What each step actually does between a photo going in and a video coming out, and which plausible-looking choices turned out to be wrong.

2026-08-20 作者 William Hsu

這篇講什麼
  1. The pipeline
  2. The ping-pong loop
  3. Three approaches that measurement proved wrong
  4. Handling failure took more work than the model did

The pipeline

  1. Receive and check: verify the format and size, run a content-safety check. Portrait motions also run a face detector. If the photograph is of a bird and a portrait motion is chosen, the model invents a human out of nowhere.
  2. Encode: the photograph and the motion instruction are encoded together. This is the slowest step, nearly 20 seconds.
  3. Sample: only 4 steps, thanks to the turbo LoRA.
  4. Decode: turn the result back into pictures — 73 frames.
  5. Post: 73 frames at 24 fps is 3.04 seconds; played forward then backward it becomes a 6.08-second seamless loop.

Why 73 frames? This model's frame count must be 17n+5, and 73 is the only usable value near three seconds. That is not in the paper; you find out by hitting it.

What going wrong looks like: the photo is a bird, but a portrait motion was chosen, so the model grows a human out of nothing (the same photo with an animal motion is fine)

The ping-pong loop

The head and tail of a generated three-second clip rarely match, so looping it directly shows an obvious cut. Playing it forward and then backward makes the two ends identical and the seam disappears.

A finished result: a 6.08-second ping-pong loop with no visible seam

The cost is a back-and-forth feel, so the motions have to avoid anything with a clear direction. "Wave" is fine — the hand comes back anyway. "Turn to look at you" has to be designed to stop once it has turned, or playing it backwards turns the head away again.

Three approaches that measurement proved wrong

1. Telling the model not to move the camera makes it push in harder

The instinct is to add "camera does not move" to the prompt. Measurement said the opposite: with that line the background displacement went from 6.6 to 39.2. Most likely the phrase puts the model's attention on the idea of a camera in the first place. The final prompt never mentions the camera and describes only what the person does.

2. Let the model draw a picture frame and the frame comes out crooked

The plan was for the result to come with a vintage frame built in. The model can draw one, but it is different every time and never symmetrical. Overlaying it with CSS in the page instead means the frame is always right, the viewer can change or remove it whenever they like, and the downloaded file stays clean.

3. Low-resolution photographs do not necessarily break

I assumed a very small photograph would fail, and even wrote that on the guide page. Then I tested a 96-pixel image: it animated perfectly well, just with soft detail. I overturned my own claim and rewrote the page.

Handling failure took more work than the model did

Getting the model running was a small part of the work. The rest was handling what the real world does:

None of these is a model problem, and every one of them makes the site look broken.

這篇文章寫的是本站實際的實作與量測。工具本身在老照片動起來與變老變年輕,都可以直接試。

← 上一篇為什麼你上傳的照片不該留在雲端:一天被掃描 417 次的實際紀錄

看其他文章