← 網誌 · Running AI models on a GPU at home: the real numbers from an RTX 5070 Ti

Running AI models on a GPU at home: the real numbers from an RTX 5070 Ti

Both tools on this site run on one computer in my home. Here is what I measured: speed, power draw, where the bottleneck actually is, and when self-hosting is not worth it.

2026-08-20 作者 William Hsu

這篇講什麼
  1. Hardware and model
  2. Where the 40 seconds go
  3. Power and what it costs
  4. The bottleneck is moving the model, not the maths
  5. The queue
  6. When self-hosting is the wrong answer

Hardware and model

The reason for not renting a cloud GPU is simple. An hour of rental costs about what this machine costs in electricity for a whole day, and the photos would have to travel to somebody else's computer. The card is already bought; one more image only costs electricity.

Where the 40 seconds go

Old Photo Motion takes 42–45 seconds per clip, age-me 39–42 seconds per image. I sampled GPU power every 0.7 seconds during a run, and the shape is two peaks with a trough between them.

Dropping sampling from 6 steps to 4 saved less than you would think. Going from 60–80 seconds to 42–45 was three changes together (a smaller text encoder, fewer sampling steps, and removing an audio decoder that was never used), and most of it came from the encode and decode stages.

Power and what it costs

Measured draw: 28–32 W idle, 144 W average while generating, 276 W peak (the card's ceiling is 300 W).

RTX 5070 Ti power draw
Measured power: two peaks (encode and decode) with a trough between

The card uses about 2 Wh per image; multiply by 1.8 for the CPU, board and power-supply losses and the whole machine is about 4 Wh. At Taiwan's 3.5 to 5.8 NT dollars per kWh, one image costs between 0.014 and 0.023 NT dollars.

That is not where the money goes. Leaving the machine on all day idles at 90–120 W for the whole system, which is 2.2–2.9 kWh a day, or NT$320 to 430 a month. A month of standby electricity is worth fifteen to twenty thousand generated images.

That is why the free allowance is generous: the card is powered up either way.

The bottleneck is moving the model, not the maths

Fitting a 30 GB model into 16 GB of video memory is what the int8 quantisation is for. Even quantised, loading the model is still the single slowest step in the pipeline. The first image after an idle spell takes several times longer than the ones after it, because the model was swapped out and has to be moved back in.

The person who pays for that is whoever arrives first. So the progress text says "the first one is slower" outright, rather than leaving somebody guessing at a percentage that will not move.

The queue

The two tools share one GPU and therefore one queue. Queueing them separately would have both sides fighting over the same video memory, which makes both slower.

The queue position on screen is read from the real queue, not a timer animation.

When self-hosting is the wrong answer

If you need a few thousand images a day, one card cannot keep up and renting is cheaper. Self-hosting suits the case where the volume is modest but you do not want the files leaving the building. This site is the second case: most of what people upload is a photograph of their own family.

這篇文章寫的是本站實際的實作與量測。工具本身在老照片動起來與變老變年輕,都可以直接試。

← 上一篇全台 1,807 家露營場只有 203 家合法:把觀光署的名單做成能查的工具,以及查證過程中發現我寫錯的一句話

看其他文章