Running AI models on a GPU at home: the real numbers from an RTX 5070 Ti
Both tools on this site run on one computer in my home. Here is what I measured: speed, power draw, where the bottleneck actually is, and when self-hosting is not worth it.
2026-08-20 作者 William Hsu
Hardware and model
- GPU NVIDIA GeForce RTX 5070 Ti, 16 GB of video memory
- System RAM 62 GB
- Model MiniMax H3, int8 quantised, with a turbo LoRA that cuts sampling to 4 steps
- Runtime ComfyUI on Windows; the website itself runs in a Docker container on the same machine
The reason for not renting a cloud GPU is simple. An hour of rental costs about what this machine costs in electricity for a whole day, and the photos would have to travel to somebody else's computer. The card is already bought; one more image only costs electricity.
Where the 40 seconds go
Old Photo Motion takes 42–45 seconds per clip, age-me 39–42 seconds per image. I sampled GPU power every 0.7 seconds during a run, and the shape is two peaks with a trough between them.
- First stage: encoding the photo and the text instruction into something the model can read. Nearly 20 seconds, the longest stage of all, and yet the card only runs at about half power
- Middle stage: the actual sampling. Only 4 steps, and the shortest part
- Last stage: decoding the result back into frames, which pushes the card to full load again
Dropping sampling from 6 steps to 4 saved less than you would think. Going from 60–80 seconds to 42–45 was three changes together (a smaller text encoder, fewer sampling steps, and removing an audio decoder that was never used), and most of it came from the encode and decode stages.
Power and what it costs
Measured draw: 28–32 W idle, 144 W average while generating, 276 W peak (the card's ceiling is 300 W).
The card uses about 2 Wh per image; multiply by 1.8 for the CPU, board and power-supply losses and the whole machine is about 4 Wh. At Taiwan's 3.5 to 5.8 NT dollars per kWh, one image costs between 0.014 and 0.023 NT dollars.
That is not where the money goes. Leaving the machine on all day idles at 90–120 W for the whole system, which is 2.2–2.9 kWh a day, or NT$320 to 430 a month. A month of standby electricity is worth fifteen to twenty thousand generated images.
That is why the free allowance is generous: the card is powered up either way.
The bottleneck is moving the model, not the maths
Fitting a 30 GB model into 16 GB of video memory is what the int8 quantisation is for. Even quantised, loading the model is still the single slowest step in the pipeline. The first image after an idle spell takes several times longer than the ones after it, because the model was swapped out and has to be moved back in.
The person who pays for that is whoever arrives first. So the progress text says "the first one is slower" outright, rather than leaving somebody guessing at a percentage that will not move.
The queue
The two tools share one GPU and therefore one queue. Queueing them separately would have both sides fighting over the same video memory, which makes both slower.
The queue position on screen is read from the real queue, not a timer animation.
When self-hosting is the wrong answer
If you need a few thousand images a day, one card cannot keep up and renting is cheaper. Self-hosting suits the case where the volume is modest but you do not want the files leaving the building. This site is the second case: most of what people upload is a photograph of their own family.