How this is actually built, including a few decisions that are not obvious until you have run into the problem they solve.

Model and pipeline

MiniMax H3 running image-to-image, int8 quantised, with a turbo LoRA that cuts sampling to 4 steps. It runs on our own RTX 5070 Ti and takes about 40 seconds per image. The photograph never leaves this machine, and it is deleted along with the result after seven days.

Three kinds of change need three sets of instructions

Age editing is really three different jobs, and one block of text does all three badly:

So there is a routing table that picks a set of wording from the current and target ages, along with how loose a parameter called source fidelity should be: 0.40 when the bone structure has to move, 0.80 and above when only skin changes. Any looser and the model simply hands back a different face.

A 10 year gap and a 40 year gap cannot share a sentence

The going-older path had always been written in four bands by size of gap. The going-younger path had not, and said "nasolabial folds become fainter" regardless. Asked to take 70 down to 30, the model did as it was told: slightly fainter folds, a face still old, and darker hair.

Split into bands, anything over a 20 year gap now states absolutes rather than degrees: no nasolabial folds, no crow's feet, no eye bags, none of the lines visible in the original. That is what makes the model repaint the skin.

The first run after that change removed the moustache along with everything else. Push the instruction hard and the model reads facial hair as something to smooth away, so there is now a line holding the shape, size, position and thickness of beards, moustaches and eyebrows in place.

Expressions drift, and it is not solved

While changing age, the model occasionally changes the expression too: a stern face ends up faintly smiling, the eyes narrowed. The instructions have always said the expression, eyes and mouth must match the original, and that lines are carved into the skin rather than smiled into it. A harder version of that rule (do not narrow the eyes, do not smile if the original is not smiling) gave the same output on the same seed.

The problem is not the wording. This is what the model does at the current fidelity setting. If you hit it, run the image again — every run uses a different seed.

Why child-to-adult checks how much of the frame the face fills

If the body is visible, the growth template has to say that the whole body grows and the head becomes proportionally smaller. On a head-and-shoulders shot, that same wording makes the model invent a body and print the child's face onto the clothing, which has happened in testing. So a face detector measures the face against the frame on upload, and the full-body template only comes out below 20%.

Where the 40 seconds goes

About 40 seconds an image, and it is not spread evenly. Measured node by node, encoding the photograph and the instruction into something the model can read takes nearly 20 seconds and is the longest stage. The sampling itself is only 4 steps and is the shortest. Decoding the result back into an image takes another ten-odd seconds.

Cutting sampling steps therefore helps far less than intuition suggests. The bulk of the time is encoding, decoding, and moving model weights into video memory. A 16 GB card cannot hold a 30 GB model, so the weights are moved on every run, which is also why the first image after an idle spell is slow.

Why a general image-to-image model rather than a dedicated age model

There are models built specifically for face age transformation, usually GANs trained on face datasets. They do well on front-facing head shots, with two catches: they take a cropped face, so background and clothing are discarded, and whatever age and ethnicity distribution the training set had shows up in the output.

Using a general image-to-image model with carefully written instructions costs speed and means controlling fidelity by hand. What it buys is that the background, clothing and framing all survive, and that one pipeline handles monochrome prints, three-quarter views, hats and glasses. For looking at photographs from a family album, keeping the picture intact matters more than a perfect face.

The same photograph twice gives different results

Every run uses a different random seed, so the same photograph with the same settings varies between runs. The variation sits in the details and the expression; the size of the age change does not wander.

That is useful: if the expression or some detail is off, running it again usually fixes it without touching the photograph or the ages. It also means a particular result cannot be reproduced on request.

How the lower limit of 16 is enforced

In the routing check on the server, not by hiding the option in the page. A number field in the browser can be worked around; the server cannot. A target age below 16 is rejected and never reaches the queue.

Face detection is used in two places

A YuNet face detector runs on upload and returns the height of the largest face as a fraction of the frame. That number does two jobs: telling us whether there is a person in the photograph at all (and warning you if not), and choosing between the head-shot and full-body growth templates for child-to-adult, where the threshold is 20%.

When detection is unreliable — a profile in a hat, for instance — the upload goes through anyway. Letting a few photographs produce odd results is better than turning away normal ones.

Content safety

Uploads and results both go through an automatic check. Anything judged inappropriate is stopped and quarantined, never reaches the queue, and never appears anywhere. The details are in the privacy policy.

Back to the tool

Other tools on this site

All run on one machine at home; the data tools use Taiwan government open data and state their limits.