Most of us have shipped an image feature at some point, and the consumer AI headshot tools are a compact case study in the parts that are easy to get wrong. I dug into how they work and the engineering choices turn out to be more interesting than the marketing.
The Pipeline
The first step is not generation, it is measurement. The platform maps facial geometry, the distances between the eyes, the jawline, the nose, and a few hundred other landmarks. That map is the anchor every later step generates against, and the resolution of it is most of the reason some tools come back recognizable and others produce a stranger with the same bone structure.
Then a pre-trained diffusion model gets briefly retrained on the specific photos. Minutes to about four hours depending on the platform. Longer runs tend to be more consistent, which is the usual tradeoff between throughput and quality showing up in a place users can actually see it.
Generation swaps the surroundings, studio lighting, neutral background, business attire, while the geometry map keeps identity stable. The last stage is post processing to catch artifacts: unnatural edges, inconsistent lighting, skin tone drift. That stage is where the paid tools separate from the free ones.
The Overfitting Problem Is A UX Problem
The failure mode that surprised me is a training problem that shows up as a product complaint. When users upload fifteen near identical frames from one sitting, the model overfits to a single look and the output comes back stiff and samey.
The fix is not in the model, it is in the upload flow. Photos from different days, varied lighting, multiple angles, no filters. Tools that ask for 6 good photos beat tools that ask for 15 lazy ones. If you are building anything that fine-tunes on user supplied data, the instructions you give at upload time are doing as much work as your training config.
The Part Worth Designing Around
Facial images are biometric data. GDPR covers it in the EU, and Illinois BIPA plus the Texas and Washington statutes cover it in the US. That turns retention into an architecture decision rather than a policy page.
The platforms vary a lot here. Some delete the uploads and the fine-tuned model within 24 to 48 hours of delivery. Some hold data 30 days to allow re-downloads. Some keep it until a deletion request arrives. Aragon AI, HeadshotPro and BetterPic all state that uploads are not used to train their general models, but the retention specifics differ between them.
If you are building in this space, the per-user model artifact is the thing to think about early. It is derived data that encodes a person's face, it usually outlives the request that created it, and a deletion request has to reach it and not just the source images.
I put together a fuller comparison of the platforms, their input requirements and their data policies here: AI headshot generators
Curious whether anyone here has shipped a per-user fine-tuning flow and how you handled expiring the model artifacts.