Before and after
Image to Image, Compared
Two results side by side are only evidence if everything except the thing being tested was held still. With image-to-image that is unusually hard, because the tools do not share a scale for the one setting that matters most.
Method
Four things to pin before comparing anything
Leave any one of them floating and the comparison is measuring that instead of the tool.
The source
The same image, at the same size
Resolution and compression change what a model has to work with. Feeding one tool a full-size original and another a downscaled copy is the most common way a comparison is decided before it starts.
The prompt
Identical wording, and its limits
Identical prompts are the fair-looking choice and are not straightforwardly fair: models respond to different phrasing, so one written against a particular tool will flatter it. The honest move is to use the same words and say so.
The strength
The setting that will not line up
How much of the original survives is the dominant variable, and it is exposed differently everywhere — a 0–1 scale, a percentage, a three-position selector, sometimes nothing at all. Two tools at “0.5” are not necessarily doing the same amount of work, and this is the single biggest obstacle to a clean comparison.
The sample
More than one generation each
Output is stochastic. One generation per tool compares luck; several per tool, with all of them shown rather than the best one, compares tools. A comparison that does not say how many runs it is based on has not told you its most important number.
Honest limits
What a comparison like this cannot settle
Worth stating plainly, because a confident verdict on any of these would be overreach.
Which one is better
Beyond obvious failures, this is taste. A result that is more faithful to the source and one that is more interesting are both defensible, and which you want depends on the job rather than the tool.
How it will behave on your image
Models have subject matter they handle well and subject matter they do not. A comparison run on portraits says little about architecture, and less about diagrams or text.
Whether it will still be true next month
Hosted models are updated without notice and often without a version number. Any observation about output quality has a shelf life, and the date it was made is part of the finding.
What happens at settings nobody tried
A tool that does poorly at one strength may do well at another. Conclusions apply to the settings used, which is why stating them matters more than the pictures do.
House rules
How this will be written
No invented numbers
No scores out of ten, no star ratings, no tier boards. Where something is better, it gets said in words and the reason is shown.
Only claims we can stand behind
Nothing describes a test that did not happen. Pricing and feature claims are checked against the vendor's own pages.
No paid placements
Nothing here is sponsored, and outbound links carry no affiliate tracking and earn no commission. No company pays to be included or to be written about more kindly.
This site is new and no comparisons have been published. The method above is what they will follow when they are — source, prompt, strength and run count stated every time, so anyone can tell what was actually held still.