How we test LooksCompare before trusting a result
A face-comparison tool should be challenged with cases where the expected answer is known before the test. This page explains the manual checks we use and the limits of what those checks can prove.
Why we publish our testing process
Face-comparison tools can produce a precise-looking percentage, but the practical question is whether the result stays sensible when the photos become more difficult. LooksCompare is tested manually with photo pairs where the expected relationship is already known before the tool is run. That avoids judging a result only because it “looks convincing” after the fact.
Our testing is not presented as a laboratory benchmark or a universal accuracy claim. It is a development process used to look for obvious failures, unstable behavior and photo conditions that can change a result. We prefer to describe what was tested and what we observed rather than publish an unsupported percentage.
Three types of checks we use
Same person across time
We compare known same-person photos taken years apart, including childhood-to-adult examples. This checks whether normal aging, headwear, glasses and camera differences cause an obvious identity failure.
Visually similar, different people
We also use known different-person pairs that appear similar to a human observer. These are useful because a system should not call two people the same simply because several visible traits overlap.
Photo-condition changes
We repeat comparisons with changes in lighting, angle, sharpness, crop and expression. The goal is to understand how much the image itself influences the score.
Family comparisons
For mom-dad-child comparisons, we look for stability and sensible movement rather than treating a family resemblance score as genetic proof. Close scores are deliberately interpreted as close.
What we consider a useful result
- Same-person checks: the result should remain consistent enough that ordinary changes in age or appearance do not create an obvious false “different person” result.
- Different-person checks: visually similar people should not be treated as the same person merely because they share a few facial traits.
- Repeated-photo checks: the score may move, but the movement should have an understandable reason such as pose, lighting or image quality.
- Family checks: the result is a visual resemblance estimate. We do not treat a 60/40 result as biological evidence and we do not force a strong winner when the two comparisons are close.
Examples from manual development checks
During manual development checks, LooksCompare has been challenged with known same-person photos separated by many years, including childhood and adult photos. We have also used examples where the newer photo includes glasses or the older photo includes headwear. Those checks are useful because the system must rely on facial structure rather than one accessory or one moment in time.
Another useful check is a “look-alike” pair: two different people who appear unusually similar. A good result is not simply the highest possible similarity score; it is a result that still distinguishes identity when resemblance is strong. These stress tests are one reason we keep the same person feature conceptually separate from the family-resemblance feature.
What we do not claim
We do not publish a 95%, 99% or 100% accuracy claim without a controlled, representative benchmark that supports it. Manual testing can tell us whether a change appears stable and useful, but it does not replace a formal biometric evaluation. We also do not use LooksCompare for legal identity, paternity, medical, immigration, security or forensic decisions.