Back to paper
Two abstract profile silhouettes made of glowing dot-matrix points, facing each other across a warm amber glow, on a dark background

When someone rates a stranger's face, and that stranger happens to be their preferred gender — does the rating go up or down?

Scroll ↓

The Bias That Points the Wrong Way — Twice

Once you separate who's rating from who's being rated, preferring your own gender really does nudge attractiveness ratings up by 0.037 points on a 7-point scale — small, statistically airtight, carried almost entirely by men, and easy to get backwards if you cut a corner.

The Number That Points the Wrong Way

Ask the simple question the obvious way — for each of 1,990 raters, average their ratings of faces matching their stated preference and subtract their ratings of faces that don't — and the answer is -0.139 points. People appear to rate their preferred gender's faces lower. The result is not a fluke: a Wilcoxon signed-rank test puts the odds of this happening by chance at about 1 in 10 sextillion (p = 9.3e-21).

But that comparison is confounded, and once you fix the confound, the sign flips. A crossed statistical model that accounts for both who is rating and which specific face is being rated finds the true congruence effect is +0.037 points on the 1-7 scale (95% CI 0.028 to 0.047, p = 2.6e-13, across 225,420 ratings from 2,210 raters) — small, positive, and just as statistically certain as the naive number, in the opposite direction.

This is the same shape as the most famous reversal in social-science statistics. In 1973, UC Berkeley's graduate admissions showed men admitted at a rate 9 points higher than women overall — a gap "unlikely to be due to chance." Broken down by department, the pattern flipped to a small bias favoring women, because women had disproportionately applied to more competitive, lower-acceptance departments. Here, the confound isn't department — it's which faces happen to fall into the "congruent" and "incongruent" buckets for each rater. Faces are not randomly attractive; some are just more appealing than others regardless of anyone's preference, and that swamps the real signal until it's modeled out.

The data behind this is the Face Research Lab London Set: 102 adult faces photographed in London in April 2012, rated for attractiveness on a 1-7 scale by 2,513 people aged 17 to 90, released by psychologists Lisa DeBruine and Benedict Jones under a CC-BY-4.0 license.

Point-range chart. Unadjusted (per-rater) has no reported CI in the source table.

A cracked mirror shard reflecting warm amber light in one direction and cool blue light in the other, symbolizing a reversal, on a dark studio background
The same evidence, pointing two different directions depending on how you look at it.

Why You Need Two Random Effects, Not One Average

The fix here has a fifty-year-old name. In 1973, psycholinguist Herbert Clark showed that experiments where many people rate a shared, fixed set of stimuli — in this case, faces — routinely treat those stimuli as if their identity doesn't matter, which inflates false positives when generalizing beyond the specific items tested. The modern answer is a "crossed" model: give both the rater and the face their own random-intercept term, so neither one's idiosyncrasies get mistaken for the effect you're trying to measure.

The data itself explains why this matters here. Rater identity accounts for 28.0% of all variance in ratings; face identity accounts for another 24.3% — nearly as much. The two sources are comparably large, which means a method that adjusts for only one of them (as the naive per-rater average implicitly does) is always going to let the other one leak through and distort the result — precisely the mechanism behind the reversal in the previous section.

The raters skew female (61.8% vs. 38.0% male) and, by stated preference, split roughly along expected lines: 49.4% prefer men, 29.8% prefer women, 8.5% report either gender, and 0.3% report neither — with 12.1% not answering the preference question at all and consequently excluded from the primary analysis. The 102 faces themselves are close to evenly split, 53 male and 49 female.

Segments sum to 100% of total rating variance: rater 28.0%, face 24.3%, residual 47.7%.

An abstract visualization of the crossed rater-by-face design: two independent grids of structure, both needing to be accounted for.

The Effect Belongs to Men

Split the +0.037-point effect by rater sex and it stops looking like a single, uniform finding. Male raters show a congruence effect of +0.091 points (95% CI 0.074 to 0.108, p = 3.3e-25, n = 844) — more than double the pooled estimate. Female raters show -0.014 points (95% CI -0.031 to 0.003, n = 1,362): the confidence interval spans zero, meaning the data cannot distinguish this from no effect at all. The interaction between sex and congruence — 0.105 points — is itself nearly three times larger than the headline effect it's splitting apart.

This tracks an established pattern in face-perception research: a 2020 replication study found gay men show significantly stronger preferences for masculinized male faces than straight men do, and straight men show significantly stronger preferences for feminized female faces than gay men do — direct evidence that sexual orientation is a real, replicable moderator of facial preference, not noise.

The picture gets more specific still: a separate study of transgender individuals found that facial preferences track gender identity, not sex assigned at birth — the preferences of male-to-female individuals mirror those typical of their identified gender — which is part of why this analysis is built on stated preference rather than assumed from any demographic proxy.

Whatever is driving the male-only pattern here — and the dataset alone can't say what that is — it is not a small statistical artifact. It survives at a p-value with 24 zeros after the decimal point.

One of these groups shows a real preference-congruence effect. The other shows none at all. Which is which?

It Grows With Age

The congruence effect isn't fixed across a rater's lifetime, either. It rises by 0.0037 points per year of rater age (95% CI 0.0028 to 0.0047, p = 3.8e-14). At the sample's mean age of 26.6, the effect sits at 0.035 points; projected to age 40 it's 0.085; projected to age 60 it's 0.160 — more than four times its value at the mean, and nearly double its value at 40.

Run that same line backward and it crosses zero at roughly age 17 — right at the youngest age recorded in this rater sample. That's a striking coincidence for a straight-line projection, but it is exactly that: a projection, not a measurement at that age. Nobody in this dataset was actually observed transitioning from "no effect" to "effect" at 17; the age pattern is a slope fit across the whole sample, extrapolated to its edge.

Why age would matter here is a genuinely open question the data doesn't answer on its own — the honest reading is that congruence effects, like the sex split before it, are not a constant of human perception but something that varies systematically with who's doing the rating.

Hollow marker = projected zero-crossing, not an observed data point.

Four tones, one per marked point on the chart above, pitch rising with rater age from 220 Hz (age 17, no effect) to 660 Hz (age 60, +0.16 points).

17 27 40 60

Guess the Wrong Shortcut and It Flips Again

Restrict the analysis to only raters with an unambiguous single-gender preference — dropping the "either" and "neither" categories — and the effect barely moves: 0.036 points versus the primary 0.037, a -4.1% shift that's well within noise. So far, so robust.

But swap out how congruence is defined, and everything changes. Instead of using what raters actually said about their preference, assume it from their sex under a heterosexual default — male raters "should" prefer female faces, female raters "should" prefer male faces — and the effect doesn't just shrink, it flips to -0.031 points (95% CI -0.040 to -0.022, p = 2.7e-12). That's not a weaker, noisier version of the same finding; it's a different, opposite, equally significant one.

The lesson lines up with what the broader literature already shows: sexual orientation and gender identity are real, independent predictors of facial preference, not a fact you can safely infer from someone's sex. Guess the shortcut here and you get a statistically confident answer — for the wrong effect.

Two reversals, two different mechanisms: the first (naive vs. adjusted) was about confounding by face identity; this one is about definitional substitution. Both look, on the surface, like solid findings with tight confidence intervals and vanishingly small p-values. Both are wrong if you don't know which one you're looking at.

T3b sits on the opposite side of zero from the primary and T3a estimates.

Checking the Math Twice

It's fair to ask whether a number this small — 0.037 points on a 7-point scale — is just an artifact of a complicated model's numerical optimizer. It isn't: every one of the seven fixed-effect estimates reported here was independently cross-checked against a simpler, optimizer-free benchmark (a within-rater, within-face fixed-effects slope). The two methods agree to within 0.002 points on every single term, including the headline T1 estimate (0.0376 vs. 0.0373 — a difference of 0.0002).

Nor is this a multiple-comparisons trick. All five pre-registered tests — the congruence effect, the sex interaction, the age interaction, and both robustness checks — survive Holm-Bonferroni correction across the full family, with corrected p-values ranging from 3.8e-14 to 2.7e-12. None of them are borderline.

The sample sizes shift a little from test to test — 2,210 raters have a usable stated preference, 1,990 have an unambiguous single-gender one, 2,507 have their sex recorded — because different tests need different fields to be present. None of that thins out enough to threaten any of the numbers above.

Dashed line = perfect agreement (y = x) between the two independent estimation methods.

A Small Number, Loudly Amplified

Zoom back out to the headline number — 0.037 points, 0.62% of the rating scale, roughly a fortieth of the variation between any two individual ratings — and it is genuinely too small to matter to any one person's snap judgment of a single face.

But the same class of bias this dataset isolates in careful, statistically-controlled isolation doesn't stay isolated once it's absorbed into software that makes judgments at scale. Dating-app matching algorithms are widely reported to learn from and reproduce exactly this kind of biased human rating behavior, amplifying it across millions of comparisons rather than averaging it away. A 2025 study of multimodal AI systems found facial attractiveness measurably swayed model outputs on stereotyped jobs and trait judgments in 92.6% of tested scenarios — even when attractiveness carried no legitimate information at all.

An effect too small for any one rater to feel is still large enough, multiplied by millions of ratings and encoded into an algorithm, to become a pattern someone eventually has to explain.

A single point of warm light radiating outward into an expanding network of fine glowing lines on a dark background
A signal too small to notice individually, large enough to shape a pattern once repeated at scale.
0.62% of the 1-7 rating scale — the entire congruence effect, once every confound is removed