Based on historical distributions from ~1500 years ago. Estimates appearance, not DNA ancestry or personal identity.

AI & TechnologySeptember 24, 20267 min read

Guess Ethnicity From a Photo: What Gets Measured First

Before a face model names anything, it measures. Here is the geometry stage: the landmark schemes tools rely on, the photo standards that fixed those tolerances, and the two things a phone camera rewrites first.

A catalogued Gracile Indid pattern recorded along the Ganges plain and the Deccan in South Asia

The Measuring Stage Comes First

Every face pipeline begins the same way, and it does not begin with a classifier. A detector finds the head in the frame, an alignment step rotates and scales it to a canonical pose, and a landmark model then places a fixed set of points on the brows, eyes, nose, mouth and jaw. The two schemes most tools draw on are the 68-point annotation set from the 300-W challenge, published by Sagonas and colleagues in 2013, and the 468-point face mesh released with MediaPipe, which adds a denser contour and, in a 478-point variant, landmarks for the iris. From that point onward the classifier never looks at a photograph. It looks at coordinates.

Those coordinates are the same quantities the anthropometric literature spent a century collecting. Interpupillary and intercanthal distance, bizygomatic breadth across the cheekbones, nasal breadth set against nasal height, the height of the upper lip, the slope of the forehead. Leslie Farkas's Anthropometry of the Head and Face, in its 1994 edition, lists more than a hundred such measurements taken on living subjects, which is roughly what a landmark model reproduces in a few milliseconds and without a caliper in sight. The vocabulary changed. The quantities did not.

Which Numbers the Standards Fix

Passport photography settled the tolerances long before machine learning arrived. ISO/IEC 19794-5, the face image standard behind travel documents and biometric enrolment, specifies a full frontal pose with head rotation under five degrees from frontal in pitch and yaw and under eight degrees in roll, an interocular distance of at least ninety pixels, an image aspect ratio between 1.25 and 1.34, and a head length between 70 and 80 per cent of the frame height. ICAO Doc 9303, which governs machine-readable travel documents, works from the same geometry.

Those figures exist because comparison fails without them. A system matching one photograph against another needs both images reduced to the same frame of reference, or the distances it computes are measuring the photographer. Classification of regional resemblance inherits the requirement without inheriting the justification. When an interface asks you to face the camera squarely, keep your expression neutral and avoid a tilt, it is reproducing the passport standard's geometry by other means, and the model downstream reads the same landmarks that standard was written to stabilise.

The Lens Distorts Before the Model Runs

Distance from lens to face is not part of the standard, and it should be. In a research letter published in JAMA Facial Plastic Surgery in March 2018, Boris Paskhover and colleagues modelled average male and female heads as stacks of parallel planes and computed the distortion at different camera positions. At twelve inches, the usual selfie distance, the nose projects thirty per cent larger for men and twenty-nine per cent larger for women. At five feet, standard portrait distance, the distortion is negligible.

Nasal projection is not a decorative detail here. Nasal breadth and nasal height are two of the quantities a landmark model reports, and their ratio is among the oldest measurements in the literature. A camera held at arm's length moves that ratio before any software runs, and it moves the apparent projection of the chin in the same direction, because both sit closer to the lens than the ears do. The model can be perfectly stable and still return a different answer for the same head.

What the Phone Does to the Face Now

The modern input has already been processed. Phone cameras stack multiple exposures, tone-map the highlights, sharpen local contrast and, on many models, apply face-aware smoothing by default. Social platform filters go further and deliberately reshape features, widening eyes and narrowing noses. Hedman and colleagues reported in Pattern Recognition Letters in 2022 that such filters degrade automated face detection and recognition, most severely when they obscure the eye region and less so around the nose. A separate study by Mirabet-Herranz, Galdi and Dugelay found that beautification measurably disrupts estimates of soft traits such as weight and apparent age.

The traits a classifier leans on for regional resemblance are largely the soft ones: pigmentation, nasal width, the shape of the eye aperture, the fullness of the lips. Those are precisely what a beautification filter is built to adjust. A retouched selfie is therefore a picture of an editing preference as much as of a face, and the resemblance it invites may belong to the filter's own template. The practical rule is unglamorous: switch the filter off before asking a tool this kind of question.

One Entry, One Point on a Continuum

The catalogue's own entries argue most clearly against treating a label as a kind. Nilo Hamitic is recorded as an intermediate between Nilotid and Ethiopid, dominant west of Lake Rudolf and north-east of Lake Victoria. Gracile Indid is described as anthropometrically close to Mediterranid while differing in lips, nose, forehead and eyelids, which is a compact way of saying that metric similarity does not imply visual identity. A closed label set still has to choose the nearest name, and the nearest name is a statement about the other entries in the list rather than about the face.

Where an entry describes an isolated population the problem sharpens. Omotic is recorded in south-western Ethiopia and northern Kenya, named for the Omo Valley, and described as a probable relict of pre-Neolithic periods. Its recorded range is small. Rare classes are the ones training sets sample least well, which is why the least familiar pattern is often the one a model handles worst, and why an unusual face can be sorted with confidence into whichever adjacent bucket happens to be best stocked.

What a Measured Face Cannot Carry

Plains Pamirid shows how far a single entry can spread. It is recorded from Uzbekistan across the Ferghana Valley into Xinjiang, formed by contact between Turanid groups and Tungid-influenced populations and dispersed further by pastoral movement. That is a distribution spanning thousands of kilometres, described under one name because the source literature grouped by morphology rather than by politics. Feeding one photograph into a system that returns such a name asks which shelf the geometry fits, not where anyone has been.

The catalogue states its own baseline, describing patterns as documented at roughly 1500 years ago. Measurement can be exact and the conclusion still bounded, because the question a face answers is which documented historical range this geometry most resembles. A single percentage cannot carry that qualification, and the interfaces that show one rarely show how wide the spread of outcomes behind it was.

How to Give the Model a Fair Input

If the question is going to be asked at all, the input is worth getting right. Photograph at a metre and a half or more rather than at arm's length, so perspective stops inflating the nose. Face the lens squarely, inside the tolerance the standards allow. Use even, diffuse light from the front, and remove tinted glasses. Leave the beauty filter off. Make sure the eyes are separated by enough pixels to survive cropping, which is what the ninety-pixel interocular minimum protects against.

None of that makes the output authoritative. It removes a set of avoidable errors large enough to change the answer, which is a smaller claim and an honest one. A reader who wants to know which author recorded a pattern, in what year, and where the sample was taken will get further with a sourced entry than with a percentage over seven buckets.

Frequently Asked Questions

What does a face model measure before it classifies anything?
It places a fixed set of landmark points on the brows, eyes, nose, mouth and jaw, after detecting the head and rotating it into a standard pose. Those points yield the classic anthropometric quantities: interpupillary and intercanthal distance, bizygomatic breadth across the cheekbones, nasal breadth against nasal height. The classifier then works from coordinates rather than from the image, which is why a change in lens distance or a beauty filter can shift the result while nothing about the person has changed.
Do passport photo standards apply to guessing ethnicity from an image?
No, and the distinction is worth keeping. ISO/IEC 19794-5 and ICAO Doc 9303 fix pose, head size and interocular distance so that two photographs of one person can be compared reliably, which is an identity task. A system classifying regional resemblance inherits those preprocessing conventions because it aligns faces the same way, but its output is a similarity judgement against a catalogued pattern. A conforming photograph guarantees stable geometry, not a meaningful label.
How much does camera distance change what a photo reports?
Enough to change the answer. At the twelve-inch distance typical of a selfie, a 2018 modelling study in JAMA Facial Plastic Surgery put the apparent increase in nasal size at about thirty per cent for men and twenty-nine per cent for women. Nasal breadth relative to nasal height is one of the measurements a landmark model reports, so that ratio moves before any software runs. At five feet the effect is negligible, which is why portrait distance is the safer choice.

Related Phenotypes

Faces from the encyclopedia that appear in this article. Open any entry for its full description, distribution, and references.

  • Gracile Indid

    South Asia

    The most widespread and frequent Indid subtype, one of the most populous phenotypes in the world. Common along the Ganges river and Deccan, ...

  • Nilo Hamitic

    East Africa

    Intermediate type between Nilotid and Ethiopid that is associated with the Nilo Hamites or Half Hamites. Similar to Maasai. Should not be co...

  • Omotic

    East Africa

    Isolated type of Southwestern Ethiopia and Northern Kenya, probably belonging to a group of ancient Ethiopids. Probably a relict of Pre-Neol...

  • Plains Pamirid

    Central Asia

    The most widespread Pamirid variety, typical for the steppes and lowlands from Uzbekistan to Xinjiang and the Ferghana Valley. Somewhat infl...

Test what you learned

Put your eye for regional faces to work in the daily quiz — a composite face, a world map, and your best guess.