Based on historical distributions from ~1500 years ago. Estimates appearance, not DNA ancestry or personal identity.

AI & TechnologyOctober 4, 20267 min read

Ethnicity Guesser Photo: What Happens to the File First

A photo handed to a face tool stops being a photograph almost at once: the file is decoded, rotated, colour-converted and resized before any model runs. Knowing where the losses happen changes how you read the score.

Portrait photograph cropped to a square frame beside enlarged pixels showing colour and texture lost to resampling

The upload is a channel, not a handover

Clicking upload feels like handing over an image. It is closer to pushing a file through a chain of filters, and every filter is permitted to change something. The bytes are encoded, transmitted, decoded, checked for a face, rotated, cropped to a square, resampled to a fixed size, colour-converted and finally turned into an array of numbers. A model never sees your photograph. It sees the last array in that chain, and the changes introduced earlier are invisible in the response.

This matters because the questions a face tool is asked are visual while a large share of the answers are pipeline artefacts. Two uploads of the same person from the same phone can produce different scores if one was edited and re-saved and the other was not. Knowing where the losses happen is the difference between reading a score and guessing at one.

A phone photo is a container, not a picture

Since iOS 11 and macOS High Sierra in 2017, iPhones and Macs have saved photographs as HEIC by default, a format that stores images in the HEIF container defined by ISO/IEC 23008-12 and compresses them with the HEVC codec borrowed from video. The container holds more than one image: a primary frame, a thumbnail, EXIF and GPS metadata, and often a depth map from the dual cameras. The file is roughly half the size of an equivalent JPEG, and it carries 10-bit colour in the Display P3 gamut rather than 8-bit sRGB.

That combination is why so many uploads do not arrive as they left. Very few browsers open HEIC directly, Safari having done so since version 17 and most others not at all. When a service accepts a HEIC and returns a JPEG, a transcode has occurred, and the transcode is where wide gamut gets flattened and the extra bit depth is discarded. A face tool describing itself as reading a photograph is often reading a re-encode of a re-encode.

Rotation lives in metadata, not in pixels

When a phone is held upright, the sensor still records a landscape frame. Rather than rotate millions of pixels, the camera writes an integer into the EXIF Orientation tag, number 0x0112, which has eight defined values: 1 for normal, 3 for a 180-degree turn, 6 for 90 degrees clockwise, and so on through the mirrored variants. The pixels stay exactly as captured and the flag does the work at display time.

Software disagrees about how seriously to take that flag. Browsers changed their default behaviour around 2020, so a photo can look upright in a preview and sideways to a library that reads the pixel grid literally. Python's Pillow needs an explicit transpose call, and the common JPEG decoder applies no orientation at all. Hand a face detector a sideways frame and the alignment stage either fails, crops a shoulder, or rotates the image and guesses. Each of those outcomes changes the numbers that follow.

Colour arrives with a profile attached

Colour is not a property of pixels alone. It is a property of pixels plus a description of what the numbers mean, and that description is the ICC profile. The web default is sRGB, standardised as IEC 61966-2-1 in 1999. Adobe RGB, defined a year earlier, covers a wider range of greens and cyans, and Display P3, which Apple adopted for phone cameras, is wider again in reds and greens. Convert a P3 image to sRGB carelessly and saturated colours shift visibly.

A classifier trained on sRGB JPEGs learned whatever saturation those images happened to carry. Feed it a wide-gamut frame without conversion and the skin tones it receives are not the skin tones the training set contained. In a task where pigment is the strongest available cue, that is not a rounding error. The failure also runs the other way when an image is edited in a photo app, exported, and exported again, each pass nudging the same colours further from the original.

Resizing deletes texture before the model sees it

Encoder and resampler each cost resolution, and they cost different things. JPEG divides the frame into blocks of 8 by 8 pixels and transforms each one, then stores brightness at full resolution while halving the resolution of colour, a scheme known as 4:2:0 chroma subsampling. Fine colour boundaries therefore survive worse than fine brightness boundaries. Then the model resizes to its own fixed input: 112 by 112 pixels for the ArcFace family, 160 by 160 for FaceNet, 224 by 224 for the VGG-Face architecture. A 12-megapixel frame loses well over 99 percent of its pixels on the way in.

What survives downsampling is low spatial frequency: overall proportion, the ratio of face to skull, the broad tone of skin. What dies is high frequency: pore texture, hair strand detail, the fine grain of a scar. Those are precisely the features that separate neighbouring entries, which is why a pipeline tends to return broad regions rather than close distinctions. Two catalogue pairs show why that is a real loss. Fenno-Nordid of northern Europe and Equatorial Sudanid of central Africa sit at opposite ends of the pigment gradient, while African Alpinoid of the Atlas Mountains demonstrates that cranial shape varies independently of pigment. Nesiotid is a Polynesian entry noted for Europoid-like features with no close genetic relationship to Europe at all.

The number that comes back is a ranking score

Most face tools end in a softmax layer, which turns a vector of raw scores into numbers that sum to one. Those numbers are routinely read as probabilities, and they are not. Researchers led by Chuan Guo showed in a paper presented at ICML in 2017 that modern deep networks are systematically overconfident: expected calibration error for a ResNet-110 on CIFAR-100 was close to 13 percent, meaning stated confidence and observed accuracy sat far apart. Dividing raw scores by a single learned parameter, a method called temperature scaling, cut that error below 1 percent without changing a single prediction.

The same paper found that depth, width and batch normalisation all made calibration worse while making accuracy better, which is an awkward combination for anyone reading confidence off a screen. The practical consequence is narrow but useful. A tool reporting 92 percent is saying that one label ranked above the others in its training distribution. It is not saying that the label is right 92 times in a hundred, and it has no way to represent the possibility that the input belongs to nothing on its list at all.

Five checks before you upload

None of this makes a face tool useless, and a surprising share of it sits under your own control. Most of the error that reaches the model is introduced before the model runs, which means it can be reduced by preparing the file rather than by searching for a better model. The five habits below are the ones that reliably change what comes back.

  • Start from the original file, never a screenshot or a re-share, because a screenshot has already been resampled, recropped and re-encoded.
  • Check for rotation. If the preview looks upright but the only record of orientation is a JPEG flag, bake the rotation into the pixels first.
  • Flatten the colour to sRGB before uploading, so the model receives the gamut it was trained on.
  • Crop to the head and shoulders with a little margin, because a detector given a wide frame will cut the face out itself, and usually less carefully.
  • Read the output as a ranking rather than a measurement: ask which entries came closest, not how certain the tool claims to be.

Frequently Asked Questions

Why does the same photo get different answers on different tools?
Because each tool applies its own pipeline before the model runs. Detection and alignment differ, colour conversion differs, and the input size differs: about 112 by 112 pixels for one model family, 160 by 160 for another, 224 by 224 for a third. Each stage is lossy in a different way. When one file produces two labels, the disagreement usually begins in preprocessing and in the label set rather than in the part of the process people picture as the artificial intelligence.
Does uploading a photo keep its metadata?
It depends on the service, and the metadata is worth thinking about before you decide. A modern phone photo is not only pixels: the container holds camera settings, a timestamp and, unless location sharing is switched off, GPS coordinates. Services that re-encode an uploaded file typically drop that data, and some strip it deliberately for privacy. Re-encoding is not guaranteed, though, so the safe assumption is that anything attached to the file may travel with it.
What can a photo actually tell a face tool?
Broad appearance at low spatial resolution. Resizing a 12-megapixel frame down to a few thousand pixels keeps overall proportion and skin tone but discards pore texture, hair detail and fine marks, which are exactly the features that distinguish neighbouring catalogue entries. That is why results cluster around broad regions instead of close matches. A photo can show that a face sits near one documented pattern and far from another. It cannot show nationality, a legal status, or a personal identity.

Related Phenotypes

Faces from the encyclopedia that appear in this article. Open any entry for its full description, distribution, and references.

  • Fenno Nordid

    Northern Europe

    Ancient type of Northeastern Europe, sometimes placed in Nordid, sometimes East Baltid. Associated with Finno-Ugric people, once widespread ...

  • Equatorial Sudanid

    West Africa

    Central African variety with Sudanid and Congolid elements, living in the Savanna regions of East Cameroon, Central African Republic, South ...

  • African Alpinoid

    North Africa

    The most brachycephalic Maghrebi type, shows close morphological similarities to Alpines and in a lower degree to Armenids and Berberids. Pr...

  • Nesiotid

    Polynesia

    Mediterraniform Polynesid variety. Has been noted for its Europoid features, although no close genetic relationship exists. Typical in Samoa...

Test what you learned

Put your eye for regional faces to work in the daily quiz — a composite face, a world map, and your best guess.