Based on historical distributions from ~1500 years ago. Estimates appearance, not DNA ancestry or personal identity.

GamesOctober 5, 20268 min read

Guess My Ethnicity: What One Guess Can Carry

When someone types guess my ethnicity, they are asking a stranger to place a bet, and the odds depend far more on the list of labels than on the face. It asks what the answer key is and what one answer can carry.

Gorid, a central European mountain pattern recorded as intermediate between West Alpinid and East Europid

The request is aimed at the wrong person

Guess my ethnicity is an imperative pointed at a stranger, and that is already an unusual arrangement for a question. In an ordinary examination the person who holds the answer writes the key and the person being tested supplies the working. Here the asker writes the key and asks somebody else to guess it. The exchange is a hybrid of test and performance, in which one person guesses and the other grades, and the grader's authority never has to be justified or even stated.

The key deserves scrutiny before the guessing starts, because it is not a measurement. In most versions of the request the answer is a label the asker selected, drawn from a list that came from somewhere else: a census form, a family story, a DNA test's reference panel, a school register. None of those is a property of the face in the photograph. The key is a decision, and a guess can be marked wrong by a decision that no amount of looking at the image would overturn.

A guess is a bet on a prior

Every guess opens with a prior, which means how likely each label is before the face is examined at all. Amos Tversky and Daniel Kahneman built the standard demonstration of this in the 1970s around a hit-and-run accident. A city runs 85 per cent green cabs and 15 per cent blue ones. A witness who is correct 80 per cent of the time says the cab was blue. Across hundreds of subjects the median and modal answer was 0.8, which is simply the witness's stated reliability; the correct figure is 12/29, about 0.41, because a city full of green cabs manufactures a great many mistaken blue sightings.

The same arithmetic runs underneath a face guess, and it explains a result that startles people who take part. When a label list holds one category covering a large share of the faces a guesser has met, that category gets returned far more often than it deserves, not because the face resembles it but because the prior is heavy. Naming the prior does not remove it. What naming it does is make a disagreement legible: two guessers can study one photograph while starting from very different weights.

How many options one answer can hold

Information theory sets a ceiling that has nothing to do with skill. Claude Shannon's 1948 paper on communication formalised the bit as a unit of choice, and for a menu of N equally likely options a single answer can carry at most log2(N) bits. Four options are worth two bits. Six buttons, the shape of most consumer interfaces, are worth about 2.6. The catalogue behind this site holds 209 documented patterns, which is close to 7.7 bits. The ceiling is fixed by the interface, and it applies before any question of accuracy is raised.

This is why a tool answering in continents and a tool answering in named patterns are not doing the same job, even when they agree. The first has already discarded most of the available answers by printing six buttons. A reader who wants to know what an answer was worth has to ask how large the answer space was, and that number is rarely printed beside the confidence percentage.

When many guesses beat a single one

Combining guesses is the usual repair, and it rests on an argument from 1785. Condorcet's essay on applying probability to decisions taken by plurality held that if each voter is more likely right than wrong, a majority grows more reliable as the group grows. Francis Galton supplied the empirical version in Nature in 1907. At a Plymouth livestock show, 787 legible sixpenny tickets had been sold to people estimating the dressed weight of a fat ox. The middlemost estimate came to 1,207 pounds against an actual 1,198.

What gets left out of that story is the scatter. A single ticket carried a probable error of 37 pounds, roughly 3.1 per cent, and a full quarter of the guesses sat more than 45 pounds above the true weight. The crowd was accurate because those errors pointed in both directions and cancelled. Galton's competitors also enjoyed one large structural advantage that is easy to overlook: their target was a weight. It could not be redefined after the answers had been collected.

Why a face breaks both conditions

Neither condition survives the move to faces. The target drifts, because the answer is a label the asker holds rather than a quantity sitting on an animal, so there is nothing for the errors to scatter around. And the errors stop being independent. Guessers work from shared prototypes and a shared short list of names, so their mistakes line up. Averaging a hundred correlated errors converges on the shared bias rather than on anything true. David Green and John Swets set out the formal vocabulary for this in 1966 with signal detection theory, which separates sensitivity, how far apart two patterns sit for a given observer, from criterion, how willing that observer is to say one name rather than another. Averaging cancels noise. It does not move a criterion everybody shares, and it cannot manufacture sensitivity that nobody has.

The catalogue shows what sensitivity costs in practice. Two entries can sit close enough that telling them apart is the entire task. Gorid is recorded in central Europe and described in its own record as metrically intermediate between West Alpinid and East Europid. Indo Brachid is described as a contact type produced by old Pashto, Saka and Scythian migrations and reported from southwestern India to Afghanistan. Neither is a kind waiting to be recognised. Each is a description of where a distribution sat at a period, and a guess that separates the two is doing something a guess that returns a continent is not.

What makes one guess better than another

The informative question is not whether a guess was right but how finely it discriminates. Landing correctly on the largest bucket costs nothing and proves nothing, because the prior made that outcome likely before the face arrived. Separating two neighbouring documented patterns is a different order of claim. Ordinary scoring hides the gap, since it counts both as one correct answer, which is why a quiz scored by distance behaves differently from one scored right or wrong.

Two properties of the catalogue make the difference concrete. A pattern can occupy a small range and then leave the record altogether. Tasmanid is described here as an insular, slightly gracilised Australid subtype of Tasmania, isolated after the island separated from the mainland, and the catalogue records that the last individuals in whom it was documented died in 1869 and 1876. Sandawe is recorded as a relict of Khoisanid populations that were more widely spread before the Bantu expansion. An answer naming either is a far narrower claim than one naming a hemisphere, and it is a claim about a date as much as a place.

The answer key has been reissued

The lists keep being rewritten and the people keep moving, which changes the prior any guess is placed against. Administrations redraw their categories, so one person can be sorted differently by two consecutive censuses without anything about them changing. Populations relocate on a scale that makes a fixed map a weak predictor, and most of that movement happens inside countries rather than between them. China's 2020 census counted 492.76 million people living outside the township where their household was registered, an 88.52 per cent rise in ten years, of whom 124.84 million had crossed a provincial boundary.

Labour migration adds a second layer. The International Labour Organization estimated 169 million international migrant workers in 2019, 4.9 per cent of the global labour force and as much as 41.4 per cent of it in the Arab States. Where a guesser's prototypes were built is therefore a poor guide to who is likely to be standing in front of them, because the mixture has changed faster than the prototypes have. A heavy prior is not wrong about the past. It is dated.

What survives the arithmetic is a narrower claim, and it is the one a catalogue entry can support: this face resembles patterns the literature recorded in this region at a baseline of roughly 1500 years ago, and here are the sources with their dates attached. That claim can be checked, argued with and revised. A guess that a stranger must grade, and only the asker can overturn, cannot.

Frequently Asked Questions

How accurate is a guess about my ethnicity?
It depends on the size of the menu before it depends on anything else. A guess returning one of four continents has about two bits of room, and a guess returning one of 209 documented patterns has close to eight, so the two answers are not comparable even when they agree. Accuracy also has no fixed target here, because the answer is a label the asker chose rather than a quantity that can be weighed. The version of the question with an answer asks which documented pattern a face most resembles.
Why do crowds guess faces less well than they guess weights?
Because the crowd trick needs two things faces do not supply. Galton's 787 guesses about a fat ox converged because the errors pointed independently in both directions and the true weight was a single number nobody could renegotiate after the fact. Face guessers share prototypes and a short list of names, so their errors correlate, and averaging preserves the shared bias rather than cancelling it. The target is also a chosen label, which lets a crowd be consistent while answering a question that was never fixed.
What does a catalogue entry give me that a score does not?
A region, a description and at least one source, with a stated baseline of roughly 1500 years ago. Sandawe, for instance, is recorded as a relict pattern of the Tanzanian savannahs tied to Khoisanid populations that were more widely spread before the Bantu expansion. A score compresses that into a single label and hides the date behind it. The entry cannot tell you which pattern you are, and it never claims to, but it can be checked, argued with and corrected.

Related Phenotypes

Faces from the encyclopedia that appear in this article. Open any entry for its full description, distribution, and references.

  • Gorid

    Central Europe

    East Alpinid variety, named after the Polish word for mountain (Gora). Metrically intermediate between West Alpinid and East Europid. The nu...

  • Indo Brachid

    South Asia

    Indo Brachid proper, relatively recent Indid-Turano-Armenid contact type, probably the result of old Pashto, Saka, and Scythian migrations. ...

  • Sandawe

    East Africa

    Distinctive East African type of the Tanzanian savannahs, especially the Dodoma region. A relict of Khoisanid populations that were more wid...

  • Tasmanid

    Australia

    Insular, slightly gracilised Australid subtype of Tasmania. Similar to Barrinean, sometimes also associated with Melanesid or Negritid. Most...

Test what you learned

Put your eye for regional faces to work in the daily quiz — a composite face, a world map, and your best guess.