Based on historical distributions from ~1500 years ago. Estimates appearance, not DNA ancestry or personal identity.

AI & TechnologyOctober 9, 20268 min read

Free AI Ethnicity Guesser: The Cheap Part Is the Compute

A free AI ethnicity guesser is not a charity, and the reason it can be free says more about the arithmetic than about the answer. The arithmetic is a few billion operations; the label list was paid for decades ago.

Norid, a Central European appearance pattern recorded from the Alps to the Carpathians and named after a Roman province

Two costs, and only one is free

The word free in front of an AI tool describes one of two very different costs. Building a vision model means paying for training: compute rented by the hour for days or weeks, plus every repeat run while the settings are tuned. Emma Strubell, Ananya Ganesh and Andrew McCallum measured that bill and published it at ACL in 2019. A single transformer-big training run emitted an estimated 192 pounds of carbon dioxide equivalent. The same programme taken as a whole, including hyperparameter tuning and a full neural architecture search, reached 626,155 pounds. Their own comparison table put an average car at roughly 126,000 pounds over its entire lifetime. Training is a fixed cost, paid once, by whoever builds the thing.

Answering a question is the other kind of cost. It is marginal: incurred once per image, and paid by whoever operates the service rather than whoever built the model. That difference is why a free tier is possible at all. A fixed cost can be spread across every user who ever shows up, so a product that cost hundreds of thousands of pounds of carbon to train can be handed out at no charge indefinitely, provided the marginal cost of one more answer is small enough to vanish into the general cost of running a website. Whether that marginal cost is small is a measurable question rather than a marketing claim.

What one forward pass actually costs

Running a trained network on a single image is called a forward pass, and its size is counted in floating-point operations. Image classifiers are small by the standards of the models that make headlines, and the difference between architectures is larger than most people expect. The four figures below are the published operation counts for a single 224 by 224 colour image.

The card matters as much as the network. An NVIDIA T4, announced in 2018, is a 70-watt inference accelerator rated at 65 teraflops in half precision and 130 tera operations per second in 8-bit integer arithmetic. Divide a single ResNet-50 pass by that rating and you get around 63 microseconds of pure computation at theoretical peak. Real throughput is lower, and the reason is worth stating: accelerators of this kind are usually limited by memory bandwidth rather than by arithmetic, so the same model on the same card can run several times slower than the peak number implies. The order of magnitude survives the correction, and the order of magnitude is the point.

  • ResNet-50: 4.1 billion operations, 25.6 million parameters
  • MobileNetV3-Large: 219 million operations, 5.4 million parameters
  • EfficientNet-B0: 390 million operations, 5.3 million parameters
  • ViT-B/16: 17.6 billion operations, 86.6 million parameters

Making the bill smaller on purpose

Inference cost did not fall by accident. Two techniques did most of the work, and both were deliberate engineering rather than a happy side effect. The first is quantisation: instead of storing every weight as a 32-bit floating-point number, a serving system stores it as an 8-bit integer and uses hardware paths built for that narrower format. NVIDIA's own architecture documentation for the Turing generation lists 8-bit integer throughput at twice the half-precision rate on the same tensor cores, and the Ampere generation extends the same trick to narrower types still. The second is distillation, set out by Geoffrey Hinton, Oriol Vinyals and Jeff Dean in 2015: a small network is trained to imitate the outputs of a large one, and the student keeps most of the behaviour at a fraction of the arithmetic.

The result is that the model actually answering requests is often a few tens of megabytes rather than hundreds of billions of parameters. That is what makes a permanently free tier a stable business decision instead of a temporary promotion. It also marks the limit of the saving. The parts of a product that are cheap are the parts that were designed to be small, and a label list is not one of them.

Three published prices for one task

Published prices for recognising a face in an image sit nowhere near the compute cost, and they sit a long way from each other. Amazon charges for Rekognition at a tenth of a cent per image, which is one thousand dollars per million images. Amazon's own vision model, Nova Lite, was priced in a December 2024 comparison on the company's developer blog at about twenty-three dollars per million images at a resolution of 426 by 240 pixels. Two services, one nominal task, and a spread of more than forty times.

The distance between those list prices and the GPU arithmetic above is larger still, six orders of magnitude or more. A price is not a cost. It is a vendor's estimate of what the buyer will tolerate, set against the value of the output rather than the electricity it consumed. That matters for anyone reading a free tool at face value: the absence of a charge tells you what the seller decided to absorb, and nothing at all about whether the answer is any good.

The expensive part was never the computer

If the computation is nearly free, the money has to be somewhere else, and it is. Supervised classification needs labelled examples, and labels are made by people. ImageNet, the dataset that did more than any other to establish modern image recognition, was assembled by Jia Deng and colleagues and presented at CVPR in 2009. The full collection now holds 14,197,122 images in 21,841 synsets drawn from the WordNet noun hierarchy. The competition subset used from 2012 onward is smaller but still 1,281,167 training images spread across 1,000 classes.

Every one of those labels came from a person. Candidate images were pulled from search engines and then verified by crowd workers on Amazon Mechanical Turk, with several independent judgements combined before an annotation was accepted, and the required level of agreement adjusted class by class because some concepts are harder to confirm than others. That labour is the capital in the product. It was paid once, years ago, and it does not shrink when electricity gets cheaper. It also freezes: a label set collected in a particular year carries that year's categories forward unchanged, long after the reasons behind them have been forgotten.

Four labels and where their names come from

The catalogue on this site keeps a list assembled by nineteenth and twentieth century anatomists, and the provenance of four names shows how much is settled before any measurement happens. Norid takes its name from the Roman province of Noricum, in what is now Austria; the word was coined by Lebzelter in 1929, and Deniker in 1900 and Montandon in 1933 had called the same pattern Sub Adriatic. Transcaspian was named by Oshanin in 1927 for a position on the far side of the Caspian, and later authors proposed folding it into a different group altogether.

Samoyedic is named for a language family rather than for a face, and the entry records that the pattern reached the Yamal peninsula only in the Middle Ages, from somewhere further south. Yakonin comes from the Japanese word yakunin, meaning a government official; Weisbach defined it in 1878, and Klenke reported it in six percent of Japanese Olympic athletes in 1938. None of those four names describes a biological kind. A classifier's output layer is the same sort of object, except that an engineer chose it in a meeting and nobody wrote down the minutes.

What a free answer cannot carry

The contemporary difficulty is not that free tools are dishonest. It is that the price of inference keeps falling while everything upstream of it stays fixed. Free tiers multiply, each returning a label from a list nobody published, and the shrinking bill removes the last pressure that might have pushed a vendor to check the list carefully. That is the reverse of how a laboratory test behaves. A genotyping array consumes a chip and reagents on every run, so a seller who is wrong pays for being wrong. A forward pass consumes electricity worth a fraction of a cent, and any incentive to be right has to come from somewhere other than the cost of being wrong.

One development points the other way. Inference also runs on the device. Phone accelerators are now rated in tens of trillions of operations per second, which is enough for a small classifier, and a tool built that way has a marginal server cost of exactly zero because the photograph never leaves the phone. Whether a particular free tool does that is a factual question with a checkable answer. The reference catalogue takes a different route entirely: an entry is text, it costs nothing to serve, and it has never needed a photograph to exist.

Frequently Asked Questions

What does a free AI ethnicity guesser actually cost to run?
For the operator, very little per answer. A cloud instance carrying an inference accelerator such as an NVIDIA T4 has been listed at roughly thirty-five to eighty cents an hour, and a small image classifier can work through many images a second on that hardware. At a conservative hundred images a second, the compute behind one label comes to well under a ten-thousandth of a cent. The costs that do not shrink are the ones already paid: the annotated images used for training and the label list chosen before training began.
Why are free AI face tools so common?
Because the marginal cost of one more answer is close to zero, and because a free tier is an inexpensive way to collect images, attention, or both. Training a model is a large one-off expense, but serving it is measured in fractions of a cent per request, so the operator can spread the fixed bill across an unlimited number of users. The commercial question is what the free tier is for, since the price of the output stopped being the interesting number a while ago.
Does a free AI tool give a worse answer than a paid one?
Not necessarily, and price is a poor guide in either direction. A paid tier may return more labels or a longer explanation rather than a more accurate one, and a free tier is sometimes the same model with a shorter display. Accuracy depends on the training labels, the test set and the metric that was published, none of which is set by the price. The checkable question is whether the tool says what its label set contains and whether it can return no match at all.

Related Phenotypes

Faces from the encyclopedia that appear in this article. Open any entry for its full description, distribution, and references.

  • Norid

    Southeastern Europe

    Central European type that closely resembles Dinarids except for lighter pigmentation. Authors interpret this to be the result of Nordid and...

  • Transcaspian

    Central Asia

    West-Central Asian Orientalid with Aralid admixture, particularly common in Turkmens. Significantly differs from other Central Asians by lon...

  • Samoyedic

    Siberia

    Northwestern Sibirid contact type, influenced by strong Tungid and possibly Lappid elements. Has migrated to the Yamal peninsula and surroun...

  • Yakonin

    East Asia

    East Asian type that has been associated with the ancient Japanese government and nobility circles. The name derives from yakunin (Jap.: bur...

Test what you learned

Put your eye for regional faces to work in the daily quiz — a composite face, a world map, and your best guess.