One photo barely beats guessing
Reading a skin tone from a single casual photo landed, on average, about two steps off on a ten-step scale. At every level of detail we tried, it beat random guessing by only a few percentage points (E1, E11, E15).
Reading colour from a photo turned out to be much harder than it looks. This page sets out every approach we tried, what each one found, and why we now believe the lighting has to be controlled.
How to read this page. Every figure comes from our experiment ledger, and each experiment on this page can be re-run from its own script. Most of the news is bad news, and we publish it on purpose. We set the pass mark before running each test and did not move it afterwards. Nothing here claims that Hrushimanju beats any other tool. Our own tests show it does not yet.
Reading a skin tone from a single casual photo landed, on average, about two steps off on a ten-step scale. At every level of detail we tried, it beat random guessing by only a few percentage points (E1, E11, E15).
The obvious fix is to correct the tint of the light. Three classic methods, and later a modern learned one, made no real difference (E2, E24).
We rebuilt the face in 3D and divided out the shading exactly. It worked perfectly, and the answer did not move (E14).
Two ideas looked promising on small samples and disappeared on bigger ones. That taught us never to report a result from one batch of people (E9, E13).
One person photographed in different lighting varies more than different people vary from each other. A second session gives a different three-way answer about 40% of the time (E16).
The human ratings we measure against are themselves imperfect, so some of the gap is not ours to close (E12, E27).
This is the finding we are surest of. We call it the mechanism, and we found it four times before we noticed it was one thing.
When the light is unknown and one-sided, the measurement is ill-conditioned. That is stronger than “noisy”. It means that changing which pixels you sample, by resizing, cropping, shifting a landmark by a few pixels, or picking a different patch of cheek, moves the answer by an amount that has nothing to do with the person. A different patch of a one-sided lit face is simply a different colour. Every step in our pipeline was individually correct, and the error appeared wherever sampling was decided.
We measured how much the reading varies at three levels, all in the same units (L*, a standard lightness scale):
Fix the camera, light, session and exposure, and the measurement holds. Change the scene, and about four and a half times more variation appears. The instrument is stable; the room is the unknown. This is why we stopped trying to out-calculate the lighting after the fact (experiments E2, E14, E24, each a different way of doing that) and decided to control it at the point of capture instead.
The honest limits. Controlled capture removes lighting variation, not camera and lens variation, and lighting is the largest term we measured. It is also “a booth, not an app”: it cannot be a casual selfie feature today. We have not claimed that every colour-analysis tool needs it. We have only measured that ours does.
Not everything failed. These are the findings we are confident in, and they are the reason we chose controlled capture over going completely free with a casual selfie.
“Chance” means what random guessing would score. A “lift” is how far above chance a result lands. “n” is the number of people. Most experiments used Meta's Casual Conversations v2 research dataset (5,566 people). Early rows used 30 to 80 people and are treated as provisional.
| # | What we tried | What we found | Verdict |
|---|---|---|---|
| E1 | The original plan: read the skin tone from one photo and match the nearest standard swatch. | About two steps off on a ten-step scale, on average. Exactly right 17% of the time. | Failed. Baseline. |
| E2 | Correct the colour cast of the light using three classic methods. | A gain of 0.01 of a step. One method made agreement worse. | Failed. The error is not mainly a colour cast. |
| E3 | Fifteen different measurements (lightness, yellowness, hue and others). | All fifteen failed in the same way, which pointed the blame away from the formulas. | Failed, and useful. Later found to have used a restricted sample (see E10, E15). |
| E4 | Re-map the scores so they match the label distribution. | Looked like an improvement, but that is what random ordering scores by construction. | Misleading. Not an improvement. |
| E8 | Reject photos whose lighting looks uneven across the cheeks. | Rejected two thirds of photos and changed the result by nothing. | Failed. Built from one photo. |
| E9 | Use the white of the eye as a lightness reference inside the photo. | Promising at 108 people. Reversed at 225. Negative at 391. | Dead at proper sample size. |
| E11 | How coarse must the answer be to work? Ten, five, three and two groups. | Barely above chance at every level, on 76 people. | Failed. This was the kill-criteria result. |
| E12 | Is the error ours or the labels'? Split results by how sure the human rater was. | Where the human rater was confident or medium-confident, the lift was +10 and +14 points. Where they were unsure, results fell below chance. | The answer key is noisy. |
| E13 | Use every frame from each person, not two. | +13 points on one batch of people, +0 on another. | Did not replicate. |
| E15 | The properly powered baseline: 391 people across all five parts of the dataset. | The signal is real (about 11% of the variation) but weak: +5 points at three groups, +0 at two. | The reference every later idea must beat. |
| E14 | Remove shading exactly, using a 3D model of the face and spherical harmonics. | The shading was removed exactly, and the result got 2 points worse. The pass mark was +2.5. | Failed. A clean null result. |
| E16 | Where does the variation live? 314 people, 1,689 readings, kept per capture. | Same person varies more than different people (11.25 vs 9.31). A second session gives a different three-way answer 39.9% of the time. The pass mark was fixed beforehand. | Noisy. Consistency is not earned. |
| E24 | A modern learned colour-correction model against no correction. | No change (0.406 to 0.404). A big-looking undertone gain was an artefact: the spread within each person got 24% worse, and 70% of people became less consistent. The measure is a ratio, and spreading people apart raises it without steadying anyone. | Null. E2's verdict stands. |
| E26 | Are the dataset's pre-picked frames representative of each video? | Mixed. But it gave the three-level split above: 2.52, 9.31, 11.25. | Split, as pre-registered. Strongest support for controlled capture. |
| E27 | How well do the dataset's two skin-tone scales agree with each other? | Agreement of 0.78 overall and 0.86 when both raters were confident. An upper bound, not true human agreement. | The labels are soft. |
Our end game is to let you design your own outfit in the software: for cosplayers who hesitate because they think they need a designer and trial fittings, for people learning fashion design, and one day a realistic version of you for a game. It is the thing we are working towards, and the experiments above tell us exactly how far away it is.
The hurdle. Matching a costume colour to a reference needs a tolerance of about 1 to 2 ΔE. Even with the camera, light, session and exposure all fixed, our readings still vary by about 2.5 L*. Across sessions it is 11.25, and across 22 seconds of casual phone photos we saw a 37-point swing. We are roughly an order of magnitude away. Published work reaches about 1.5 ΔE only with a physical colour target in frame, and about 15 without one.
The road to it is controlled capture with a colour card in the frame. Cosplayers are the one audience who would willingly hold a card up, so the end game and the method meet. It starts as a booth and grows into a kit, not a casual app. Until the capture is controlled, the product answers the easier question, whether a character's palette suits your colouring, and not whether a fabric matches.
We are looking for volunteers of every skin tone, and especially deeper skin tones where our testing is thinnest. Please do not send photos or personal details. We will not ask for a photo until the privacy policy and data-protection assessment have been reviewed by a qualified person. Email us to register your interest.