Aainahh Add to Shopify
← Back to blog
How it works

What "AR shade matching" actually means (and where it breaks)

"AR shade matching" gets used as a single phrase, but it's really three separate technical problems stacked on top of each other, and most of the disappointing try-on experiences you've used trace back to one specific layer failing quietly while the other two work fine. If you're evaluating a try-on app for your store — or trying to understand why a competitor's demo looks better than another's — it helps to know what's actually happening under the hood.

The three layers, in order

01

Skin sampling

The system has to isolate actual skin pixels from a live camera feed — separating cheek and jawline skin from hair, shadow, makeup already being worn, background, and the lips or eyes themselves — and sample color from multiple points, not one. A single sample point is fooled easily by a stray shadow or a red undertone from ambient lighting.

02

Lighting correction

Raw camera color is not the same as true skin color — a phone camera under warm indoor lighting will read someone's skin tone several shades warmer than daylight would. Before any shade can be recommended, the sampled color has to be corrected against a white-balance reference, usually by identifying a known-neutral point in frame (sclera of the eye, teeth, or a calibration step) and adjusting the whole sample against it.

03

Undertone + depth scoring against the catalog

Once you have a corrected skin tone, it gets scored on two independent axes — depth (how light or dark) and undertone (warm, cool, neutral, olive) — and matched against your actual product catalog's shade data. This is where the app's judgment matters most: a system that only sorts by depth will recommend a shade that's the right lightness but the wrong undertone, which is exactly the kind of "close but wrong" match that still gets returned.

Where each layer actually breaks in the wild

Skin sampling breaks on facial hair, makeup already applied, and low-end cameras

Multi-point sampling across the cheek, jaw, and forehead is table stakes at this point — anything sampling from a single spot is going to get fooled regularly. Where it still goes wrong: a customer testing a lip shade while already wearing foundation two shades off from their bare skin, or trying on a shade in poor light on an older device with a lower-quality sensor. Good implementations sample continuously and re-average rather than locking onto a single frame.

Lighting correction breaks under mixed or colored lighting

Daylight from a window on one side of someone's face and a warm ceiling light on the other is a common real-world condition that trips up naive white-balance correction — the two halves of the face can read as different undertones. This is genuinely one of the harder unsolved edges of the category; the honest answer is that no camera-based system fully solves mixed lighting today, and the best you can do is correct aggressively and communicate confidence, not pretend precision you don't have.

Undertone scoring breaks when it's trained on a narrow shade range

This is the failure mode that matters most and gets talked about least. If the underlying color model and reference dataset were built and tested mostly on lighter skin tones, the undertone math simply degrades at the deeper end of the range — not because deeper skin is "harder" in some fundamental sense, but because the system was never properly calibrated against it. We cover this specific failure in more depth in does virtual try-on actually work for every skin tone, because it deserves its own answer rather than a footnote here.

The tell that a shade-matching system is doing real work versus performing a filter effect: ask it to explain why it recommended a shade — undertone and depth reasoning, not just a swatch. If a vendor can't show you that reasoning, they're probably not doing the scoring in step three at all, just overlaying a fixed set of options.

What "real-time" is actually buying you

The other claim worth unpacking is "real-time." The value isn't speed for its own sake — it's that sampling continuously, frame over frame, lets the system average out noise from a single bad frame (a blink, a shadow, a hand moving past the camera) instead of committing to one snapshot. A system that samples once and locks is more fragile than one that keeps re-measuring as the shopper moves. If a demo looks static and doesn't visibly adjust as the person turns their head or the lighting shifts slightly, that's usually a single-snapshot system, not a continuous one.

What this means if you're evaluating a try-on app for your store

Ask three questions before you install anything:

The bottom line

"AR shade matching" is doing three distinct jobs — sampling, correcting, and scoring — and a system is only as good as its weakest layer. When you're comparing options, don't evaluate the demo video; evaluate whether it holds up under bad lighting, movement, and a range of real skin tones, because that's the actual product your customers will be using on their own phones, in their own bathrooms, in whatever light they happen to have.

Related reading

See the full sampling-to-recommendation pipeline live

Aainahh's engine re-samples continuously and scores undertone and depth independently — try it with your own camera in under a minute.

Try with your camera