How accurate is HeadShape AI?
Across 3,349 unique photos submitted from May 2024 through July 2026, HeadShape AI's three-way read of a head — long, typical, or wide — agreed with the reference measurement 85.9% of the time, and on the photos where the reference recorded a clearly noticeable left-right difference, ours recorded one too 80.5% of the time. This window excludes an earlier batch of 910 photos, from before the reference measurement's own method had settled — we explain why below, and publish that excluded window's numbers too, along with the combined total, so you can judge the exclusion yourself. The rest of this page is the full breakdown: what we tested against, why an earlier window is excluded, the agreement figures in full, the known failure mode, and where this study's evidence ends.
What these numbers do and don't describe. Everything on this page measures the automatic measurement working on its own — one photo in, three numbers out, nothing else involved. That is deliberate: it is the part worth scrutinising, and the part we can put a number on.
It is not the whole product. The reference values we compare against were themselves reviewed by a person before they were ever delivered to a family, and every full report we deliver is checked by a person before it reaches you. So the figures below are the floor the automatic measurement reaches by itself — the free scores you see instantly are exactly that, and a full report has a person between the measurement and you.
What did we test it against?
4,264 unique images submitted for comparison — the full non-rejected history of the earlier version of this tool, from January 2024 through July 2026. HeadShape AI's current measurement produced a usable reading for 4,259 of them. Every figure on this page compares that current measurement against the earlier system's reference value for the same photo, photo by photo.
Those reference values were reviewed by a team that has spent more than ten years working with families on babies' head shape and everyday positioning, across thousands of cases — all 4,264 unique photos in this dataset passed through that review. That is real experience, though it is not the same as an assessment by independent, blinded professionals, and we repeat that distinction in the limitations below.
Why does the headline window exclude photos from before May 2024?
Because the evidence for the cut comes from the reference side, not from how the split makes our own numbers look. Across the two windows, our side is unchanged: our width-to-length reading has an identical median in both, and so does the quality of our segmentation.
| Window | Our reading — median [IQR] | Reference reading — median [IQR] | Our ellipse residual — median | Signed bias |
|---|---|---|---|---|
| Before May 7, 2024 (910) | 81.9 [77.7, 86.0] | 78.9 [75.2, 85.1] | 0.0247 | +2.58 |
| From May 7, 2024 (3,349) | 81.9 [78.2, 86.0] | 81.9 [78.3, 85.7] | 0.0247 | +0.00 |
Readings are the width-to-length ratio on the tool's internal 0–100 scale; IQR is the 25th-to-75th-percentile range.
Our own median being identical in both windows — 81.9 either way — rules out the explanation that our algorithm simply performs worse on the older photos. What changes is the reference side: its median shifts by 3.0, and the shift lands as an abrupt step around May 6–7, 2024 rather than a gradual drift — consistent with the earlier system's own measurement approach still being adjusted in that period.
Excluding those 910 photos is a judgement call, not the only defensible way to handle this data. So we publish the excluded window's numbers in full, next, and the combined total across both windows too — so you can weigh the decision yourself instead of taking our word for it.
What do the excluded photos show?
On the 910 photos from before May 7, 2024, HeadShape AI's current measurement disagreed with the reference value by an average of 3.86 on the width-to-length ratio, with a median error of 2.86. The three-way long/typical/wide read matched the reference bucket 74.5% of the time, and on photos where the reference recorded a large left-right difference, ours recorded one too 62.5% of the time. A deviation of more than 5 points occurred on 17.3% of these photos — worse across the board than the headline window.
| Metric | Primary window (3,349) | Excluded window (910) | Combined (4,259) |
|---|---|---|---|
| Width-to-length mean error | 2.46 | 3.86 | 2.76 |
| Width-to-length median error | 0.56 | 2.86 | 0.73 |
| Three-way agreement | 85.9% | 74.5% | 83.4% |
| Large left-right difference agreement | 80.5% | 62.5% | 77.0% |
| Deviation over 5 points | 12.0% | 17.3% | 13.1% |
The rest of this page reports the primary-window figures unless stated otherwise. The combined column above is what you'd get if you disagreed with excluding the earlier window — we'd rather you see it than have to ask for it.
How close are the numbers?
The width-to-length ratio row is measured on the tool's internal 0–100 scale — the same ratio described on the methodology page, just not divided by 100 here (a ratio of 0.80 there reads as 80 in this table).
The left-right difference row is already a percentage — the gap between the two diagonals as a share of the longer one, the same quantity behind the Symmetry Score on the methodology page — so a mean error of 1.30 here means 1.30 percentage points, not a ratio scaled by 100.
| Measurement | Correlation (r) | Mean error | Median error | Within ±1 | Within ±2 |
|---|---|---|---|---|---|
| Width-to-length ratio | 0.647 | 2.46 | 0.56 | 67.6% | 80.4% |
| Left-right difference | 0.480 | 1.30 | 0.36 | 79.7% | 88.1% |
The width-to-length ratio's systematic bias across this window is +0.00 — on average it runs neither high nor low. Its error still grows further out: 84.1% of photos land within ±3, 88.0% within ±5. That last, wider band matters — it's the subject of the known failure mode below.
Every figure on this page, for all three windows, is also published as a summary CSV you can check our arithmetic against. It contains aggregate rows only — no photos and nothing identifying anyone.
How often does the overall read — long, typical, or wide — agree?
Sorting each photo into one of three buckets — reads long, typical, or reads wide — HeadShape AI's bucket matched the reference bucket 85.9% of the time across the 3,349 photos in the primary window. The excluded window and the combined total are in the table above.
How closely does it agree on photos with a large left-right difference?
Of the 938 photos in the primary window where the reference measurement recorded a large left-right difference, our measurement recorded one too on 755 of them — 80.5% agreement. On 228 further photos ours recorded a large difference where the reference did not. Both figures describe how two measurements of the same photo line up, not anything about the babies in them.
Why are these numbers different from what we published before?
This page has now been revised twice, and both earlier versions had real problems — not just less flattering numbers.
The first published version used 400 photos, with a width-to-length mean error of 1.58, three-way agreement of 88.2%, and 92% agreement on photos with a large left-right difference. That sample was pulled most-recent-first — the most recently reviewed records at the time — which biased it toward recent, apparently easier photos.
The second version, the one this page replaces, used a much larger sample — reported as 6,027 photos, 6,021 of them measured — with a width-to-length mean error of 2.37, three-way agreement of 85.0%, and 79% agreement on photos with a large left-right difference. That version had two separate problems. First, it pooled two different reference conventions together, treating photos from before and after the earlier system's own measurement approach changed on 2024-05-06/07 as one consistent reference. Second, its "6,027 photos" were actually 6,027 database rows: the earlier system writes more than one row per submission, so repeated images were counted, and their agreement counted, more than once — and repeated images skew toward better-looking photos, so the double-counting flattered the result.
This version corrects both problems: 3,349 unique photos, all sharing the same post-May-2024 reference convention, each one counted once. Its numbers are a width-to-length mean error of 2.46, three-way agreement of 85.9%, and 80.5% agreement on photos with a large left-right difference.
The lesson we're taking from this: a small sample with pretty numbers is a liability, but a large sample can just as easily hide data measured under different conventions, or the same record counted twice. Any decision to split or deduplicate data has to be justified by evidence in the data itself, not by which cut produces the better-looking number — and whatever gets excluded should be published alongside what's kept, which is why the excluded window's numbers are above, not omitted.
What are the known limitations?
About 12.0% of the 3,349 photos in the primary window — 402 of them — came back with a larger deviation: more than 5 points off the reference measurement on the width-to-length ratio. Strip those out and the remaining 2,947 average an error of just 0.77. Most of the overall error above comes from this smaller group, not from broad measurement noise.
About two-thirds of that group has an explanation; the rest does not. On about 67.9% of the 402, the traced outline is a poor fit to an ellipse — consistent with it having picked up bedding, a swaddle or clothing along with the head and measured that instead. On the other 32.1%, our outline fits an ellipse cleanly, and we cannot yet say why the two measurements disagree. We would rather say that than fold all 402 into one tidy explanation.
The automatic outline check described on the methodology page is what catches the first group — it compares the traced outline against an ellipse fitted to its own shape and flags a result when the two disagree too much. Measured across the primary window, it flags 27.5% of photos and catches 69.9% of the large deviations. That leaves it well short of a safety net: about a third of the large deviations look perfectly elliptical to it, and most of what it flags turns out to be fine. Its verdict is recorded on every scan and gates nothing — and it also corrects an earlier published claim that this same check caught 100% of the large deviations in a 400-photo sample; that figure does not hold up against the fuller data.
A few boundaries worth stating plainly:
- This is a measurement-agreement study, not a clinical validation. It tells you how closely two measurement methods agree with each other, and nothing more.
- The reference values were reviewed by a team with more than ten years of experience working with families on head shape and everyday positioning — not an assessment by independent, blinded professionals.
- The sample is real uploads through a single product, not a sample deliberately balanced across ages or across how mild or pronounced each head shape was.
- We haven't yet tested how consistent repeat photos of the same baby are across different phones, distances, or lighting.
- Excluding the earlier window is a reasoned judgement call, not the only defensible approach — its numbers are published in full above, so you can weigh that decision yourself.
What can't this tool tell you?
HeadShape AI reads two proportions off a photo and turns them into scores. It does not evaluate, screen for, or make claims about any physical condition, and none of the figures on this page should be read that way — they describe how closely two measurement methods agree with each other, not what any measurement means for your baby. See what this tool can't tell you for the full picture of what HeadShape AI is, and read how the measurement works for the pipeline behind these numbers.
For informational and entertainment purposes only. Consult a professional if you have health concerns.