The blind spot no one announced
A fingertip oxygen clip - the pulse oximeter - is one of the most trusted devices in medicine. It is also a quiet lesson in what goes wrong when a tool is built around one kind of body.
In a 2020 NEJM study, Black patients had nearly three times the frequency of dangerously low oxygen the device failed to detect, compared with White patients - 17.0% versus 6.2% in the larger of two cohorts. During COVID-19, that gap tracked with who got flagged for supplemental oxygen and treatment. The device wasn’t broken. It was calibrated on a population that didn’t represent everyone it would be used on - and it never announced its own blind spot.
Now imagine that same principle, but learning from data, at the scale of AI.
The pattern underneath
The uncomfortable question behind most health AI is simple: who is the data for? Whose records taught the model what “normal” looks like? Across the field, the answer is too often “the people already well served.” The same shape repeats:
-
Skin cancer AI. A systematic review in Lancet Digital Health examined 21 open-access skin-image datasets totalling 106,950 images. Of the images with a recorded skin type, only ten were from brown skin (Fitzpatrick V) and just one from dark brown or black skin (Fitzpatrick VI). In the subsets with ethnicity data, none were from people of African, Afro-Caribbean, or South Asian background.
-
Chest X-rays. In Nature Medicine, AI classifiers trained on large public datasets consistently and selectively underdiagnosed female, Black, and low-income patients - labelling sick people as healthy. The failure was worst for intersectional groups, such as Hispanic female patients. A silent error: no alarm, just care that never gets triggered.
-
Genomics. Martin and colleagues in Nature Genetics found that roughly 79% of participants in genome-wide association studies were of European descent - a group that is about 16% of the world’s population - and that the non-European share had stagnated or declined since around 2014. The consequence: polygenic risk scores are several times more accurate for people of European ancestry than for everyone else. Deployed as-is, they would systematically hand their benefits to the already-advantaged.
-
The column nobody records. In many programs, donor and ministry reporting captures age and sex but omits disability entirely. A datapoint that is never collected can never reveal a disparity - and an algorithm built on that data can’t account for what it was never shown.
The thread: this is a data problem before it is an algorithm problem. The models mostly did what they were built to do. The failure was upstream - in who was counted.
Why this matters more in LMICs, not less
Everything above was documented in well-resourced health systems with relatively rich data. Now transplant the logic to low- and middle-income settings, where high-quality datasets are scarcer still, and the risk compounds.
Pooling data and expertise across high-income and low-income countries is a reasonable response to that scarcity - but it is not a fairness guarantee. A 2024 study in Scientific Reports, using a real-world COVID-19 screening case, showed that collaborative models can produce sharply divergent performance across HIC and LMIC settings, especially where data is imbalanced. A model that looks strong on paper can still fail the patients it was exported to serve. Population diversity is the variable these systems are most sensitive to, and generalizability collapses precisely at the margins where equity is decided.
The barriers here are as much about trust and infrastructure as about code. Workshops held in the UK and Uganda in 2024, reported through the University of Cambridge’s Accelerate Programme, surfaced the familiar obstacles: complex ethical approvals, public mistrust of researchers, fears of data misuse, and limited secure infrastructure for sharing and storage. Sustainable AI starts by strengthening the data ecosystem - not by importing a finished model and hoping it travels.
The takeaway
Before adopting any AI or analytics tool, ask four questions:
- Who is in the training data - and who is missing?
- Where was it validated, and was it on people like the ones you serve?
- What are you not recording in your own data?
- Does the tool report its own uncertainty, or answer everything with false confidence?
If a vendor can’t answer the first, that is the answer.
And one to argue about over coffee: if a model works worse for an underrepresented group but better than no tool at all, is deploying it progress - or entrenching a two-tier standard of care? I don’t think that one has a clean answer. I’d like to hear yours.
One link
The STANDING Together consensus recommendations (Lancet Digital Health, 2024) set out what transparency about representation in health datasets should look like - a useful north star if you’re evaluating a tool this year. For governance, the WHO’s 2024 guidance on AI for health pairs well with it, calling for rigorous validation, continuous monitoring, and genuinely multi-stakeholder oversight.
The promise of AI in health is real. But it is only as representative as the data beneath it - and built carelessly, it delivers its benefits to those already best served and its errors to everyone else. If you’re weighing an AI investment and want to think through representativeness first, that’s exactly what we help with.