← All articles
Applied AI

Who is the data for?

Health AI is only as representative as the data beneath it. When whole populations go uncounted, the models quietly fail the people who were never in the data — and it starts upstream, not in the algorithm.

A model cannot serve a population it was never shown. Absence in the data becomes absence in the care.

From this article
By Mike Bitok · August 17, 2026 · 4 min read
AIHealth equityData

The blind spot no one announced

A fingertip oxygen clip - the pulse oximeter - is one of the most trusted devices in medicine. It is also a quiet lesson in what goes wrong when a tool is built around one kind of body.

In a 2020 NEJM study, Black patients had nearly three times the frequency of dangerously low oxygen the device failed to detect, compared with White patients - 17.0% versus 6.2% in the larger of two cohorts. During COVID-19, that gap tracked with who got flagged for supplemental oxygen and treatment. The device wasn’t broken. It was calibrated on a population that didn’t represent everyone it would be used on - and it never announced its own blind spot.

Now imagine that same principle, but learning from data, at the scale of AI.

The pattern underneath

The uncomfortable question behind most health AI is simple: who is the data for? Whose records taught the model what “normal” looks like? Across the field, the answer is too often “the people already well served.” The same shape repeats:

The thread: this is a data problem before it is an algorithm problem. The models mostly did what they were built to do. The failure was upstream - in who was counted.

Why this matters more in LMICs, not less

Everything above was documented in well-resourced health systems with relatively rich data. Now transplant the logic to low- and middle-income settings, where high-quality datasets are scarcer still, and the risk compounds.

Pooling data and expertise across high-income and low-income countries is a reasonable response to that scarcity - but it is not a fairness guarantee. A 2024 study in Scientific Reports, using a real-world COVID-19 screening case, showed that collaborative models can produce sharply divergent performance across HIC and LMIC settings, especially where data is imbalanced. A model that looks strong on paper can still fail the patients it was exported to serve. Population diversity is the variable these systems are most sensitive to, and generalizability collapses precisely at the margins where equity is decided.

The barriers here are as much about trust and infrastructure as about code. Workshops held in the UK and Uganda in 2024, reported through the University of Cambridge’s Accelerate Programme, surfaced the familiar obstacles: complex ethical approvals, public mistrust of researchers, fears of data misuse, and limited secure infrastructure for sharing and storage. Sustainable AI starts by strengthening the data ecosystem - not by importing a finished model and hoping it travels.

The takeaway

Before adopting any AI or analytics tool, ask four questions:

  1. Who is in the training data - and who is missing?
  2. Where was it validated, and was it on people like the ones you serve?
  3. What are you not recording in your own data?
  4. Does the tool report its own uncertainty, or answer everything with false confidence?

If a vendor can’t answer the first, that is the answer.

And one to argue about over coffee: if a model works worse for an underrepresented group but better than no tool at all, is deploying it progress - or entrenching a two-tier standard of care? I don’t think that one has a clean answer. I’d like to hear yours.

The STANDING Together consensus recommendations (Lancet Digital Health, 2024) set out what transparency about representation in health datasets should look like - a useful north star if you’re evaluating a tool this year. For governance, the WHO’s 2024 guidance on AI for health pairs well with it, calling for rigorous validation, continuous monitoring, and genuinely multi-stakeholder oversight.

The promise of AI in health is real. But it is only as representative as the data beneath it - and built carelessly, it delivers its benefits to those already best served and its errors to everyone else. If you’re weighing an AI investment and want to think through representativeness first, that’s exactly what we help with.

The Enkop dispatch

Field notes and tutorials, twice a month.

Practical data & AI methods for mission-driven organizations — the same rigor we bring to engagements, written to be used.

Twice a month at most. Unsubscribe any time via the link in every email.

Not sure which engagement fits?

Book thirty minutes. You will leave with a clearer view of your options and an honest recommendation on the right next step.

Book a 30-minute call →