Access to a capable AI model is no longer the scarce thing. Almost anyone can reach one now. The scarce thing is a model that was ever built to understand the population in front of it, and that gap is where the quiet damage happens.
An earlier field note argued that the value in AI lives in the layers beneath the model, and that you cannot fine-tune your way around missing data, broken integrations, or weak governance. This one follows a single consequence of a weak foundation, the one that does the most damage precisely because it is so easy to miss. When a model is trained in one context, its people, but also their disease patterns, their equipment, their clinical routines, and then pointed at another, it does not simply underperform. It fails with confidence, and it fails hardest exactly where there is the least room to absorb the mistake.
A narrow slice of the world
Start with where medical AI actually comes from, because the concentration is starker than most procurement conversations assume, and it holds no matter how close you look.
Nature Communications 2024; NEJM AI 2024; Stanford HAI 2020 / OODH 2024
The research — who produces the world's medical AI
The datasets — where the training data comes from
One country — where US clinical AI is trained
Read the figure above from the top down, and the same story repeats at three widths. Start with the research itself. The United States and China together produced almost half of the world’s AI life-science research over two decades, while Africa and Latin America combined accounted for less than five percent, even though those two regions are home to more than a quarter of the world’s people and carry more than half of the global burden of disease. The middle band moves to the data that trains the models: a 2024 systematic review of clinical text datasets found that seventy-three percent came from the Americas and Europe, regions holding just twenty-two percent of the world’s population, with more than half of the datasets in English. The bottom band narrows to a single wealthy country, and the picture does not soften. Among United States clinical AI studies where the origin of the data could be identified, roughly seven in ten drew on patients from just California, Massachusetts, or New York, and thirty-four states contributed no data at all.
The obvious objection is that this must be old news, that a field moving this fast has surely diversified by now. It has not. The three-state concentration first documented in 2020 was still being described as a well-known systemic bias in a 2024 review. What has moved is the conversation about fixes, not the underlying geography. The medicine we are teaching machines is still learned from a handful of places, and that narrowness is not a rounding error. As one Stanford researcher put it, geography correlates with a great many things that matter to health, from diet and environment to the equipment in the room when a scan is taken. A model that has only ever seen three states has, in a real sense, only learned three states.
Representation is not a dataset problem
The reflex, once the numbers land, is to call for more data. Collect more, from more places, and the gap closes. It is an intuitive fix and a mostly wrong one, because it misdiagnoses where representation actually lives.
Representation is a property of the entire lifecycle, not a box ticked at the dataset stage. This is the reframing at the center of a 2026 opinion in Frontiers in Artificial Intelligence, and it is the single most useful idea for anyone about to sign a contract. Data representation matters, but so does contextual representation, whether the social and environmental conditions of the place a model will run were considered while it was built. So does epistemological representation, whether local clinical knowledge and the judgment of the people who will use the system were treated as legitimate inputs. And so does infrastructural representation, whether the monitoring, appeal, and governance needed to catch a failure even exist where the model lands. A model that looks demographically balanced in development but arrives somewhere with no way to validate it, monitor it, or contest its output has not achieved representation. It has only extended the reach of a system nobody on the ground can question.
This is why compute pointed at unrepresentative data can be worse than no model at all, a point the earlier field note made in passing and this evidence makes concrete. A model that has learned the wrong signal does not hedge. It produces an answer that is fluent, confident, and wrong, and it produces it at scale. The failure modes are well documented. Models appear accurate in validation and then collapse in a new hospital because they learned the quirks of one site’s equipment rather than the disease. Fairness that holds where a model was built does not travel with it across a border. When a United Kingdom model was moved to hospitals in Vietnam, its performance did not carry over cleanly, because local disease prevalence, infrastructure, and clinical workflow differed from the place it was trained. When a dermatology model validated at its source was tested against a teledermatology clinic in Colombia, subgroup performance gaps that had been small at home grew substantially larger. The model announced none of this. Someone had to go looking.
The gap is not only a choice
It would be neat to say the concentration is simply a matter of where people chose to collect data, but part of it is written into law, and this cuts in a direction worth being honest about. The instruments that protect people, privacy rules, copyright reservations, data-localization and data-sovereignty regimes, also govern what data can be pooled to train a model, and different jurisdictions have drawn those lines very differently. The European Union restricts training on personal and copyrighted data far more tightly than the United States does, through the GDPR, a copyright opt-out for text and data mining, and the AI Act’s ban on untargeted facial scraping. Analysts at the Information Technology and Innovation Foundation, who favor a more permissive approach, tie this to a widening “model gap,” with United States organizations producing forty notable foundation models in 2025 against the European Union’s three.
For the settings this piece is about, the tension runs both ways, and neither end serves representation. Where data protection is weak, populations are exposed to extraction they cannot control, which is its own harm. Where the rules are strong but blunt, data can be walled off so completely that the population never enters the training set at all. The mistake is to treat this as a simple case for looser rules. Data sovereignty, an institution keeping control of its own patients’ records, is not the obstacle to representation. It is part of what representation means. The task is to make protection and participation compatible rather than trade one for the other, which is exactly what the better technical fixes are built to do.
One clarification matters here, because it is easy to blur. These legal constraints explain a good deal about why a wealthy bloc like Europe underproduces relative to the United States. They do not explain why Africa and Latin America are nearly absent, which owes far more to thin digitized records, limited compute, and scarce research funding than to any law forbidding participation. Both stories are true. Neither is closed by simply collecting more.
The failure is not “over there”
It is tempting to file all of this under Global South, a problem of poorer places catching up. That framing is comfortable and it is wrong, and the cleanest way to see why is to look inside the richest country in the world.
Rural America is underrepresented in health AI the same way a low-resource clinic is: thin in the data, thin in infrastructure, and thin on the institutional capacity to notice when a model is failing. A 2026 analysis from the Federation of American Scientists describes exactly this “rural invisibility,” and names the structural detail that makes it dangerous. Rural systems tend to run with less margin for error and fewer backup resources when something goes wrong, so a confident wrong answer has less to absorb it. The same analysis notes that the major AI governance frameworks, strong as they are on principles, do not require testing on small or geographically distinct populations. A model can pass every general fairness check and still have never been examined on the people who will actually use it.
Put the rural-Kansas picture beside the clinic-in-an-LMIC picture and the real thesis comes into focus. This was never a problem of one region lagging another. It is a problem of any population that was absent when the model was built and is powerless to correct it afterward, whether that population sits in the Rift Valley or the Great Plains. What “Global South” really names here, as the Frontiers authors argue, is not a place on a map but a position: being on the receiving end of a system you did not shape and cannot contest.
That reframing changes who the argument is for. This is not a request to be charitable to distant others. It is a warning to anyone deploying a model into any setting it was not built to understand, which, given where the data actually comes from, is nearly everyone.
The compounding trap
The deepest harm comes from two failures stacking on top of each other. A population is underrepresented in the training data, and that same population lacks the compute, the digitized records, and the regulatory capacity to build an alternative or to detect that the imported model is failing. Underrepresentation is the first cut. The absence of a correction loop is what lets the wound go unnoticed.
This is the part the excitement at the top of the stack tends to skip. A model can be imported far faster than the capacity to monitor it can be built. When that happens, the populations least represented in the original data become the least protected from silent failure, because there is no drift monitoring to flag the degradation, no subgroup evaluation at the threshold clinicians actually use, and no appeal through which a patient or a clinician can push back and, in pushing back, expose the error. The failure stays invisible precisely where it matters most. That is not imperfect generalization. It is inequity that compounds in silence, not by anyone’s intent, but through everything left out.
What actually holds
None of this is an argument for waiting until the data is perfect, which would freeze every useful program forever. It is an argument for building the layer that makes a model correctable, and for being honest that this layer, not the model, is the hard part.
The corrective has a shape, and the recent literature is converging on it. Data documentation should reach past demographic counts to record how and where data was captured and, explicitly, what population it represents. Evaluation should disaggregate performance by subgroup at the operating threshold actually used in practice, not report a reassuring aggregate. Deployment should be conditioned on monitoring that detects drift over time, with defined thresholds for review and withdrawal, and on appeal pathways that let affected people contest a decision. Governance should establish who owns the model and the data it generates locally, and who has the authority to inspect, retrain, or switch it off. This is context proofing across the lifecycle, and it is the same discipline whether the context is rural, national, or foreign.
Two developments are worth naming because they show the fix is practical, not aspirational. The first is technical, and it resolves the sovereignty tension directly rather than wishing it away. Federated learning lets institutions, including those in lower-resource settings, contribute to a shared model on data that never leaves home. It improves representativeness without forcing anyone to surrender their patients’ records, which is the whole point: representation without giving up control. The second is organizational, and the evidence for it is unusually direct. The same Nature Communications atlas that maps the concentration also found that international collaboration produces more relevant, higher-impact research, and it singled out Kenya, which collaborates internationally on two-thirds of its highest-ranked work. Co-development is not a moral concession to the places a model will serve. It is the pattern the data says produces better AI.
Five questions to ask before you deploy
For anyone approving an AI tool for a setting it was not born in, the argument reduces to a short interrogation. Not “how accurate is it,” but “accurate for whom, and what happens here when it is wrong.”
- Was this validated on a population like the one it will serve? Not demographically balanced in the abstract, but tested on the disease prevalence, equipment, documentation practices, and clinical workflow of the actual deployment setting.
- How will we know when it is wrong for our people? What is monitored, at what threshold, and who reviews it? A model with no drift monitoring in a low-margin setting is an unmonitored risk, not a tool.
- Can someone contest a bad output, and will that feedback reach anyone? Appeal is not a courtesy. It is often the only channel through which a localized failure becomes visible before it scales.
- Who owns the model, and the data it generates here? Can the deploying institution inspect, retrain, or withdraw it, or does it depend on a vendor who bears none of the consequences of being wrong?
- Were the people who will use this involved in building or adapting it? If the answer is no, ask what conditions of their setting the model has therefore never seen.
A vendor who can answer these calmly is worth more than one with a more impressive benchmark. A model earns the right to deploy through accountability to the place it will serve, not through the demographics of the data it happened to be trained on.
Build accordingly
The question was never whether your organization can get its hands on a powerful model. Increasingly, everyone can. The question is whether that model was ever built to understand the people you are about to point it at, and whether the system around it can catch the moment it is confidently wrong about them.
You can import the model. You cannot import the context it was never trained on, the monitoring it was never given, or the accountability to a population it never saw. Those you have to build where the model will actually work. Build accordingly.
This is a field note from our practice. If you’re weighing an AI tool for a setting it wasn’t built for, and want an honest read on whether it will hold there, start a conversation.
Sources
-
Schmallenbach L, Bärnighausen TW, Lerchenmueller MJ. “The global geography of artificial intelligence in life science research.” Nature Communications, vol. 15, 12 September 2024. — The concentration of the research enterprise, the under-5% share for Africa and Latin America against more than half the global disease burden, the international-collaboration premium, and the Kenya figure.
-
Wu J, et al. “Clinical text datasets for medical artificial intelligence and large language models: a systematic review.” NEJM AI, 2024. — The 73% / 22%-of-population dataset figure (as cited in Hilling et al., below).
-
Hilling DE, et al. “The imperative of diversity and equity for the adoption of responsible AI in healthcare.” Frontiers in Artificial Intelligence, vol. 8, 16 April 2025. — The lifecycle framing, federated learning, and the STANDING Together reference.
-
Kaushal A, Altman R, Langlotz C, as reported by Stanford HAI. “The Geographic Bias in Medical AI Tools.” 21 September 2020. — The three-state concentration in US clinical AI, and the geography-correlates-with-health point.
-
Open Data Health systematic review, 2024 (doi:10.1093/oodh/oqaf027). — Documents that the three-state concentration persists.
-
Tinco Aliaga CA, Tan MJT, Hinostroza Fuentes VG, Abdul Karim H, AlDahoul N. “AI without representation is just inequity at scale.” Frontiers in Artificial Intelligence, vol. 9, 29 July 2026. — The “representation as infrastructure” reframing and the fluent-but-wrong failure cascade.
-
Yang J, et al. Generalizability assessment of AI models across hospitals in a low-middle and high income country. Nature Communications, 2024 (doi:10.1038/s41467-024-52618-6). — The UK-to-Vietnam transfer study.
-
Schrouff J, et al. Diagnosing failures of fairness transfer across distribution shift in real-world medical settings (doi:10.52202/068431-1403). — The Colombia teledermatology validation.
-
Qi Z, Lin T, Olagoke A. “Making Rural Communities Visible in Artificial Intelligence.” Federation of American Scientists, 8 June 2026. — “Rural invisibility,” “less margin for error,” and the governance-framework gaps.
-
PLOS Digital Health analysis, reported by STAT. “More than half of data used in health care AI comes from the U.S. and China.” 6 April 2022. — The earlier US-and-China dataset finding across more than 7,000 clinical AI papers.
-
Castro D. “How Rules for Publicly Available Data Are Shaping the Future of AI.” Information Technology and Innovation Foundation, 13 March 2026. — The “model gap” figure and the US-versus-EU divergence (a think tank that argues for a more permissive approach).
-
Greenberg Traurig. “EU AI Act’s Opt-Out Trend May Limit Data Use for Training AI Models.” 3 July 2024. — The EU AI Act’s text-and-data-mining opt-out.