Global Effort to Overhaul Regulatory Framework for AI Medical Devices

By Deborah Borfitz 

September 8, 2026 | Data routinely captured to track health and measure wellness do not reflect ground truth, as is well known in the field of medical informatics. That imperfect healthcare data is nonetheless being used to train artificial intelligence (AI) models powering devices used in patient care, often having never been evaluated on their clinical effectiveness.    

“What happens in the end is the models are simply learning not about the disease or the person but about how we collect the data, who we collect it from, and what devices we use to collect this data,” says Leo Anthony Celi, M.D., clinical research director and senior research scientist at the MIT Laboratory for Computational Physiology and an ICU doctor at Beth Israel Deaconess Medical Center. The realization has turned him into a vocal advocate for upending the status quo and creating an entirely new system into which AI models can be safely deployed. 

The coming disruption is expected to arise from MIT Critical Data, an international consortium Celi founded in 2014 to put data and learning at the center of healthcare. It has organized more than 100 events since that time with funding bootstrapped from across multiple sources, creating an ever-growing global footprint spanning the Americas, Europe, Asia, Africa, and Australia.  

Through Medical Information Mart for Intensive Care [MIMIC], its shared open health datasets, MIT Critical Data has facilitated thousands of scientific publications worldwide. The latest, finding most AI medical devices cleared for use were not tested on patient outcomes, was published recently in PLOS Digital Health (DOI: 10.1371/journal.pdig.0001597).  

Alarmingly, the analysis revealed that only three of 1,357 AI-based medical devices authorized by the U.S. Food and Drug Administration (FDA) had been tested on whether they improved patients’ health. They were cleared for marketing under the 510(k) regulatory pathway based on their substantial equivalence to a predicate device.  

The problem with benchmarking against a legally marketed predicate device is that these, too, were approved based on predicates that are approved based on other predicates, says Rodrigo Gameiro, M.D., MPH, who is currently studying biomedical informatics at Harvard Medical School and an active participant in MIT Critical Data. “We are now investigating the pedigree of all of those devices and building a chain to see when does that end.” 

The investigation has already uncovered a recent approval that tracked back to a device “grandfathered in” when the modern FDA approval process for medical devices took effect in 1976, since the product was already being sold at that time, Gameiro says. “We don’t know the last time that someone thoroughly assessed the evidence for the effectiveness of that device.” 

Open Data-Sharing 

MIT Critical Data is a group that grew organically over time, and MIMIC was the seed, says Celi. The datasets contain deidentified electronic health record data from Beth Israel, which have been publicly available to researchers at large since 2001 with more than 100,000 pairs of viewing eyes.  

That level of open data-sharing is unprecedented in healthcare, he points out, and made plain early on that “simply providing the dataset is not enough ... [since] it is not a neutral representation of biology.” Since 2014, the consortium has been actively building and bridging scientific communities to scrutinize the data before it is applied to patient care. The effort crosses disciplines, ages, and geography.   

More than 200 projects are underway currently via these collaborators, says Celi, and the “consequences” include inquiries from high school and graduate students from around the world wanting to get involved in research. MIT Critical Data is interested in role-modeling a joy-of-learning ethos that educational systems have “successfully squashed” in favor of standardized testing, impressive CVs, and landing at elite universities and top-tier companies.   

Celi says his interest in data science grew out of frustration early in his career as a doctor with clinical guidelines based on research performed predominantly on white males in a handful of rich countries. Since health records at the time were slowly being digitized, he naïvely thought this would lead to the creation of local knowledge systems that understand the context of how health, disease, and wellness are experienced in different parts of the world. 

In his view, one of the biggest legacies of AI to date has been in “allowing us to be more critical of ourselves and of each other.” Importantly, Celi says, this includes vigilance against what medical devices in real-world use are doing—and who gets to make decisions about how such products are designed, implemented, prototyped, and evaluated.  

The other footprint left by AI is promising a future for solving some inequities, but also “showing that the emperor has no clothes,” says Celi. “In the original story ... the emperor just stiffens up and finishes the parade without clothes, as he was doing before, and if we take that pathway I think we are doomed.”   

‘Factory Reset’ Needed  

Gameiro is now leading the examination of how medical devices are being approved for deployment in the United States to understand the historical roots of why the regulatory system currently in place is not working and, ultimately, help re-architect it. More than 90% of all AI medical devices have been approved by the FDA via the 510(k) pathway, he says, allowing potential flaws to perpetuate down a chain of older technology.  

“The implications are huge because we are telling doctors that those devices work without giving them enough information on how they were tested or who they work for,” Gameiro says. Even when device performance is formally evaluated, it is typically not in a generalizable population but a specific subset of it (i.e., white males).  

It’s a problem that permeates the entire clinical trial enterprise, says Celi. Everywhere he has traveled, including the Philippines and New Zealand where he previously practiced, descriptions of stroke and heart attack have invariably been of white men experiencing the medical emergencies and the basis of guidelines issued by professional societies. That’s why the conditions are so poorly recognized in women and why they have worse outcomes when treated with drugs recommended by the American Heart Association.    

AI has made it completely obvious that existing health data do not provide reliable, accurate benchmarks, he says. Scientists worldwide are excited to be at this inflection point that could enable a “factory reset” of the medical knowledge and regulatory systems.  

“The good news is change is inevitable; we don’t even have to lift a finger for the dismantling of legacy systems,” says Celi. “AI is doing that all on its own and the onus is on us to come up with a kinder replacement system.” 

Data Quality Concerns 

Among the eye-opening studies emerging from collaborations forged through MIT Critical Data is one by a researcher in Canada finding that the birthweights of premature Black infants are more likely to be rounded off to .0 or .5 compared to white infants, a clear indicator of systemic disparities and implicit bias in medical documentation, Celi offers as an example.  

In another study he was personally involved in, it was discovered that standard protocols are not routinely followed regarding the proper positioning of patients for chest X-rays in the ICU, Celi shares. AI is therefore more likely to misclassify those images, affecting the quality of the data entered into electronic health records.  

Among patients admitted for heart failure and treated with diuretics, another study found that treatment effectiveness among people with limited English proficiency was confounded by the inconsistent times at which they were weighed day to day, he continues. This tended to happen at random times, perhaps coinciding with the availability of an interpreter or family member in instructing those patients to step on the scale.    

Since machines don’t understand context, they’re going to use features like these that even clinicians do not recognize as problematic, says Celi. “This truly raises the urgency with which we tell people they have to understand the data provenance,” he adds, getting back to one of the core missions of MIT Critical Data.  

Communities need to be purpose-built for conversations about data provenance because it isn’t something that can be learned in a classroom, he stresses. The ability to teach on the topic is “way beyond one or even a small group of professors” but, without that understanding, AI poses a real danger.   

For industry, regulatory input on this front would be economically viable, says Gameiro, notably more specific guidance on how AI medical devices should be approved. Documents issued by the FDA to date have been too vague for companies to know how to confidently proceed.  

‘Overlapping Blind Spots’ 

In a meeting convened by the U.S. Department of Health and Human Services—attended by folks from the FDA and the Office of the National Coordinator for Health Information Technology as well as industry and academia—it was “refreshing to hear everyone say they don’t know how to do this,” says Celi. “The government people are saying we’re decades behind those in research and the research people are saying we’re decades behind in evaluation science.” 

Pressed to offer concrete recommendations for the latest published paper, Celi and his team proposed a framework for addressing the recognized failures in properly testing AI medical devices. In the spirit of transparency, they suggested that every patient-facing AI evaluation gets prospectively registered with a publicly accessible protocol and analysis plan. But that FDA certification is “only 1% of what needs to happen,” he says.  

“What we should also be investing in is how to monitor this ... [and] go back to the drawing board if suddenly we’re finding that outcomes are actually not getting any better, as you’d think they would be with the adoption of AI,” says Celi. This is where society overall needs to chip in, since there will never be enough AI experts. “The field is moving so fast, 10 to 20 papers a week; no classroom [and] no conference will be able to keep up.”  

Celi says he and his colleagues told FDA officials that it is not enough to look at how AI models are being evaluated, but “the systems into which the models are being deployed in.” That requires leveraging resources, possibly students, to produce the necessary research work in lieu of relying on traditional sources of expertise as well as breaking down disciplinary silos. “We are going to have to figure this out together ... health systems cannot hire 10 data scientists and data engineers, which they probably need.” 

The FDA and its partners have developed problematic “overlapping blind spots,” says Celi, who has also been critical of academia’s myopia in carrying out its historical gatekeeping role. Those blind spots explain why pulse oximeters have been allowed to be used clinically despite reports since the 1970s that they don’t work well for people of color.  

“There must have been something wrong with the pipeline with which we establish regulation,” he says, adding that his group was roadblocked from pursuing this line of inquiry. Specifically, they had hoped to find the trail back to a 2013 FDA policy saying that having 15% of the total study pool be people of color was a good approach to diversity. For the pulse oximeter, as was later acknowledged, it was a dangerous failure.  

The agency last year updated its guidance, still in draft form, to require that 25% of the study pool be people with the darkest levels of skin pigmentation, which Celi says is not enough to rectify the underlying systemic problem. The documented trail back to the people and meetings where those original guidance decisions were made should be followed, but there has been neither will nor incentive to do so. History, therefore, is bound to repeat itself.  

“We need to sort of chip away at existing power grids,” says Celi. “It is going to be hard to remove the power structures in academia, but we can perhaps do it in small groups ... [and influence] bigger and bigger communities of people.” 

Importantly, young people need seats at the table, he says, because they’re more open-minded than their older expert counterparts. They have a “beginner’s mind,” marked by an attitude of fresh wonder rather than autopilot thoughts and a status quo bias.   

Dialoguing as Equals  

The reluctantly proposed regulatory redesign framework—a prerequisite to publication of the PLOS Digital Health paper, notes Celi—also included a proposal that journals and funders enforce adherence to reporting standards requiring disclosure of outcomes, data sources, and limitations. Much of the reporting happening currently is “performative” and not altering the culture.  

He likens the phenomenon to workplace sexual harassment training that is often treated as a legal checkbox. “The culture hasn’t changed, so ... now it’s being used as evidence that diversity is not valuable because the DEI movement failed.” 

Checklists and reporting aren’t enough because “humans are natural gamers,” says Celi. “Somehow, we need to make taking shortcuts miserable ... and we have to make the long process more enriching and pleasurable.”  

The final published recommendation was to stage evidence generation from phase 0 (pre-clearance), including retrospective validation on diverse, representative datasets; to phase 1 (peri-clearance), requiring prospective studies involving at least 500 patients; to phase 2 (post-clearance) mandating multi-center outcome trials involving at least 2,000 patients, with patient-centered endpoints and pre-specified subgroup analyses across equity-relevant strata. The idea here was not to be prescriptive, which Celi abhors, but to provoke and inspire more ideas.  

The parties who need to be engaged in the dialogue are industry, academia, and doctors as well as patients, says Gameiro, “all coming together as equals to try to come up with better metrics” for measuring AI medical device performance. When invited to the decision-making table, patients in the past tended to feel like they had to live up to the expectations of whoever was paying the check. 

MIT Critical Data has a chance of creating much-needed change, whatever that ultimately looks like, Celi says. Its partners procure all the needed funding for events and initiatives, but the consortium helps neighboring countries connect—e.g., Chile with Ecuador and Colombia, and Uganda with Ghana and South Africa. The group is getting better via word of mouth and “people are saying they are not just adding publications in their CV but actually having fun.”  

The work of MIT Critical Data, he says, can be summed up in three Rs: reflection, reimagination and reinvigoration. “Reinvigoration is the most important part ... we need to be dancing and singing while we’re learning because that’s the only strategy for this to be sustained and to be scaled.” 

That means putting aside the fear of being proven wrong, says Celi, in quoting a portion of Richard Feynman’s first principle that “the easiest person to fool is yourself.” Embracing uncertainty creates stress, but that discomfort is also what helps humans thrive.  

By learning together, everyone can become “better versions” of themselves with AI, he quips. “We see AI as our appendage, allowing us to discover new agencies ... [rather than] worrying about down-skilling or never-skilling.”    

Load more comments
comment-avatar