One big concern from studies in the United States is that AI models often work worse when used outside the controlled places where they were trained. A review by Rajurkar and others showed that many FDA-approved AI models for medical imaging, like those that find cervical spine fractures, did not work as well on real clinical data. This difference can be risky for patients and can make doctors lose trust in AI.
The FDA approval is usually seen as proof that AI software is safe and works well. The 21st Century Cures Act defines Software as a Medical Device (SAMD) as software that works alone for medical purposes without any hardware. The approval process mostly checks how well the software works during development and early tests. But the FDA does not usually test how AI works in everyday healthcare. The models are rarely tested in future studies or at different hospitals with different patients and imaging methods.
A big part of the validation problem is the difference between data from clinical trials and data from real healthcare settings. Real-world data comes from everyday healthcare and shows many types of patients, variable image quality, and different clinical ways of doing things. Clinical trials usually use carefully chosen, controlled data that do not show this natural variation.
Researchers like D. Navarro-Garcia and their team report that combining AI with real-world radiology data can improve cancer diagnosis support systems. But they also mention problems like uneven image quality, limited data access, and difficult data processing. Without standard ways to collect and process data, AI models may not work well in different places. This causes mixed results when AI is used in real life.
One suggested fix is to create standard rules for handling real-world imaging data. Standardization makes sure that data is processed the same way everywhere. It helps tests to be reliable and results to be compared across different sites. Without this, AI models might read clinical data differently at each place and work in different ways.
Transparency is also important for AI in medical imaging. The Coalition for Health AI (CHAI), a nonprofit group in the United States, supports tools like the Applied Model Card. This card works like a food label and gives key details about AI algorithms, such as how they were trained, possible biases, limits, and who made them. This helps healthcare managers and doctors evaluate AI tools before use and check how they perform afterward.
CHAI’s Responsible AI Guide stresses ideas like fairness, safety, trustworthiness, and responsibility. Since AI can have bias—from training data that does not represent everyone or from how the algorithm is built—ongoing checks and clear communication about these risks are needed to keep patient care fair.
One challenge with medical AI is bias in machine learning models. Matthew G. Hanna and others say AI systems often face three main types of bias:
These biases can make AI give unfair or wrong advice, which can harm patients. So, managers and IT staff should think carefully about these issues when choosing AI tools. They should also ask for audits and studies that test for biases in their own patient groups.
CHAI points out that a big problem is the lack of checking AI after it is sold and used in real healthcare. Most FDA approvals do not require ongoing reviews of AI software once it is in use. Without this, AI tools that do not work well or have bias might keep being used without anyone noticing.
CHAI suggests building a nationwide quality assurance network for AI. This network would collect data on AI performance from different healthcare places and set common standards. Managers could use these results to keep AI safe and effective and to know when it needs to be updated or removed.
Checking AI continuously is important because many AI models get worse over time due to what is called “temporal bias.” Medical tech, clinical rules, and disease patterns change. Older AI models trained on old data might become less accurate or useful. Regular testing and updates can lower risks and protect patients.
Apart from validation issues, AI has clear uses for making medical imaging work better. AI automations can cut down on paperwork and make front-office jobs easier, which helps healthcare run smoothly.
For example, AI phone systems, like those from some U.S. companies, can answer patient calls more efficiently. These systems can set appointments, answer questions, and direct communication between patients and staff. This lets office workers focus on other tasks.
In radiology departments, AI can help process images and do early analysis. It can alert radiologists quickly if something like a tumor or fracture is found. This shortens delays and speeds up reports. Automated workflows also reduce human errors in entering data or communicating.
Using AI in both office work and imaging scans can improve overall healthcare experiences. But managers must make sure these AI tools are checked well in real clinical settings and watched regularly to make sure they work properly.
People in charge of medical AI use should keep these points in mind:
AI in medical imaging is growing and could help improve diagnosis and patient care. But it also shows the need to be careful and check how AI performs in real life. Medical leaders in the United States must balance interest in new tech with strong evaluation to make sure AI truly helps patients and clinics.
SAMD is software intended for one or more medical purposes that performs these purposes independently, without being part of a hardware medical device. The 21st Century Cures Act updated its classification and regulation.
No, CDS tools are only regulated if they meet four criteria specified by the 21st Century Cures Act; otherwise, they are considered non-device CDS and do not require FDA approval.
1. It cannot analyze medical images or signals. 2. It must display or print medical information. 3. It must support recommendations for diagnosis or treatment. 4. It must allow independent review by healthcare professionals.
The FDA’s involvement in clinical validation of SAMD has been limited, and many approved algorithms are rarely tested in real-world settings.
Their analysis indicated that AI models are often not tested outside their training environments, leading to poorer performance when validated on external data.
CHAI’s AI Action Plan aims to create standardized performance benchmarking and address regulatory gaps for varying risk levels of AI applications in healthcare.
CHAI suggests that high-risk applications, such as diagnostic tools, should undergo stronger oversight, while lower-risk applications should have fewer regulatory requirements.
CHAI’s principles include usefulness, fairness, safety, transparency, and privacy, which guide the development and evaluation of AI in healthcare.
The Applied Model Card provides detailed information about healthcare algorithms, including developer identity, bias mitigation, training data sources, and model limitations.
Continuous monitoring ensures that AI applications remain effective and safe in clinical settings, helping mitigate risks associated with their use in patient care.