The Challenges of Real-World Validation in AI Applications for Medical Imaging: Lessons from Recent Research

One big concern from studies in the United States is that AI models often work worse when used outside the controlled places where they were trained. A review by Rajurkar and others showed that many FDA-approved AI models for medical imaging, like those that find cervical spine fractures, did not work as well on real clinical data. This difference can be risky for patients and can make doctors lose trust in AI.

The FDA approval is usually seen as proof that AI software is safe and works well. The 21st Century Cures Act defines Software as a Medical Device (SAMD) as software that works alone for medical purposes without any hardware. The approval process mostly checks how well the software works during development and early tests. But the FDA does not usually test how AI works in everyday healthcare. The models are rarely tested in future studies or at different hospitals with different patients and imaging methods.

Real-World Data and Its Role in AI Validation

A big part of the validation problem is the difference between data from clinical trials and data from real healthcare settings. Real-world data comes from everyday healthcare and shows many types of patients, variable image quality, and different clinical ways of doing things. Clinical trials usually use carefully chosen, controlled data that do not show this natural variation.

Researchers like D. Navarro-Garcia and their team report that combining AI with real-world radiology data can improve cancer diagnosis support systems. But they also mention problems like uneven image quality, limited data access, and difficult data processing. Without standard ways to collect and process data, AI models may not work well in different places. This causes mixed results when AI is used in real life.

The Importance of Standardization and Transparency

One suggested fix is to create standard rules for handling real-world imaging data. Standardization makes sure that data is processed the same way everywhere. It helps tests to be reliable and results to be compared across different sites. Without this, AI models might read clinical data differently at each place and work in different ways.

Transparency is also important for AI in medical imaging. The Coalition for Health AI (CHAI), a nonprofit group in the United States, supports tools like the Applied Model Card. This card works like a food label and gives key details about AI algorithms, such as how they were trained, possible biases, limits, and who made them. This helps healthcare managers and doctors evaluate AI tools before use and check how they perform afterward.

CHAI’s Responsible AI Guide stresses ideas like fairness, safety, trustworthiness, and responsibility. Since AI can have bias—from training data that does not represent everyone or from how the algorithm is built—ongoing checks and clear communication about these risks are needed to keep patient care fair.

Ethical Considerations Surrounding AI Bias

One challenge with medical AI is bias in machine learning models. Matthew G. Hanna and others say AI systems often face three main types of bias:

  • Data Bias: Happens when training data does not include all patient groups or clinical situations. For instance, AI trained mostly on data from one race or age group may not work well for others.
  • Development Bias: Happens from choices made when designing the algorithm, like picking features or settings that favor certain outcomes without meaning to.
  • Interaction Bias: Comes from changes in how doctors use AI or shifts in clinical practice over time.

These biases can make AI give unfair or wrong advice, which can harm patients. So, managers and IT staff should think carefully about these issues when choosing AI tools. They should also ask for audits and studies that test for biases in their own patient groups.

The Role of Post-Market Monitoring

CHAI points out that a big problem is the lack of checking AI after it is sold and used in real healthcare. Most FDA approvals do not require ongoing reviews of AI software once it is in use. Without this, AI tools that do not work well or have bias might keep being used without anyone noticing.

CHAI suggests building a nationwide quality assurance network for AI. This network would collect data on AI performance from different healthcare places and set common standards. Managers could use these results to keep AI safe and effective and to know when it needs to be updated or removed.

Checking AI continuously is important because many AI models get worse over time due to what is called “temporal bias.” Medical tech, clinical rules, and disease patterns change. Older AI models trained on old data might become less accurate or useful. Regular testing and updates can lower risks and protect patients.

AI and Workflow Automations in Medical Imaging Practices

Apart from validation issues, AI has clear uses for making medical imaging work better. AI automations can cut down on paperwork and make front-office jobs easier, which helps healthcare run smoothly.

For example, AI phone systems, like those from some U.S. companies, can answer patient calls more efficiently. These systems can set appointments, answer questions, and direct communication between patients and staff. This lets office workers focus on other tasks.

In radiology departments, AI can help process images and do early analysis. It can alert radiologists quickly if something like a tumor or fracture is found. This shortens delays and speeds up reports. Automated workflows also reduce human errors in entering data or communicating.

Using AI in both office work and imaging scans can improve overall healthcare experiences. But managers must make sure these AI tools are checked well in real clinical settings and watched regularly to make sure they work properly.

Lessons for Medical Practice Administrators, Owners, and IT Managers in the United States

People in charge of medical AI use should keep these points in mind:

  • Do Not Rely Only on FDA Approval: FDA clearance is important but doesn’t guarantee that AI will work well in your real practice. Look for extra validation from places with similar patients and imaging setups.
  • Demand Transparency: Use AI tools that give full details about their training data, limits, and known biases, like the Applied Model Card from CHAI. This helps assess risks properly.
  • Focus on Standard Workflows and Data Quality: Make sure your tech supports consistent data collection and processing. This is key for AI to work reliably.
  • Monitor AI Continuously: Set up ongoing checks after AI is in use. Use networks that share quality data when possible. Regular reviews keep patients safe and AI working well.
  • Address Ethical Concerns: Know that AI can have bias. Work with teams of clinicians, data experts, and ethicists to check fairness and how it affects care for all patients.
  • Use AI Thoughtfully in Workflows: Consider AI for both data analysis and automating office tasks. This can improve efficiency without hurting care quality.

AI in medical imaging is growing and could help improve diagnosis and patient care. But it also shows the need to be careful and check how AI performs in real life. Medical leaders in the United States must balance interest in new tech with strong evaluation to make sure AI truly helps patients and clinics.

Frequently Asked Questions

What is FDA’s definition of Software as a Medical Device (SAMD)?

SAMD is software intended for one or more medical purposes that performs these purposes independently, without being part of a hardware medical device. The 21st Century Cures Act updated its classification and regulation.

Are all clinical decision support (CDS) tools regulated by the FDA?

No, CDS tools are only regulated if they meet four criteria specified by the 21st Century Cures Act; otherwise, they are considered non-device CDS and do not require FDA approval.

What are the four criteria for non-device CDS?

1. It cannot analyze medical images or signals. 2. It must display or print medical information. 3. It must support recommendations for diagnosis or treatment. 4. It must allow independent review by healthcare professionals.

What limitations does the FDA have regarding AI validation?

The FDA’s involvement in clinical validation of SAMD has been limited, and many approved algorithms are rarely tested in real-world settings.

What did Rajurkar et al find about AI applications in medical imaging?

Their analysis indicated that AI models are often not tested outside their training environments, leading to poorer performance when validated on external data.

What shortcomings does the Coalition for Health AI (CHAI) address?

CHAI’s AI Action Plan aims to create standardized performance benchmarking and address regulatory gaps for varying risk levels of AI applications in healthcare.

What does CHAI recommend for high-risk AI applications?

CHAI suggests that high-risk applications, such as diagnostic tools, should undergo stronger oversight, while lower-risk applications should have fewer regulatory requirements.

What foundational principles does CHAI emphasize for AI?

CHAI’s principles include usefulness, fairness, safety, transparency, and privacy, which guide the development and evaluation of AI in healthcare.

What is the purpose of CHAI’s Applied Model Card?

The Applied Model Card provides detailed information about healthcare algorithms, including developer identity, bias mitigation, training data sources, and model limitations.

Why is continuous monitoring of AI models important?

Continuous monitoring ensures that AI applications remain effective and safe in clinical settings, helping mitigate risks associated with their use in patient care.