A new generalist AI trained on raw clinical reports matches expert radiologists, proving that manual data labeling is no longer the bottleneck for medical imaging.
For years, building medical AI meant paying armies of radiologists to painstakingly label thousands of medical images. It was slow, expensive, and limited the software to narrow tasks like spotting a single type of nodule. What happens when you bypass manual labeling entirely and let the AI learn directly from raw clinical text?
The launch of RADAR, a generalist vision-language model, suggests the era of single-task medical AI is ending. By training on raw, unstructured reports, researchers bypassed the annotation bottleneck that has stalled clinical AI deployment for a decade. This shifts the debate from whether AI can generalize to how we validate systems that interpret hundreds of findings simultaneously. It challenges the assumption that medical AI must be built brick-by-brick for every individual disease.
The scale of the data
The sheer volume of the training data sets a new benchmark for abdominal imaging. Researchers trained RADAR on more than 400,000 contrast-enhanced abdominal CT examinations. The training also included 15 million anatomy-wise image-text pairs. Crucially, the model did not rely on manual annotations, learning directly from existing clinical reports to evaluate 18 anatomical structures and 146 different imaging findings.
This approach allowed the system to absorb the complex, nuanced language of real-world clinical practice. It did not just learn to spot abnormalities. It learned how radiologists describe them, bridging the gap between pixel data and medical terminology.
How the model performed
The system underwent internal and external evaluations across multiple clinical centers to test its real-world viability.
- It demonstrated robust generalization across varied clinical scenarios and imaging centers.
- In a reader study, RADAR assistance increased the diagnostic sensitivity of 26 radiologists by approximately 10%.
- The model successfully matched human experts in both general and highly complicated diagnostic tasks.
The clinical reality check
A 10% boost in sensitivity for 26 radiologists is a massive margin in clinical practice. It means fewer missed cancers, fewer overlooked internal bleeds, and faster triage in emergency departments. Because RADAR handles 146 different findings, it acts as a true generalist safety net rather than a hyper-specialized tool that only looks for one specific disease. It changes how we judge what counts as a working clinical assistant.
We must also consider the cognitive load on the clinicians. While a 10% increase in sensitivity is impressive, we do not know if this assistance came at the cost of diagnostic speed or increased false positives. If an AI flags 146 different findings, radiologists might face alert fatigue, spending more time dismissing minor suggestions than focusing on critical pathology.
While RADAR generalizes across multiple centers, the paper does not detail how it handles rare, edge-case pathologies outside its 146 classified findings. Relying on unstructured clinical reports also means the model may have absorbed the historical biases and writing idiosyncrasies of the original dictating radiologists. If the training reports contained errors, the AI may have learned those errors too.
If generalist models can consistently boost expert performance by double digits without manual training labels, the economic equation of medical software changes. The bottleneck is no longer data curation, but clinical trust.
Read the full study in Science.



