Despite showing promise in early studies, recent findings do not support the use of vision-language AI models (VLMs) for detecting lung nodules on x-rays, according to a team in Japan.
Researchers at Kobe University tested the “zero-shot” performance (no task-specific prompting or fine-tuning) of six VLMs on a dataset of 247 images, and although the models had moderate or even excellent specificity, sensitivity ranged from less than 1% to 27%, noted Mizuho Nishio, MD, PhD, and colleagues.
“The findings therefore do not support use of the tested VLMs for lung nodule detection on radiographs and highlight the need for caution in their application for chest radiography interpretation,” the group wrote. The study was published August 5 in the American Journal of Roentgenology.
Lung nodules are clinically important but commonly missed on x-rays, particularly when small or obscured by overlapping anatomy, the authors wrote. Although VLMs have shown ability to generate free-text interpretations of chest x-rays, their performance for flagging specific findings has remained poorly characterized, they added.
To further test the models, the researchers used the Japanese Society of Radiological Technology (JSRT) chest radiograph database, which includes 247 posteroanterior chest radiographs: 154 containing a single nodule (mean size, 17.3 mm; confirmed by CT scans) and 93 without. The six VLMs tested were RadVLM, GPT-4o-mini, Qwen3-VL-8B-Instruct, MedGemma-4b-it, LLaVA-Rad, and CheXpert Plus.
Five of the models received a generic prompt to generate a "concise report" with no nodule-specific instructions, while the CheXpert Plus received no text prompt at all. Two board-certified radiologists reviewed each model's output and classified it as positive or negative for nodule presence.
LLaVA-Rad led on sensitivity at 27%. CheXpert Plus detected just one nodule in 154 cases. Full results appear in the table below.
Performance of VLMs for lung nodule detection on 247 chest radiographs | |||
VLM | Sensitivity | Specificity | Accuracy |
RadVLM | 11% | 100% | 44.5% |
GPT-4o-mini | 6.5% | 93.5% | 39.3% |
Qwen3-VL-8B-Instruct | 7.8% | 88.2% | 38.1% |
MedGemma-4b-it | 5.2% | 100% | 40.9% |
LLaVA-Rad | 27.3% | 71% | 43.7% |
CheXpert Plus | 0.6% | 95.7% | 36.4% |
“This study’s findings indicate poor zero-shot performance of six VLMs for lung nodule detection on chest radiographs,” the authors wrote.
Because of incomplete public documentation of the VLMs’ original training data, it could not be determined with certainty whether the VLMs had been exposed during development to the JSRT radiographs or to other resources developed using the JSRT radiographs, the group noted.
Limitations of the study included the small sample size, use of a single dataset, lack of information regarding nodule type, incomplete public information regarding potential overlap between JSRT radiographs and VLM training data, lack of a human reader study, and lack of comparison with dedicated computer-aided detection systems, they added.
“Future studies could evaluate whether task-specific prompt engineering improves nodule detection performance,” the authors concluded.
The full study is available here, including an author video.
Whether you are a professional looking for a new job or a representative of an organization who needs workforce solutions - we are here to help.