Enhancing Chest X-ray Diagnosis with a Multimodal Deep Learning Network by Integrating Clinical History to Refine Attention

通过整合临床病史,利用多模态深度学习网络增强胸部X光诊断,以提高注意力

阅读:1

Abstract

The rapid advancements of deep learning technology have revolutionized medical imaging diagnosis. However, training these models is often challenged by label imbalance and the scarcity of certain diseases. Most models fail to recognize multiple coexisting diseases, which are common in real-world clinical scenarios. Moreover, most radiological models rely solely on image data, which contrasts with radiologists' comprehensive approach, incorporating both images and other clinical information such as clinical history and laboratory results. In this study, we introduce a Multimodal Chest X-ray Network (MCX-Net) that integrates chest X-ray images and clinical history texts for multi-label disease diagnosis. This integration is achieved by combining a pretrained text encoder, a pretrained image encoder, and a pretrained image-text cross-modal encoder, fine-tuned on the public MIMIC-CXR-JPG dataset, to diagnose 13 diverse lung diseases on chest X-rays. As a result, MCX-Net achieved the highest macro AUROC of 0.816 on the test set, significantly outperforming unimodal baselines such as ViT-base and ResNet152, which scored 0.747 and 0.749, respectively (p < 0.001). This multimodal approach represents a substantial advancement over existing image-based deep-learning diagnostic systems for chest X-rays.

特别声明

1、本页面内容包含部分的内容是基于公开信息的合理引用;引用内容仅为补充信息,不代表本站立场。

2、若认为本页面引用内容涉及侵权,请及时与本站联系,我们将第一时间处理。

3、其他媒体/个人如需使用本页面原创内容,需注明“来源:[生知库]”并获得授权;使用引用内容的,需自行联系原作者获得许可。

4、投稿及合作请联系:info@biocloudy.com。