Enhancing Representation in Radiography-Reports Foundation Model: A Granular Alignment Algorithm Using Masked Contrastive Learning

Huang, Weijian; Li, Cheng; Zhou, Hong-Yu; Yang, Hao; Liu, Jiarun; Liang, Yong; Zheng, Hairong; Zhang, Shaoting; Wang, Shanshan

doi:10.1038/s41467-024-51749-0

Computer Science > Computer Vision and Pattern Recognition

arXiv:2309.05904 (cs)

[Submitted on 12 Sep 2023 (v1), last revised 3 Sep 2024 (this version, v3)]

Title:Enhancing Representation in Radiography-Reports Foundation Model: A Granular Alignment Algorithm Using Masked Contrastive Learning

Authors:Weijian Huang, Cheng Li, Hong-Yu Zhou, Hao Yang, Jiarun Liu, Yong Liang, Hairong Zheng, Shaoting Zhang, Shanshan Wang

View PDF HTML (experimental)

Abstract:Recently, multi-modal vision-language foundation models have gained significant attention in the medical field. While these models offer great opportunities, they still face crucial challenges, such as the requirement for fine-grained knowledge understanding in computer-aided diagnosis and the capability of utilizing very limited or even no task-specific labeled data in real-world clinical applications. In this study, we present MaCo, a masked contrastive chest X-ray foundation model that tackles these challenges. MaCo explores masked contrastive learning to simultaneously achieve fine-grained image understanding and zero-shot learning for a variety of medical imaging tasks. It designs a correlation weighting mechanism to adjust the correlation between masked chest X-ray image patches and their corresponding reports, thereby enhancing the model's representation learning capabilities. To evaluate the performance of MaCo, we conducted extensive experiments using 6 well-known open-source X-ray datasets. The experimental results demonstrate the superiority of MaCo over 10 state-of-the-art approaches across tasks such as classification, segmentation, detection, and phrase grounding. These findings highlight the significant potential of MaCo in advancing a wide range of medical image analysis tasks.

Subjects:	Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:2309.05904 [cs.CV]
	(or arXiv:2309.05904v3 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2309.05904
Journal reference:	Nature Communications 15, 7620 (2024)
Related DOI:	https://doi.org/10.1038/s41467-024-51749-0

Submission history

From: Weijian Huang [view email]
[v1] Tue, 12 Sep 2023 01:29:37 UTC (1,483 KB)
[v2] Mon, 18 Sep 2023 01:23:52 UTC (1,591 KB)
[v3] Tue, 3 Sep 2024 01:40:52 UTC (798 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:Enhancing Representation in Radiography-Reports Foundation Model: A Granular Alignment Algorithm Using Masked Contrastive Learning

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Enhancing Representation in Radiography-Reports Foundation Model: A Granular Alignment Algorithm Using Masked Contrastive Learning

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators