计算机科学
嵌入
编码(内存)
管道(软件)
水准点(测量)
人工智能
编码器
自然语言处理
可视化
程序设计语言
大地测量学
操作系统
地理
作者
Yan Yang,Jun Yu,Jian Zhang,Weidong Han,Hanliang Jiang,Qingming Huang
标识
DOI:10.1109/tmm.2021.3122542
摘要
Medical image report generation (MeIRG) aims at generating associated diagnosis descriptions with natural language sentences from medical images, which is essential in the computer-aided diagnosis system. Nevertheless, this task remains challenging in that medical images and linguistic expressions should be understood jointly which however show great discrepancies in the modality. To fill this visual-to-semantic gap, we propose a novel framework that follows the encoder-decoder pipeline. Our framework is characterized by encoding both deep visual and semantic embeddings through a triple-branch network (TriNet) during the encoding phase. The visual attention branch captures attended visual embeddings from medical images with the soft-attention mechanism. The medical report (MeRP) embedding branch predicts semantic report embeddings. The embedding branch of medical subject headings (MeSH) obtains semantic embeddings of related medical tags as complementary information. Then, outputs of these branches are fused and fed into a decoder for the report generation. Experimental results on two benchmark datasets have demonstrated the excellent performance of our method. Related codes are available at https://github.com/yangyan22/Medical-Report-Generation-TriNet .
科研通智能强力驱动
Strongly Powered by AbleSci AI