判决
计算机科学
嵌入
人工智能
图像(数学)
自然语言处理
模式识别(心理学)
匹配(统计)
代表(政治)
数学
政治学
政治
统计
法学
作者
Lin Ma,Wenhao Jiang,Zequn Jie,Yu‐Gang Jiang,Wei Liu
出处
期刊:IEEE Transactions on Circuits and Systems for Video Technology
[Institute of Electrical and Electronics Engineers]
日期:2019-01-01
卷期号:: 1-1
被引量:25
标识
DOI:10.1109/tcsvt.2019.2916167
摘要
In this paper, we propose a novel multimodal matching model for the image and sentence based on their multiple representations. Each representation of the image or sentence undergoes an independent neural network, consisting of multiple layers of nonlinear mappings to yield the corresponding embedding. Besides exploiting the image and sentence relationship based on their embeddings, we propose one novel loss to further exploit the relationship within each single modality, namely, image and sentence based on the yielded multiple embeddings, which is used to train the neural networks simultaneously. The experimental results demonstrate that multiple representations can help to capture the image contents and the sentence semantic meaning more precisely, thus making comprehensive exploitations of the complicated image and sentence matching relationship. More concretely, the proposed matching model significantly outperforms the state-of-the-art approaches in bidirectional image-sentence retrieval on the Flickr8K, Flickr30K, and Microsoft COCO datasets.
科研通智能强力驱动
Strongly Powered by AbleSci AI